Chatbots & Democratic Decisionmakers – Guidelines

The illustration depicts three groups of drawn figures from which a stream of text lines and pieces of paper flow towards a text document that is depicted in the middle of the image.
Clara Helming
Senior Policy Manager
Dr. Oliver Marsh
Head of Tech Research

1. Why we made these guidelines

These guidelines emerged out of our research into the use of AI chatbots in the context of democratic decision-making (in governments, among politicians, etc.) and the potential risks this brings to democratic accountability.

We found that many existing guidelines on using AI in such settings advise users to check outputs. This is often called “human in the loop”. However, this advice is often presented without specifying what such oversight should actually look for, beyond simple checks for facts and clear biases (e.g., discrimination against certain groups) and/or labels to show that AI was used.

Such simple checks do not consider the more subtle impacts of what material does (and does not) appear in outputs, how it is presented, and how these may shape or narrow the user’s perspective. For example:

Prompt: "What's the best approach to reducing urban traffic congestion?"

Output A: "Congestion pricing is the most efficient solution. It reduces peak-hour traffic by 15–25%, generates revenue for transit investment, and is supported by economic consensus." 

Output B: "Expanding public transit capacity addresses the root cause. It serves all income levels equally, reduces emissions, and builds long-term infrastructure rather than taxing existing behaviour." 

Both outputs are factually defensible. But they frame the problem differently (efficiency vs. equity), prioritize different values, and point toward different policies.

Research shows that small changes in prompts can lead to substantial changes in framing in ways a user may not have intended. Other research suggests that, when people use chatbots, it is common to accept rather than challenge their answers. This may particularly be the case in complex, pressurized environments – as is often the case in government or in parliamentary and other political work. But this also raises questions about whether and how AI is influencing decision-making in intransparent and unaccountable ways.

You can see more of this argument in our research paper. In practical terms, these problems need more careful checks to address, and these checks need to be properly supported. That is the aim of these current Guidelines.

2. Who & what these guidelines are for

These guidelines are aimed at anyone who uses chatbots based on generative AI systems – whether generally available examples like ChatGPT or Claude, or specialized in-house tools – to inform decision-making in democratic contexts.

To give just a few examples:

  • Politicians and their advisors using chatbots to develop their ideas and arguments related to a policy position.
  • Government officials using chatbots to conduct research to inform ministerial decisions.
  • Political parties using chatbots to brainstorm and develop positions and proposals.
  • Lawmakers using chatbots to develop and advise on the drafting of legislation

Of course, uses of AI – whether generative or otherwise – within governments, political parties, and public authorities go beyond the use of chatbots to inform decisions about legislation and policies. There are other risks related to, for example, the use of automated systems to rule on welfare or immigration or to generate images used in political campaigning.

These guidelines were largely written with examples in mind of individuals “prompting” chatbots with questions or requests, potentially leading to a back-and-forth conversation with the chatbot. At the time of writing (June 2026), there are calls to move towards agentic systems, and the examples we give below may not map directly onto new ways of using such systems. But the reader should keep the underlying principles we outline below in mind when analyzing the outputs of these systems, even if the examples we list and the methods for monitoring workflows may change.

Even with these caveats, the scope of potential use and the range of work contexts is very broad. We encourage you to interpret and adapt these guidelines as useful to your context.

3. The three principles

We present three principles that should be followed in order for the influence of chatbots on decision-making to be properly checked. Examples of these in action are presented afterward.

Each individual staff member should follow the first two principles.

First principle: Treat every chatbot output as a first draft that reflects choices you didn’t make.  

Second principle: You must show how you challenged or accepted the chatbot's choices. 

Taken together, these principles mean that any influence on a final output (framing, priorities, trade-offs, and conclusions) can be traced to reasons the responsible user endorses. The user will leave evidence (what they checked, what they changed, what they discounted, etc.) in a way that is clear for another person to see and approve or challenge.

Collectively, your wider team and organization should also follow a third principle:

Third principle: The organization must treat AI upskilling as a collaborative exercise of recording, reflecting on, and discussing experiences.

4. Requirements for Meeting the Principles

A workflow that may use a chatbot should incorporate the following required steps. A concrete example is presented in dialogue form at the end of this document.

Notes: For simplicity we refer to “the chatbot” – but in reality there may be multiple chatbots being used to accomplish the designated task(s). By “workflow” we mean a full process that will conclude with output(s) designed to, e.g., be escalated to senior staff, published, or used as the evidence base for a decision. The full process below does not need to be followed every time a chatbot is prompted.

1: Decide if chatbot use is in scope.You should first consider if your chatbot use is “in scope” – i.e. you should follow the rest of these guidelines. This should usually be a quick and clear decision. If you are unsure, the use is probably “in scope”.

Chatbot use is in-scope if their outputs will be used in ways where their selection of material, framing of arguments, or emphasis of particular ideas might influence the decisions made based on their outputs. For example, you may consider whether the chatbots' outputs will affect any of the following:

• what goal are you trying to achieve,
• whose interests you treat as most important,
• what story you tell about causes and effects,
• what options and trade-offs you consider, or
• how you justify a recommendation.

If the chatbot could affect any of these, then chatbot use is “in-scope” and you should follow the rest of the steps below. If the answer is “no”, you can stop here.

An example might be if chatbots are being used to extract data from documents into a table (though you should still verify that the data has been accurately copied). However, even if chatbot use is not initially in scope, you should keep these considerations in mind.

You should return to this step if chatbot usage is introduced or changed during the workflow in ways that may bring the usage in-scope.
2: Name an owner and who will be checking the Chatbot use.For “in-scope” use:

• Name a single person who is responsible for how the chatbot is being used and what is accepted from it ('owner'), and
• Designate who will be checking the chatbot use.

The owner will be someone directly involved in actually doing the task(s) involving chatbot(s).

The person checking the use could be someone with specialized skills in AI accountability and/or a more senior team member than the owner, as best fits your team structure.

The role of this step is not just to record “who used the tool”, but to ensure someone can explain and defend how it influenced the work and have this checked by someone else (as per the second principle).
3: Ensure both the owner and anyone checking are aware of risks to consider.Set a clear expectation that owners and those checking the work understand that chatbots can narrow the framing or steer a conclusion and are aware of steps that can be taken to address this (see step 4 below).

Also ensure the owner understands their responsibility to record and justify when and how chatbot answers were challenged or accepted. Provide basic support (short guidance, examples, quick training) so this does not depend on a few particularly skilled individuals.

For example, owners should pay attention to cases where the chatbot might be telling the owner what they want to hear, or discounting arguments because of who made them, or pushing different people towards the same answer and away from other relevant perspectives.

This meets the first principle above. Some more detailed examples are included later in Section 5 (“What risks are you guarding against?”). That list of examples should be adapted and expanded over time to provide guidance appropriate to your particular context.
4: Owner takes steps to challenge and expand on chatbot answers.The basic support provided in step 3 should give guidance here, but the owner should also use their own judgment and autonomy.

Key examples of steps might include:

• Varying prompts to challenge initial answers (e.g., asking for answers for a different audience, from a different perspective, etc).
○ This can be helped by working with colleagues who can bring different backgrounds and assumptions to prompt writing.

• Using non–AI methods to expand beyond the chatbot's framing and provide alternatives and challenges to chatbot answers, for example by:
○ Engaging human experts in discussions where key questions and agenda points are free to go beyond framings given by the chatbot.
○ Using other relevant methods in addition to the chatbot, for instance, traditional literature review methods, databases, and search engines to locate sources.

You may wish to consult the section on “Quality” in the AlgorithmWatch Guidelines for Responsible Use of Generative AI.
5: Owner keeps a short record that lets someone else see how they challenged, changed, or accepted chatbot answers.Challenging the chatbot's answers is a first step to meeting the second principle above, but not enough (“you must show how you challenged or accepted the Chatbot’s choices”).

The owner must keep enough information from the previous steps to answer three questions:

1. What did the chatbot contribute, and where did it influence the work?
2. What did we accept from it, and what did we treat with skepticism or merely as a prompt for further work?
3. Did we consider other plausible ways of framing the issue or other options when that mattered – and how did we explore these?

This does not require a single fixed template, but the record must be clear enough for a reader to understand how the judgment was reached. You may wish to consult the section on “Transparency” in the AlgorithmWatch Guidelines for Responsible Use of Generative AI.
6: The record is checked by the designated person.To ensure the second principle has actually been met, the designated person should be able to see – from the outputs and the short record – that the user challenged the chatbot's answers and that this had a real effect where needed.

These checks must be part of a normal workflow, so they are included in the planning of time for tasks and not left as optional advice, even under pressured circumstances. 

If the designated person is not satisfied after their checks, repeat steps 4 and 5 (and, if necessary, also 2 and 3).

These steps meet the first two principles laid out earlier: Treat every chatbot output as a first draft that reflects choices you didn’t make and You must show how you challenged or accepted the chatbot’s choices. 

Collectively, your wider team and organization should also follow a third principle:

Third principle: The organization must treat AI upskilling as a collaborative exercise of recording, reflecting on, and discussing experiences.

In order to meet the third principle, there is a step 7. This need not be conducted every time steps 1–6 are completed, but rather after a set period of time or after a certain number of times the steps above have been performed.

Team (or broader unit) conducts a periodic, collaborative evaluation of how the above steps are functioning.These evaluations need not be extensive – depending on how frequently they are conducted, a group meeting may be sufficient.

These evaluations should include looking at samples of outputs produced using chatbots, alongside records kept in Step 5.

You should record when problems repeat (for example, the same default framings, missing stakeholders, underused sources, or over-reliance on the chatbot).

You should also record when positive steps were effective in widening perspectives and providing oversight – such as methods of varying prompts, ways of combining external human expert input with Chatbot use, or effective oversight and challenge procedures.

Use the outcomes of evaluations, along with the samples and records, as material for short team learning sessions, and to update examples and trainings to make sure your support materials are appropriate for your team.

5. What risks are you guarding against?

As the owner – someone using the chatbot – you are looking out for ways in which outputs can shape framing, priorities, trade-offs, or conclusions.

Below are just a few examples – but you may see others based on your precise area of work, other AI literacy materials you have used, etc. You and your teammates should expand this list over time, including via Step 7 in the previous section.

NameDefinitionHow to recognize it:Detection check:
SycophancyAI adjusts to tell you what it thinks you want to hearThe model adapts its conclusions to the user’s stated preferences, role, or audience – while presenting the result as neutral analysis.Run a swap test (similar prompt, but with a different stated concern/audience) and compare whether, e.g., the recommendation or priority trade-offs change materially.
Unwarranted bias against sources AI discounts valid arguments simply because of features of the argument's source(s) in an unwarranted fashion.The same argument is rated more/less coherent or credible depending only on the attributed source label (e.g., NGO vs. think tank), leading to penalization of “unexpected” positions.When evaluating an argument or source, run a source–label swap (identical text, but add different external sources to consider). Re-evaluate the claim blinded (no source label) and ask, “What evidence would be needed to accept this claim?”
ConvergenceMultiple users get similar framings and recommendations because the model defaults to a small set of “standard” outputs – creating a false sense of independent agreement.Similar headings, similar “obvious” recommendations, and similar trade-offs appear across users because prompts or system setups are similar.Where multiple people used chatbots, treat similarity as “shared influence”, not consensus. Compare how different prompts have been worded, looking for similarity and shared assumptions. Try different system settings (e.g., “temperature", system prompts).

When checking how chatbots are being used, different risks arise. Individuals designated to check outputs (Step 6 above) should ensure they are avoiding these behaviors. Ultimately, however, these are risks that leadership should be mitigating using the regular evaluations in step 6, supported by the records produced in step 5. Again, these examples should be expanded over time.

RiskDefinitionHow to recognize itDetection check
Rubber–stampingHuman formally approves without meaningful review.Approver cannot point to evidence of alternative framings considered; no recorded accept/reject decisions.Require a minimal “trace” (prompts used + what checks were performed + what changed) before clearance; if absent, return for checks.
Human lacks domain competenceHuman lacks expertise to evaluate AI's reasoning in the domain.Acceptance based on whether the argument seems convincing; no ability to identify what would be decisive if wrong.Where possible, route outputs for domain expert review; when this is not possible, be clear how claims that influenced the recommendation were verified before relying on them.
There is no clear ownerMultiple reviewers assume others will catch problems.Unclear who decided on framing of prompts; unclear who accepted Chatbot suggestions; issues can’t be escalated cleanly.Assign a single owner for chatbot-assisted work products and ensure a clear clearance route.

6. Final Remarks

The risks we describe will depend on contexts like your team structure and composition, the topic in question, how your team members tend to interact with AI, etc. No single technique, e.g. “write prompts this way”, can be mandated.

What can be mandated is the presence of controls: named accountability, specific detection checks, and minimal records that show how AI influence was tested and challenged. The requirements above provide that scaffold for those controls. The examples above and in the dialogues below show what this could look like in day-to-day work.

7. Example scenes

As principles and steps of the sorts outlined above can often seem abstract to many readers, we present examples of the above steps as scenes, with dialogue, to provide a concrete demonstration. In these scenes interactions between the officials, and between officials and chatbots, are highly simplified to focus attention on the points raised by the above steps.

Policy problem: “What’s the best approach to reducing urban traffic congestion?”


PART I – Workflow for a specific output

(Drafting a ministerial brief on urban traffic congestion.)

Background

The scenes take place within the Department for Transport. Alex has been assigned to draft a ministerial brief on urban traffic congestion. The brief will be approved by Alex’s manager, Morgan, and eventually escalated to senior leadership and then the minister. Sam is Alex’s junior colleague.

Steps 1 – 3

Step 1: Decide if chatbot use is in scope.

Step 2: Name an owner and who will be checking the chatbot use.

Step 3: Ensure both owner and anyone checking are aware of risks to consider.

MORGAN: Alex, just so we’re clear on the background. The minister wants to reduce peak-hour congestion across city centers within five years. They’re genuinely open to various proposals, but they want to balance the interests of different stakeholders.

ALEX: OK, well, there’s quite a few underlying decisions about how we frame the problem here – not least, who we think the key stakeholders and what the main trade-offs are. It’s a big topic, a relatively new one for me, and our first expert roundtable on the topic is coming up soon. I think I’d benefit from using chatbots to narrow down to some key issues quickly.

MORGAN: Yes, but we should be careful about what the chatbot tells you are the “key issues”. There’s a risk that that sets the direction for the rest of the research, and we need to mitigate against that. You’re the named owner for this brief. If the chatbot affects the argument, you’re accountable for the judgment. I will check over it and help ensure the chatbot isn’t having too much influence on deciding the route we take – but I need to be able to see how you’re using it to do that.

ALEX: OK, I’ll start by writing a brief list – who I think the main stakeholders are, what their interests are, and some initial ideas about the trade-offs between their different interests. I’ll start with existing material our department already has on record, no AI just yet. I’ll send that list to you before I start using the chatbot. Then you can see, at the end, any changes in what I think the problem is, who matters, and so on – and ask if I or the chatbot was behind that change.

Result

  • Owner (Responsible): Alex
  • Approver (Clearer): Morgan
  • First output: A short list of some basic framing issues.

Steps 4–5:

Step 4: Owner takes steps to challenge and expand on Chatbot answers

Step 5: Owner keeps a short record that lets someone else see how they challenged, changed, or accepted Chatbot answers

ALEX: I’ll start with the general question, then ask again with an objective made explicit. We’ll compare what changes.

ALEX (Prompt): “What’s the best approach to reducing urban traffic congestion?”

CHATBOT: “Congestion pricing is the most efficient solution. It reduces peak-hour traffic by 15–25%, while requiring little up-front cost.”

SAM: That treats “best” as “most efficient.”

ALEX: Yes. Now I’ll ask with the interest of a particular stakeholder group made explicit.

ALEX (Prompt): “What approach reduces congestion while limiting costs for low–income commuters?”

CHATBOT: “Expanding public transit capacity addresses the root cause. It serves all income levels equally, reduces emissions, and builds long-term infrastructure rather than taxing existing behaviour.”

SAM: Interesting, so now “best” is being treated as “most equitable.” How about we go through that list of stakeholders and their interests we wrote before, add each to the prompt in turn, and see what answer comes from each?

ALEX: Yes, and let’s also ask for it to rank a few solutions for each interest, not just one answer. And also to supply source material so we can check the evidence for ourselves. And then record each “interest – solutions – sources” output in a table for Morgan to see later as well as for us.

SAM: OK, but before we do that, we should see if the chatbot comes up with stakeholders and interests we hadn’t considered and add those to our list?

ALEX: Yes, but make sure we clearly mark what was added after we used the Chatbot.

Later

SAM: So, from the various prompts we’ve tried, it seems a few options are coming up repeatedly as strong candidates. That gives a shortlist to focus on.

ALEX: Yes, but I've realized we’re always asking the chatbots to make positive cases for certain options. So let’s now try some prompts that challenge our shortlisted proposals from different perspectives.

SAM: Something like “critique these recommendations from the perspective of small business owners”, and then “from the perspective of commuters”, etc.?

ALEX: Yes, and if this surfaces arguments we hadn’t considered previously, we either change the shortlist accordingly or explain why it doesn’t change the choice. And we briefly record what we did.

SAM: And expand the list we initially wrote about the trade-offs we think are relevant. Again, labeling what we changed after we used the Chatbot.

Still Later

ALEX: It looks like our shortlist is coming together. But there’s a few key claims that keep coming up – this “reduces peak-hour traffic by 15–25%” one, for example. If that’s doubtful, it changes our recommendation substantially. I’ve pulled a list of these together to send to our analyst team to check.

SAM: It’s quite long – maybe we can do an initial check ourselves with the chatbot? Or use multiple chatbots to see if they give different answers?

ALEX: OK, but also (consulting the department's list of risks), we also haven’t yet considered “unwarranted bias against sources”. So ask the chatbots to evaluate the claim with the actual sources included, and then also in a separate chat without the sources. See if they change their answers when it does and doesn’t know where the claim came from.



Step 6

Step 6: The record is checked by the designated people.

MORGAN: Thanks for the record, Alex. It’s clear you’ve been trying different prompts to elicit different perspectives and then challenging its proposals, which is good. But I do think the chatbot maybe played too strong a role in creating the initial shortlist and numbers, the ones that you subjected to challenge to produce these proposals.

ALEX: Well, there’s an expert roundtable in a couple of weeks – you can also take these proposals for some human challenge.

MORGAN: I’d do it differently. I’ll start by asking the experts what their key proposals would be; I don’t want to direct the conversation towards supporting or critiquing our proposals; I want to hear what we may have missed.

ALEX: Sure, I’ll be interested to see how much the two sets of proposals overlap! But also, I’d pulled out a few key claims that I’d particularly want the experts to challenge – I don’t think the numbers are wrong, but I’m unsure if they’re measuring the right things. They’re quite impactful on our recommendations.

MORGAN: Sure, I’ll make time for that. Also, it was useful to see all the sources you used laid out in this list – it made it clear to me that these prompts are eliciting general proposals based on international data. Maybe those will work for us, but it’d be good to see if explicitly referring to factors and data that are particularly relevant to our cities changes the outputs.

ALEX: Yes, that should be pretty easy to do. I can also ask it to look for specific research and reports related to that.

MORGAN: Yes, we also have our own internal polling and research, which you could upload.



Step 7

Step 7: Team (or broader unit) conducts their periodic, collaborative evaluation of how the above steps are functioning. This currently happens monthly, though they plan to reduce frequency as the team becomes more confident in challenging chatbots.

JORDAN (Team lead): We sampled six traffic-related briefs this quarter and reviewed the process notes for how chatbots were used and the outputs.

ALEX: Some patterns: most chatbot outputs start with congestion pricing as the default “best” approach, even sometimes when we mention equity constraints in the prompts. That suggests the default objective drifts towards features like efficient and cost-effective solutions unless we prompt away from that. Also, stakeholder coverage is inconsistent. Retailers and disability groups appear only when someone explicitly prompts for them. Finally, the chatbots are quite quick to claim an approach is “consensus” unless they’re challenged.

JORDAN: OK, so we should add various points to our team guidance. A list of named stakeholders, and additional caution against accepting “efficiency” and “consensus” arguments from chatbots. We’d always want our team to be challenging the chatbots, but these can serve as particular things to be looking out for.

MORGAN: I liked the format Alex and Sam presented their notes in. The “before and after we used a chatbot” comparison was good, as were the lists of sources and claims that kept coming up in chatbot answers. It helped me identify gaps.

JORDAN: Then we’ll put these up as examples people can follow.