Across this section we see that some things change and some stay the same, depending on the choices made. Here we summarise what we found, and what we think is needed to evaluate AI-assisted evidence.
Stable
Sensitive
Not supported
Some stability is document length. A document's category is a vote across its passages, so long reports hold still and short articles move.
| Stage | What evaluation needs to ask |
|---|---|
| 1Question and intended use | What are we asking, why are we asking it and what decision might the answer support? |
| 2Evidence base | What was included, what was left out, whose records shaped the analysis and what does the collection represent? |
| 3Analysis and assessability | How did the model, categories and counting rules transform the evidence? Can those choices be inspected and tested? |
| 4Interpretation and use | What claim does the result support? Does it make sense in context, and is it adequate for the proposed use? |
| 5Consequence and correction | Who is affected, who can challenge the finding and what happens when a challenge succeeds? |
Some findings held, some moved and some were withdrawn. The lesson is not which model to trust. It is that evaluation must follow the whole chain from question and evidence to use, consequence and correction.
What changed when we tested the analysis, which choices caused the change, and what evaluation has to cover as a result.
Broad patterns and precise results behave differently. A pattern can survive a change that moves its size, its ranking or what it can be said to mean.
Whether the categories mean the same thing across countries, whether the collected records represent the wider debates, or whether a finding is adequate for a particular use.
Evaluating the model alone is not enough. A finding is produced through a chain of decisions, and following that chain takes technical expertise, people who know the system and people affected by it. It needs a visible record of the choices made, tested against reasonable alternatives, and a route to challenge and correct the result.