2.3·What evaluation needs

What does meaningful AI evaluation require?

Across this section we see that some things change and some stay the same, depending on the choices made. Here we summarise what we found, and what we think is needed to evaluate AI-assisted evidence.

Where are the insights stable, and where do they change?

Stable

Some broad patterns held

  • Main country differences across several analytical changes
  • Country patterns retained using one combined model instead of three separate models
  • Same five categories in each country’s top five under different counting methods

Sensitive

More precise results moved

  • Changes to exact shares, rankings and margins
  • Larger effects from removing a source type than removing documents at random in 10 of 12 tests
  • Sharp fall in England’s Inspection and Accountability share without education media
  • Some topics merged when the model changed
  • More Scottish and Irish text left without a topic
  • Changes in category size under different topic groupings
  • Scotland’s Equality, Rights and Inclusion still prominent, but substantially smaller under one reasonable alternative

Not supported

One claim disappeared

  • Teaching Profession and Workforce as a clear Irish distinction
  • Shared Irish category absent under one reasonable alternative grouping
  • Ireland and Scotland gap reduced to 0.3 percentage points among documents used to build the models
  • Same underlying evidence, but a different comparison supported by it

Some stability is document length. A document's category is a vote across its passages, so long reports hold still and short articles move.

The evaluation chain

StageWhat evaluation needs to ask
1Question and intended useWhat are we asking, why are we asking it and what decision might the answer support?
2Evidence baseWhat was included, what was left out, whose records shaped the analysis and what does the collection represent?
3Analysis and assessabilityHow did the model, categories and counting rules transform the evidence? Can those choices be inspected and tested?
4Interpretation and useWhat claim does the result support? Does it make sense in context, and is it adequate for the proposed use?
5Consequence and correctionWho is affected, who can challenge the finding and what happens when a challenge succeeds?

What else should evaluation cover?

What to take from this

Some findings held, some moved and some were withdrawn. The lesson is not which model to trust. It is that evaluation must follow the whole chain from question and evidence to use, consequence and correction.

What this shows

What changed when we tested the analysis, which choices caused the change, and what evaluation has to cover as a result.

How to read it

Broad patterns and precise results behave differently. A pattern can survive a change that moves its size, its ranking or what it can be said to mean.

What this cannot tell us

Whether the categories mean the same thing across countries, whether the collected records represent the wider debates, or whether a finding is adequate for a particular use.

Why this matters

Evaluating the model alone is not enough. A finding is produced through a chain of decisions, and following that chain takes technical expertise, people who know the system and people affected by it. It needs a visible record of the choices made, tested against reasonable alternatives, and a route to challenge and correct the result.