2.1·Change the data

Does the finding change when the sources change?

AI can analyse thousands of documents, but only the evidence it is given. A large dataset can still present a partial picture, and in many AI systems the evidence behind an output may not be visible. Here we test one important source of uncertainty: does the finding change when a source group is removed?

How the removal test worked

The three collections contain different mixes of publishers. Media make up 54% of the England collection but only 3% of Scotland’s, while government documents make up 37% of Scotland’s and 13% of England’s.

To test whether this matters, we removed one source type at a time (media, government, unions or civil society) and recalculated the issue shares.

Nothing was retrained. The model, topic assignments, expert labels, shared categories and counting rules stayed fixed. Only the evidence included in the calculation changed. This isolates the effect of changing the evidence base.

Because removing more documents naturally creates more movement, we compared each source removal with 2,000 random removals of the same number of documents. This gives a baseline for how much change we would expect simply from removing that amount of evidence.

The same principle applies more broadly to AI-assisted analysis: changing the documents, data or retrieved context can change the result even when the model stays the same.

A separate test could rerun or retrain the model on a different evidence base. That would test a different question, because both the evidence and the model output could change.

See what happens if we remove a source

As published With media removed Change

What do you make of these two figures?

What to take from this

The mix of sources matters, not just the number of documents. In ten of the twelve tests, removing one source type changed the result more than removing the same number of documents at random.

What this shows

How issue shares change when one source type is removed, compared with an equal-sized random removal.

How to read it

England’s Inspection and Accountability category provides the clearest example. Its share falls sharply when education media are removed, by more than we would expect from removing the same number of documents at random. Media documents are also the shortest, and short documents move most easily.

What this cannot tell us

Whether the collections represent the national debates as a whole. Evidence that was never collected cannot be tested through removal.

Why this matters

Treat these as patterns in the documents collected, not as direct measures of national priorities, importance or need.