When AI Agreement Isn’t Evidence
Several models can agree because they were given the same hint. A practical way to get separate answers, compare them and check their claims.
While building the AI Scenario Explorer, one model produced a neat framework for tracing what AI systems do: creation, access, trust, portability, query rights and actionability.
All six components were already in its prompt.
The answer developed the framework, but that could not count as independent discovery. Comparing the output with its prompt exposed the problem.
Two later runs, using different models and prompts without those hints, did not reproduce the same framework. Both still discussed related problems. Because the models and wording changed together, this wasn’t a controlled test of the hint’s effect.
The original mistake was clear enough: an idea supplied in the question had been counted as a finding in the answer.
That risk survives a second opinion. If another model sees the same leading prompt—or reads the first answer before attempting its own—it may repeat the idea without providing much additional support.
For the Scenario Explorer, the workflow now separates those jobs:
- Get separate attempts first. Give models the question without the proposed answer, earlier drafts or the project’s favourite explanations.
- Then bring in an informed critic. Let it see the drafts and wider project so it can challenge assumptions, spot omissions and identify repetition.
- Compare the explanations. Keep the original responses. Look for whether they describe the same cause and consequence, and preserve disagreements.
- Check against the world. Where an argument depends on factual claims, check sources or seek relevant expertise. Agreement between models cannot do that job by itself.
This still leaves plenty of room for error. Different models can share assumptions and repeat the same mistake. A model comparing their answers can invent agreement that isn’t there.
The scenarios are possible futures to examine, not forecasts. Using several models helps generate and challenge ideas; it does not establish their likelihood.
Before counting agreement as support, ask what each model saw—and whether the answer was already in the question.