A Structured Evidence Base vs. Ad Hoc Research
When searching, reading and concluding run together, claims lose their sources and coherence becomes the stopping rule. The case for keeping them apart, with limits.
Somewhere between the fifth browser tab and the second page of notes, a quiet substitution happens. “The company says retention improved” becomes “retention improved.” Nobody decides this. It is what note-taking does under time pressure: the speaker drops out, the claim stays, and by the time the notes become a memo, the claim has been promoted to a fact that no one remembers admitting.
Multiply that substitution across three weeks and forty sources, and you get a document with a specific pathology. It reads well. It is coherent, plausibly argued, and impossible to audit. Ask which of its statements rest on the company’s own account, which on an independent record, and which on someone else’s interpretation, and the document cannot answer. The distinctions were never recorded; they dissolved at the moment of reading.
That dissolution has a cost, and diligence is not the variable that controls it. What matters is what happens to judgment when searching, reading, interpreting, and concluding all occur in one continuous motion, with no stable record of claims, sources, contradictions, and gaps in between. We organize our own analytical work around the opposite premise: before any judgment is formed, raw materials are rebuilt into a structured evidence base that keeps sources, claims, and gaps separately visible. The premise has reasons, evidence, and limits, and all three belong on the table.
What fused research does to judgment
The failure modes of ad hoc research are not folklore. Several are well-established findings in the study of human judgment, and their limits matter as much as their force.
The first is that an early hypothesis shapes what gets searched next. In Peter Wason’s classic rule-discovery experiments, participants tended to propose tests that could confirm their current guess rather than tests that could eliminate it. Later work by Klayman and Ha showed that this positive test strategy is not irrational in every environment, but it can fail to distinguish a favored explanation from plausible alternatives. Applied to research, the danger is mundane: once you suspect a competitor is struggling, your next searches instinctively look for evidence of distress. When the first plausible explanation also writes the next search query, a sequence of individually reasonable searches turns into a test of one hypothesis.
The second is stranger and closer to the bone. Holyoak and Simon had participants evaluate arguments in isolation, then again inside an ambiguous case as they moved toward a verdict. The ratings shifted. Items already read became stronger or weaker depending on the conclusion taking shape, before the decision was final. They called it spreading coherence: an emerging answer quietly re-weights the evidence that produced it. Imagine evaluating a potential acquisition, where a flattened growth curve begins to look like a temporary market adjustment as you lean toward a “buy” recommendation. In a fused workflow, there is no record of what a source looked like before the story arrived, so the shift is invisible even in principle.
The third concerns provenance. People are demonstrably poor at remembering where information came from, and repeated statements feel truer than novel ones even when the reader knows better. The public-information environment compounds the problem mechanically: one announcement becomes a wire story, three rewrites, and a dozen aggregator pages. Read casually, that feels like ten sources agreeing. Traced, it is one evidentiary origin and nine echoes. Pages are not evidentiary origins. A researcher who counts them instead of tracing them is not corroborating facts; they are measuring the efficiency of a distribution network.
The last is the stopping rule. Every research process has one, stated or not, and in ad hoc work the implicit rule tends to be stopping when the account feels coherent. The diagnostic-error literature in medicine gave this a name, premature closure, and found it among the common cognitive contributors to missed diagnoses. Coherence is a property of stories, not a property of evidence. A workflow that cannot tell the two apart will end its searches at exactly the wrong moments.
None of these studies tested an end-to-end research workflow, and it would overstate the record to claim that structured research eliminates these tendencies. The narrower claim is enough: when collection and interpretation are fused, these distortions become harder to detect because the workflow leaves little trace of when they occurred.
Rebuilding materials into evidence
The alternative is not more diligence. It is a different object. Before analysis begins, we rebuild raw materials into a structured evidence base: a representation in which every claim remains attached to whoever made it, support and contradiction have separate places, and a missing piece of evidence is recorded as missing rather than papered over by volume elsewhere.
The load-bearing choice inside that rebuild is a three-way separation. Everything that enters the base is held in one of three registers:
- Self-description: What the subject says about itself is real evidence about what the subject wants salient, even when it is weak evidence about whether the underlying claim is true.
- Verifiable record: This establishes what can be independently supported in the record. It does not by itself tell you how anyone reads those facts.
- External reception: This records how outside parties publicly describe, question, or act on the subject. It shows which readings become observable in the external record without assuming access to the audience’s underlying beliefs. A perception can be factually wrong yet institutionally decisive.
Collapse them into one narrative and each question contaminates the others; keep them apart and their disagreements become findings instead of noise.
No field we know of uses exactly this triad as a named standard. It is our synthesis, made deliberately, because each seam in it is borrowed from a discipline that learned the hard way. Auditing standards insist that management’s assertions are inputs to be tested, never self-validating proof. Historians’ source criticism treats a document as evidence of two different things at once: the event it describes and the author’s decision to describe it in a particular way. Qualitative researchers who triangulate across sources are taught that contradiction between perspectives is not a defect to clean up but frequently the phenomenon itself. The triad packages those three lessons into a single working structure.

The same information, run both ways
To see what the structure buys, take one concrete scenario: investigating an AI infrastructure startup. You have forty pieces of information drawn from company PR, public filings, media coverage, GitHub commits, and former employee commentary. Nothing about the pool changes in what follows. Only the handling does.
Run ad hoc, the pool becomes a memo, and the failure modes happen naturally along the way. The first plausible theory shapes later searches. The CEO’s claims are copied without attribution. Supporting material accumulates while a dissenting engineer’s comment is absorbed into a narrative transition. Repeated reporting from tech blogs is treated as corroboration because nobody traced the origins. Notes are progressively rewritten as the story firms up, so earlier readings disappear. And the work stops when the narrative feels complete.
Run through the evidence base, the exact same pool of information comes out shaped differently: perhaps a dozen attributed claims in the startup’s own voice; nine independently supported facts whose apparent fifteen sources collapse to three real origins; a handful of reception items tagged to the voices driving them; two contradictions left standing in plain view; and five explicit gaps.
The memo version is easier to read. The structured version is easier to challenge. It can answer questions because it preserved exactly the distinctions the memo dissolved: who said what, what survives verification, which readings appear in the external record, what conflicts, and what is simply not known.
The judgment built on the structured base is not automatically correct. It is inspectable, which means its errors are findable.
Structure at the scale of an engagement
A single evidence base is the unit. The same logic extends to the whole working process. Every engagement runs through five stages: materials, evidence, confirmation, analysis, and delivery.
Client materials and public materials enter through separate channels and land in one traceable package, so a claim’s route into the work is always reconstructable. The confirmation stage exists because some claims matter enough that analysis should not rest on them until the client has confirmed or corrected them. Analysis begins only after the base is sufficiently stable. This is a sequencing discipline rather than a bureaucratic one; it is where the first half of this argument becomes practice.
Within that pipeline, roles are deliberately kept separate. Strategic judgment, material execution, external research, and independent review are different jobs with different failure modes, and combining them recreates the fused workflow at a larger scale. Where automated tools help in this division, and where judgment stays human, is taken up in Judgment Is Not Ratification. How this analytical work sits alongside communications advisors is mapped in Narrative Architecture, Messaging, PR and IR: Different Jobs at Different Moments.
Nothing moves forward on the author’s own confidence alone; review stands between every conclusion and the client. That review is independent in the operative sense, conducted by someone standing outside the author’s chain of reasoning, because spreading coherence is one reason authors should not be the sole auditors of conclusions they helped form. What an engagement examines and what it produces, stage by stage, is described in What an NPA Engagement Examines and Produces.
What structure does not do
It would be convenient to claim that all of this makes judgment accurate. The evidence does not support that claim, and the shortfalls need stating plainly.
Structure is not neutral. A taxonomy can encode the wrong problem, and a formal representation can create false exhaustiveness; in classic fault-tree experiments, people failed to notice what a tidy diagram omitted. Structured analytic techniques do not automatically improve accuracy. Intelligence agencies institutionalized hypothesis matrices decades ago, and the experimental record on Analysis of Competing Hypotheses is mixed, with some studies finding that visible procedure raised analysts’ confidence without raising their accuracy. A completed template is often just evidence that a template was completed.
It can also be performed ceremonially: a base that counts correlated sources as independent ones will manufacture false confidence more efficiently than sloppy notes ever could. That is why source lineage is a first-class field in our practice, not an annotation.
None of this is free, and the cost is not always worth paying. A tiny, direct evidence set behind a reversible decision does not need this machinery. The value rises with the stakes and the mess: consequential judgments, large and heterogeneous sources, multiple live explanations, evidence that will be challenged, and work that changes hands. Strategic questions about cross-border perception sit firmly at that end of the range, which is why we treat the structure as the default there and not everywhere.
And in strategic work, ultimate accuracy may be unobservable. A transaction happens or does not for many reasons; an audience’s underlying beliefs may never surface. What remains assessable, even then, is the epistemic conduct of the work: whether claims were attributed, lineage preserved, contrary evidence retained, uncertainty stated at its actual size, and the reasoning left reconstructable by someone who was not in the room. What public information can and cannot establish, and how we handle that boundary, is treated at length in What Public Information Can and Cannot Establish About Narrative Risk.
A structured evidence base does not make judgment objective. It keeps the objects of judgment from collapsing into one another too early: a source stays attached to its claim, an assertion stays distinct from its verification, a perception stays distinct from the fact it may or may not describe. Contradiction stays visible. Absence stays absence.
None of that guarantees a better conclusion. It makes known routes to a worse one harder to travel unnoticed, and in work where being wrong is expensive and being unauditable is worse, that is the difference that compounds.
Understand the method: Judgment Is Not Ratification | See it applied: Why a Cross-Border Valuation Discount Can Persist After the Numbers Improve | Evaluate fit: What an NPA Engagement Examines and Produces
Sources and further reading
- Wason, P. C., “On the Failure to Eliminate Hypotheses in a Conceptual Task” (1960)
- Klayman, J. and Ha, Y.-W., “Confirmation, Disconfirmation, and Information in Hypothesis Testing” (1987)
- Holyoak, K. and Simon, D., “Bidirectional Reasoning in Decision Making by Constraint Satisfaction” (1999)
- Johnson, M., Hashtroudi, S., and Lindsay, D. S., “Source Monitoring” (1993)
- Fazio, L. et al., “Knowledge Does Not Protect Against Illusory Truth” (2015)
- Graber, M., Franklin, N., and Gordon, R., “Diagnostic Error in Internal Medicine” (2005)
- Heuer, R. J., Psychology of Intelligence Analysis (1999)
- Wilcox, J. and Mandel, D., “Critical Review of the Analysis of Competing Hypotheses Technique” (2024)
- Fischhoff, B., Slovic, P., and Lichtenstein, S., “Fault Trees: Sensitivity of Estimated Failure Probabilities to Problem Representation” (1978)
- Mathison, S., “Why Triangulate?” (1988)
- PCAOB, Auditing Standard 1105, Audit Evidence
- Kahneman, D. and Klein, G., “Conditions for Intuitive Expertise: A Failure to Disagree” (2009)