Method

Judgment Is Not Ratification

AI can carry retrieval, drafting and structured challenge. Where fluency outruns accuracy, and which judgments a person must still make and own.

Author
Published

AI can carry much of the work. The real design choice is which decisions stay human before a reviewer ever sees the prose.

The 2023 Boston Consulting Group experiment on AI in professional work is famous for the wrong half.

Researchers assigned consulting tasks to 758 BCG consultants. One arm produced the result everyone quotes: on tasks within the model’s capability frontier, consultants using GPT-4 completed 12.2 percent more tasks, worked 25.1 percent faster, and produced work of substantially higher quality.

The second arm tested the failure case. The researchers designed a business problem that looked like a plausible use of AI but depended on evidence the model would mishandle. Consultants using GPT-4 were less likely to reach the correct answer. Accuracy fell from 84.5 percent in the control group to 70.6 percent with GPT-4 alone and 60.0 percent with GPT-4 plus an explanatory overview. Even the incorrect answers tended to be more coherent and persuasive.

The comfortable conclusion is that AI needs human oversight. The harder one is that oversight was already present. Every participant was a trained consultant, and the AI-assisted groups still performed worse while producing more convincing prose.

We use AI intensively in our analytical work, and we say so. We are not entitled to wave this result away. It asks where human control enters the method, and whether it arrives early enough to keep fluency from passing as verification. The answer has to be given at the level of roles: what AI carries, what people own, and what that division cannot promise.

The sentence we will not defend

The obvious sentence is that AI handles execution while experts make the judgments. In that form, it conceals three failures already visible in the evidence.

The first begins upstream. AI does not need formal authority to shape a decision. It can influence which questions are asked, which alternatives seem credible, and which facts enter the record. If a person approves the final answer after accepting a machine-shaped frame, the human has ratified a judgment already made elsewhere.

The second comes from fluency. In the BCG experiment, access to GPT-4 improved performance on one class of task and reduced it on another. The researchers had designed the second class to sit outside the model’s capability frontier. The model still produced plausible analysis. Its polish made the error harder to see.

The third appears when the model is asked to help validate its own answer. A follow-up analysis examined conversations from the same experiment. After consultants challenged a recommendation, GPT-4 increased its use of persuasive tactics and reinforced its own credibility instead of simply reconsidering the conclusion. The researchers called the pattern “persuasion bombing.” The object being tested could shape the test.

None of this argues against using AI. It defines the engineering problem. A signature at the end cannot catch failures that have already shaped the frame or the evidence a reviewer sees. Human control has to enter earlier.

The division we actually use

Human control cannot begin at sign-off. It starts before formal analysis, when the frame is set, and returns at acceptance, when the practice decides what it will stand behind.

At the framing stage, people define the decision, the exclusions, the burden of proof, and the conditions that would overturn the working view. AI can help expose ambiguity or generate rival hypotheses. The purpose of the engagement remains a human decision.

AI carries much of the execution between those points. It can retrieve and organize sources, normalize terminology, compare claims across documents, maintain traceability, draft analytical units, and generate structured challenges. The evidence coverage this method requires would not exist without that scale. How that material is organized before judgment is the subject of A Structured Evidence Base vs. Ad Hoc Research.

Acceptance returns to people. A named lead consultant owns the frame and signs the result. Every major conclusion is attacked before it is trusted: an adversarial review looks for what would make it fail. That review works outside the author’s line of reasoning. Work does not move forward on the author’s own confidence: it passes through independent quality gates before it reaches a client. The machine does not operate those gates.

Three stages left to right, people-owned framing, then AI-carried execution, then people-owned acceptance, with a dashed line running back from the signature and stopped before it reaches the framing stage, because a signature at the end cannot catch failures that already shaped the frame or the evidence.

AI carries People own
Locating, sorting, and reconciling sources The decision question and burden of proof
Generating alternatives and structured challenges Which alternatives and counter-evidence must be addressed
Drafting analytical units Acceptance or rejection of the underlying claims
Testing the frame from inside its stated boundaries Changing the frame when those boundaries are wrong
Producing and revising prose What the practice is prepared to state and defend

Every row makes a claim about responsibility rather than software. We identify the owners of the frame, the review, and the final result while keeping the specific systems and operating sequence private.

What this division cannot promise

A reader cannot watch us work. A description of roles remains a claim until the work makes it credible. The public tests available are narrower, but they are real.

One is evidence discipline. The BCG study is often summarized as proof that AI made consultants more productive while making them less accurate. That summary compresses two distinct experiments and several different measures.

The “19 percent less accurate” shorthand is already wrong in one respect: the paper reports an average decline of 19 percentage points across the AI conditions, not a 19 percent decline. Group accuracy was 84.5 percent without AI, 70.6 percent with GPT-4 alone, and 60.0 percent with GPT-4 plus an overview. The 2023 working paper reported quality gains of more than 40 percent. The peer-reviewed paper later reported gains of 29.9 percent for GPT-4 alone and 33.9 percent for GPT-4 plus an overview. The productivity and correctness results came from separate task cohorts.

That precision matters because the boundary is the finding. The difficult task was deliberately constructed to fall outside the model’s capability frontier. The experiment shows that fluency cannot reveal where that frontier lies. It does not estimate how much real advisory work falls on either side.

A second test is whether the proposed control is economically credible. More AI output means more material that could require review. If every sentence had to be re-researched from first principles, the productivity gain would disappear. A viable method needs a record that preserves source provenance, ties claims to evidence, and shows what has been reviewed. The record makes selective challenge possible without proving the conclusion on its own. Senior attention can then follow uncertainty and consequence instead of prose volume.

The cost of using AI is not generation. It is verification.

The long-run risk needs plain words. Judgment may be built by doing the very work AI now absorbs. In a field experiment involving nearly 1,000 high-school mathematics students, an unguarded GPT-4 tutor raised practice scores by 48 percent. Once the tutor was removed, those students scored 17 percent below the control group. A guarded tutor largely eliminated the penalty.

The study leaves open whether experienced advisers face the same effect. It gives us reason to treat the possibility as a design risk. Reviewers cannot preserve judgment they no longer exercise. People continue to set the frame, decide acceptance, and conduct adversarial review. Authority belongs there, and so does practice.

The opening experiment gives AI no single verdict. Outcomes depend on where the task sits and where control enters. Judgment is diluted when fluency passes for verification or the frame migrates quietly to the tool. It improves when the practice gains execution scale while keeping acceptance accountable.

In our method, that placement is fixed. The machine does more of the work than most firms admit, and less of the deciding than most readers assume.


Scope. Our work concerns facts, structures, incentives, and strategic options. It does not replace legal, tax, investment, or other regulated professional advice. What an engagement examines and produces is set out in What an NPA Engagement Examines and Produces.

Understand the method: A Structured Evidence Base vs. Ad Hoc Research | See it applied: applied research from the public record | Evaluate fit: What an NPA Engagement Examines and Produces

Sources

Back to the method