Issue 23 · Published audit

AI Can Hallucinate a Jury

The audit ran before publication, on an essay about staged criticism. This page is the record: which objections were raised and what happened to each, which changes were adopted, cut, or declined, what could not be checked at all — and where this audit misreported itself before it was corrected.

Published audit

The essay is the judgment. This page is the inspection surface.

Canonicality: The essay Signal & Noise stands behind is Issue 23, AI Can Hallucinate a Jury, published at Signal & Noise. This page is its published audit, not a second edition of the essay. This is the first issue produced and audited under the streamlined pipeline adopted August 8, 2026 (constitution v1.3 will record the amendments).

How to read every disposition on this page: verifiers were instructed to attempt refutation and to default to refuted when uncertain — a deliberate anti-false-positive setting. “Refuted” here means this verifier, told to lean that way, did not sustain the objection. It never means the objection is false. No tally of refutations appears on this page, because a tally would be exactly the survival badge this essay argues against and this publication’s rules forbid. Three of the nine verified objections were both raised and adjudicated inside the same model family that designed this audit — a closed loop, and the sharpest limitation on this page.

The redline

Every adopted change in place — deletions struck, additions marked, and the three additions the author accepted as valid but cut from the essay, shown where they would have sat.

Read the redline →

An external critique, verbatim

The strongest outside leg’s full 1,362-word return, unedited — including the objections this audit refuted and the ones it conceded.

Read the return →

The conversation behind this

The actual author + AI conversation that produced this issue — who brought what: the embedded earlier exchange, the discarded first essay, and the author’s two-line pushback that forced the published thesis.

Read it →

Binding floor rows

A floor failure blocks publication unless the author explicitly overrides it with ownership of the risk, on the record. Two were found on the text; both were fixed. One closes at packaging.

F1 — unquantified frequency (resolved)

  • “Many essays receive no such prompt resistance” asserted an empirical rate with no support. The published text reads “An essay about policy, culture or the future may meet no such resistance at all” — frequency converted to modality, nothing else moved.

F2 — opinion identification (closes at publication)

  • The piece must be identified as opinion wherever a reader arrives, with a live link to this audit. Because the audit ran pre-publication, the status line reads “Audit complete” from day one.

F3 — attribution exceeded its anchor (resolved; reversed an earlier pass)

  • The essay’s Guardian citation supports conjecture, criticism and error-correction — but its text nowhere carries the clause “not through authority or proof of certainty.” Zero occurrences of “authorit-” on the cited page. The floor pass had scored this citation as faithful, because it checked the proposition as a whole rather than each clause against its anchor. The audit caught it; the clause was deleted — its substance survives in the adjacent interview citation (“we don’t judge theories by their source”) and in the essay’s own text. Recorded as a reversal, not a silent correction, and the floor procedure now checks clause by clause.

Objections raised, and what happened to each

Nine objections were classed as potential kills — defects that would mislead a careful reader — and each went to an independent fresh-context verifier with no edit authority, under the refute-by-default setting above. The full reasoning on both sides is preserved; compressed dispositions:

  1. “The central premise is insinuated, never owned.” Refuted: the scenario is explicitly one model wearing personas, so “overlapping training, shared defaults” is definitional there; the wider empirical extension is hedged where it appears. A residue survives: the premise’s available support is stronger than the prose suggests (see the declined-changes note on grounding).
  2. “The thesis dichotomy is self-contradicted” — if checked objections earn confidence, staging did lower confidence’s cost. Refuted: the essay forecloses the imported premise (criticism exposes errors; it does not issue certificates), so checked objections improve an essay without touching the null-result inference the subtitle prices. Raised independently by two model families.
  3. “The prescription dissolves the alarm.” Refuted: a remedy that would work if adopted does not contradict the claim that the unremedied practice misleads. Residue: why the danger persists for readers who never see a transcript — recorded among the open proposals.
  4. “The closing standard is itself a fakeable process signal” — the strongest objection anyone has raised against this essay. Contested, and published as unresolved. One leg confirmed it; the verifier refuted the confirmation (a specifics-report publishes checkable content — the difference between a certificate and a record); the confirming leg named two residuals the defense does not cover: in verdict-sparse domains readers often cannot cheaply re-run claimed checks, and an honest report of correlated, shallow scrutiny still satisfies the standard’s letter. The tiebreak was the refute-by-default rule, so this page does not call it closed. The essay’s adopted response: the standard raises the cost of fakery; it does not eliminate it.
  5. “Scarce human time is smuggled in as epistemic evidence.” Refuted: the objection applies the essay’s found-objection standard to its negative-evidence question, where search extent is genuinely evidentially relevant; the carrying clause is “outside the author’s control” — selection, not cost.
  6. “The essay mistakes a governance choice for an intrinsic property of AI.” Refuted: the essay claims no intrinsic limitation, and external administration fixes selection effects but cannot manufacture independent error distributions from one model family.
  7. “The boundary is drawn on the wrong axis” (tool-grounded vs. language-only, not synthetic vs. human). Refuted: the criticized sentence is the tool-grounded/language-only distinction, and the essay twice repudiates the synthetic-vs-human axis it is accused of using.
  8. “The ending substitutes rhetoric for a standard.” Refuted: “certificate” has a defined antecedent (the detached “survived” badge), and the closing couplet preserves the findings while discarding the badge. Two of the critic’s counterexamples the essay already credits; two (benchmark results, independent audits) have residual purchase.
  9. “The essay is unfalsifiable — its hedges protect it from any test.” Refuted on a factual error: the load-bearing claims carry no hedge, most hedges present run against the essay’s own thesis (concessions, not armor), and the central claims are negative universals refutable by a single demonstrated counterexample.

Raised and dispositioned without verification routing (recorded so the list is complete): independence-as-continuum — raised by both external legs, conceded in substance, held as an open proposal; “fake institution is loaded language” — recorded, no change; critic-overfitting lacks an operational criterion — correct, fix held as an open proposal; the prescription’s counting rule is under-specified — a real gap, no cheap remedy found; a prior equivocation charge against the subtitle — re-examined fresh and refuted; and the charge that the essay proposes a disclosure standard it does not meet — discharged by this page’s existence.

A convergence worth more than any single objection: the claim that the essay compares AI review at its most gameable against human review at its best was raised independently by all three critique legs, was never refuted, and was partially conceded — the author accepted the counterfactual sentence as true and ruled it into this audit rather than the essay (below). Objections that independent readers keep reaching are evidence about how the essay reads, whatever their logical merit.

Change rulings

Changed in the essay: the two floor fixes; “Readers are left to lean on process signals” (behavior claim recast as structural); “Burn the verdict” (was “certificate” — no friendly fire on the essay’s own reporting standard); the positioning paragraph (“This is not an outsider’s complaint… a correction of a habit, not a discovery” — the staged panel in the essay’s opening describes this publication’s own documented practice, and an earlier issue praised what survived such review); the cost-raising close (“A report like that can be faked too — but faking specifics means publishing claims a reader can check”); and one author-written bridge sentence welding the confession to the prescription.

Accepted as valid — published here, not in the essay. The author’s ruling: these three additions are correct, and they cost more than they pay — nuance that drains energy and focus from the essay’s main points, especially for a first-time reader. The essay is deliberately lossy compression; this page is where the cut nuance lives. The author accepts all three as true qualifications of the argument:

  • “For most writing, meanwhile, the realistic alternative was never a jury at all. It was nothing — and against nothing, even correlated criticism is a gain.” (The most-corroborated finding of this audit.)
  • “The two failures compound: selection poisons the survival inference even when the critics are independent, and correlation keeps the poisoning invisible.”
  • “The claim does not require the critics to be identical — only that their blind spots overlap enough that ten silences carry much less than ten tests’ worth of information. Show a persona panel whose misses are as uncorrelated as separate reviewers’, and this worry dissolves.”

Declined or deferred, with reasons: grounding the correlation premise in measured results — declined for now, to keep a conceptual essay from needing its own literature review; register and reader-experience polish — deferred unless a cold human read still flags it after the substance fixes; and a standing rule applied throughout: no change that adds a new empirical claim without a receipt. Eight further minor proposals were neither adopted nor declined and remain open in the internal record.

What could not be checked

The audit’s largest gap — and the one that matters

  • The central premise was never tested. The disconfirming experiment is fully specified: hold compute constant, seed blinded flaws, randomize essays across one-pass, repeated-same-model, same-model-persona, cross-family-with-tools, human, and hybrid review; preregister prompts and stopping rules; publish every output; measure detection rates, false positives, calibration, residual-error correlation and held-out performance — and separately randomize how the review process is described to readers. None of it was run. Every objection to the premise was adjudicated by argument. An audit that dismisses objections to a claim it never tested has established nothing about that claim.

Also unchecked or unresolved

  • Objection 4 — whether the essay’s closing standard escapes its own argument — is unresolved on the merits; the tiebreak was a rule biased toward refutation.
  • The reader-behavior claim: no case, survey, or experiment locates the deceived reader the essay posits.
  • No quantitative threshold at which the essay would concede a multi-model, tool-grounded panel is independent enough to license confidence.
  • Calibration coverage: disconfirmers, a prediction and alternatives were produced for the central claim only; four further hypothesis-typed claims are typed but uncalibrated.
  • The standing residual: this publication’s record shows a defect class — register, cognitive load, how a live page actually lands — that machine checks have not caught and human cold reads have. Nothing here covers it.

Calibration of the central claim

The typing pass scanned every load-bearing sentence for two hard-gate conditions — narrative confidence exceeding evidentiary confidence in a way that would mislead, and claims written as explanatory truth that nothing could falsify. A hit would have been a sentence asserting an empirical rate or mechanism with certainty and no support. The pass returned no hits; the one objection challenging that result was adjudicated under the refute-by-default setting and the zero inherits that bias.

The central claim (N personas of one model do not provide N independent tests): disconfirmer — the planted-error experiment above; if persona-panel misses prove statistically uncorrelated and union detection approaches the independence prediction, the claim is wrong. Prediction, roughly a year’s horizon — essays “certified” by single-model multi-persona review will yield additional errors to cross-model or human holdout review at a materially higher rate than their transcripts imply; testable on any such corpus, including this publication’s own archive. Alternative explanations for why multi-persona review feels convincing — transcript surface features (length, distinct voices, concession-and-verdict structure) may drive conviction regardless of error coverage; and the review does find real, fixable errors, so genuine local improvement is misread as global vindication.

What ran, and what a reader can inspect

Five audit legs from the Claude model family, via API — the same family that operates this publication’s editorial process and designed this audit: mechanical and anchor verification (10 findings), an adversarial battery (9), a cold-reader pass (11), claim typing and calibration (25), and prior-findings adjudication (3). Two external legs, each given the essay and five adversarial questions only — not told whose essay it was, for what publication, or what any other critic had found: deepseek-v4-flash via OpenRouter (6 findings) and gpt-5.6-sol via the author’s own provider account (12 objections; published verbatim). All verification legs were Claude instances. Both cited sources were re-fetched live at pass time and read against the sentences they anchor.

A disclosed failure and its repair: the gpt-5.6-sol leg failed twice on August 8 — identical transport timeouts — and was not substituted; the failure was recorded instead. Re-run at the author’s request on August 9 with an identical prompt, it completed, and produced the sharpest set of objections in the audit. Both the failure and the re-run are recorded because a leg that silently succeeds on a later attempt is indistinguishable, in the artifact, from one that never failed. Note the asymmetry this leaves: the essay was drafted in the author’s conversation with the GPT family, and the GPT family also supplied the strongest critique leg.

Inspection limits, stated rather than implied: the external return above is published in full. The Claude legs’ raw transcripts and the verifiers’ reasoning exist as internal artifacts and are summarized on this page but not published, so those legs are taken on this page’s assertion rather than open to outside inspection. That is a real gap against this publication’s own auditability standard, named here rather than glossed.

This page’s own corrections

The first draft of this audit page was checked against its source material before publication and misreported its own record in the direction that flattered the audit. It claimed a thrice-raised objection had “failed on the merits” when that objection was never verified and had in fact been partially conceded; it stated that objections came from outside the model family that built the audit, when three were raised and adjudicated inside it; it dropped four of the external leg’s twelve objections without disposition; it printed refutation tallies in bold and closed with a self-awarded conformance badge; it reported one transport failure where there had been two; and it filed a publication-blocking floor row under a heading that declared nothing binding. Each is corrected above.

That matters more than any finding on this page. An audit page written by the process it audits will drift toward flattering that process, and this one did — while explicitly warning against exactly that. The check that caught it was mechanical: read the page against the raw returns and count what is missing. That check is now a standing part of the audit stage.

What this audit did not do: test the essay’s central claim. What it structurally cannot see: how the published page reads to a human encountering it cold. What remains unresolved: whether the essay’s closing standard escapes its own argument. Per that standard — use the criticism; burn the verdict — this page offers no verdict to burn.