Published audit
The essay is the judgment. This page is the inspection surface.
Canonicality: The essay Signal & Noise stands behind is Issue 25, Perfect AI Alignment Is Not Alignment, published at Signal & Noise. This page is its published audit, not a second edition of the essay. High-stakes designation, author, 22 August 2026: the piece makes adverse characterizations of Anthropic and OpenAI drawn from those companies’ own published documents, and the process defaults pieces that audit someone’s claims to a pre-publication audit. Under that designation the audit’s findings arrived while the text could still change — and the author’s ruling on the standing objections is recorded below.
How to read every disposition on this page — including a fault this audit found in itself: objections were first sent to a fresh-context cross-model verifier told to attempt refutation and default to refuted when uncertain. That pass returned every objection refuted — including two kill-class objections that both critique models had raised independently. A sweep of that shape was rejected as a reading of the instrument, not of the essay: the refutations were substantive but shared one move (“the essay anticipates this elsewhere”), and against a heavily hedged essay a refute-by-default verifier holding the full text can answer almost anything that way. Every objection was then re-judged against a different question: would a fair, informed, skeptical reader land this objection reading the essay once, front to back — with answers assembled from distant paragraphs not counting? The labels below come from that corrected pass. LANDS means the objection stands and is published here. Did not land means the cross-model judge concluded a linear reader would not sustain it — never that the objection is false. No tally of refutations appears on this page.
The redline
Every change between the adopted draft and the frozen text, marked in place with its reason — including the changes the floor pass forced and the ones the author declined.
Both critique returns, in full
Issue 24 published one leg and summarized the rest, and named that as an inspection limit. This issue closes it: both legs verbatim — every Battery answer, every objection as raised, both steelmen.
The conversation behind this
The full development conversation — from the first thesis impulse through Stage A, the redline rulings, the floor pass, the pre-publication audit, and publication. Five long third-party pastes are summarized in place; distribution work after the publication report is outside the record.
The machine’s version
An experiment, ruled on by the author: the essay the AI model in this issue’s editorial process wrote with a free hand, after the essay and this audit were complete — published here with its own floor pass and its own pre-publication audit. Not a critique of the essay above, and not a correction. The author’s essay remains the one Signal & Noise stands behind.
Binding floor rows
Nine rows were raised against the adopted text. Every row was remedied in the text before freezing except one; no author override was exercised.
Remedied in the frozen text
- A ceiling stated as a contract value. “Its $200 million Pentagon contract” — the $200M is a ceiling (OpenAI’s own words), with $1.9M reported obligated in FY2025. The author cut the figure; the quoted contract scope carries the point without it.
- Two unreceipted claims in one optional sentence. A late addition called sycophancy “the best-documented failure mode in the field” and today’s assistants “worse at this than people are.” Neither had a receipt, and the sentence sat inside the essay’s closing build. The author cut it entirely. The addition had been proposed by the AI process, not the author — recorded as a process lesson: a defensive insertion can cost more than the objection it pre-empts.
- Identity block. The adopted draft had no Opinion Piece identification line, no ownership statement, and named three AI vendors against the standing unnamed-AI rule. The author ruled the naming in under the Issue 21 exception — when the conflict is specific, a reader cannot weigh it unnamed — inside a method note that also carries the ownership sentence: the human author owns every claim, word and error on the page.
Carried here, unresolved
- One quotation rests on convergent secondary reporting. “Warfighting and enterprise domains” is the Pentagon contract announcement’s language as quoted identically by four independent trade outlets; the defense.gov release itself was not reached by any tool in this run. Blocked for a tool, not for a person — and blocked is
null, not verified. If the primary surfaces and disagrees, the correction belongs on this page and in the essay.
Objections raised, and what happened to each
Two critique legs, given only the essay and the five Core Adversarial Battery questions, raised seventeen objections between them. Both legs independently converged on the same central pair of kill-class objections. Under the corrected reader test, fourteen stand and three did not land. The full text of every objection as raised is on the returns page.
The four central objections — one problem, four faces. All LAND.
- The thesis is stated universally, then narrowed without owning it. (kill) The essay concedes spec-fidelity is “a useful engineering target” for narrow tools, which makes the headline either false over its stated scope or in need of a restriction the text performs but never states. The false-or-trivial dilemma arrives mid-essay and is answered only by a pivot.
- Necessity is argued; decisiveness is asserted. (kill) The essay establishes that governance is necessary (the twin, working-to-rule, the vendors’ own revision machinery) and then claims the governance process is “the real alignment machinery” — a dominance nothing on the page ranks or measures.
- “We speak as though…” sweeps in people who don’t. (major) The unattributed “we” covers a research community whose work — cited by the essay itself — already treats alignment as processual. The essay’s real target is vendor marketing and casual discourse; the sentence doesn’t say so.
- The concession paragraph undercuts the novelty claim where a skeptic looks hardest. (major) “To be sure, serious alignment work already studies…” concedes the field studies exactly what the essay presents as its correction, then dismisses the tension with “These developments are not rebuttals” — asserted at the point it most needs arguing.
Ten further objections LAND. Compressed; full text on the returns page.
- Divergence-over-time proves too much. (major) Every persisting agent diverges, including the reader; the criterion alone cannot rank systems. The essay’s discriminator (managed vs unmanaged divergence) exists but arrives late.
- The twin is specified to need what it lacks. (major) Copy the dispositions that maintain relationships — deference, update-seeking, conflict-flagging — and the twin arrives with its own update channel; the conclusion is partly built into the premise.
- The stateless-model gap. (major) The twin’s divergence engine is accumulating history; most deployed models are stateless across sessions. The essay’s other divergence source (the world moves while the spec stands) does the work, but the reader must supply that repair.
- The product-page conflation is characterized, not quoted. (major) “The product pages treat instruction-following and collaboration as if they were the same promise” is an interpretive bridge between two documented halves; no page is quoted equating the two.
- No documented failure case. (major) The essay predicts a failure mode — harm from perfect obedience to a dated spec — and cites no instance of it. A conceptual claim by genre, but the absence is real and a skeptic notices.
- “The hardest alignment problem” is asserted, not ranked. (minor) Competing candidates (deception, multi-principal conflict, misuse) are never compared.
- The close can absorb any outcome. (minor) “Trustworthy AI will be built by governing what happens when agreement ends” can read successes as governance working and failures as governance lacking; no stopping rule is given.
- “No channel to break” vs the channels later inventoried. (minor) The essay lists shared context, feedback, permissions and monitoring four paragraphs after saying a model has no channel; the reconciliation (channel-grade vs institution-grade) is available but unstated.
- Contestability has no capture analysis. (minor) “Visible, contestable, correctable” relocates power to whoever runs the appeal channel; who writes the rule is asked, never answered.
- The conflicted-verification residue. (minor) The quotations load-bearing for the Anthropic portrait were checked by Anthropic’s own model. Disclosed in the essay’s method note and priced below — and the objection is right that disclosure is not resolution.
Three objections did not land.
- Working-to-rule imports intent. The judge found the essay flags the disanalogy at the point of use (“a model does it without meaning to, and has no channel to break”) — the reader is told exactly what does not transfer.
- The Claude Gov passage insinuates concealment. The judge found the essay links Anthropic’s own public announcement, marks its inference as inference (“it does not say which… the obvious candidate”), quotes the access limitation, and disclaims hypocrisy in the next sentence.
- “Drift we have not yet detected” contradicts “divergence is the product.” The judge found the managed/unmanaged distinction, made two sentences earlier, resolves it at the point of reading.
Author ruling, 22 August 2026 — the opinion-scope ruling. All fourteen standing objections publish here; none edits the essay. The author’s reasoning, on the record: an opinion piece is called an opinion piece because its central claim reaches farther than citations can support — if it did not, the writer would not be saying anything worth reading. The four central objections are, in the author’s words, excellent and they belong in the audit. This is the same lossy-compression ruling exercised on Issue 23: the essay stays sharp, and this page is where the price of that sharpness is visible.
Change rulings
Changed before freezing: the floor remedies above (a figure cut, a two-sentence addition cut, the identity block added), plus the Stage B revisions shown on the redline. Because the audit ran pre-publication, every change is in the frozen text; nothing was mutated after publication.
Proposed by this audit and staged: none. The author pre-ruled the standing objections to this page rather than to the text. The four central objections are thesis-level: any future response to them is Stage A material for another piece, not a correction to this one.
What could not be checked
- The Pentagon announcement primary. Four convergent trade outlets, zero direct fetch of the defense.gov release. Carried as
null. - One critique leg’s maker is undisclosed. Ox Alpha is a stealth model; OpenRouter names no lab. Published forensics put it at ~0.98 confidence of being a GLM-5.3 variant (Zhipu) — which is why GLM 5.3 itself was not used as the second leg. If the forensics are wrong in a specific direction (the provider is actually one of the four companies in this essay’s pipeline or pages), the independence claim below degrades. Judged unlikely; not checkable today.
- Both critique legs ran over one transport (OpenRouter). A transport-level fault would have correlated both legs. Disclosed because this process has seen exactly that failure once before.
- The standing residual: register, cognitive load, and how the published page lands for a human reading it cold. This run read the text and a locally rendered page, not the live site through a reader’s eyes.
- “Did not land” is not “false.” Three objections carry that label on one cross-model judgment under one test. They are preserved in full on the returns page.
Calibration of the central claim
The load-bearing claim was typed as a category judgment, not an empirical rate: installed fidelity to a specification and maintained agreement between agents are different kinds of things, and the second — which the “coworker” promise invokes — requires governance that no architectural property supplies by itself.
Disconfirmers: a deployed system whose spec-fidelity alone sustains teammate-grade trust across shifting circumstances and principals, without update machinery, adjudication, or a party bearing losses; or vendor governance documents converging toward less revision machinery as capability rises.
Prediction (roughly a year): vendor specs and constitutions keep their revision cadence or accelerate it, and authority hierarchies inside them deepen rather than dissolve. Checkable against the same documents the essay quotes.
Alternatives, from the two steelmen (both published in full): (1) technical alignment names robust intent-tracking under distribution shift — if achieved, governance becomes optional rather than constitutive, and the essay has refuted only a snapshot-certification position nobody defends; (2) the field already studies governed change under alignment’s own name, so the essay’s correction is to marketing language, and the title claims a larger target than the argument engages.
What ran, and the full method accounting
The essay’s method note ends: “the full accounting — every model, every stage — is in the audit.” This is that accounting.
- Stage A (research, prior-art sweeps, source verification): Anthropic’s Claude. Anthropic is discussed in the essay.
- First draft: OpenAI’s ChatGPT. OpenAI is discussed in the essay.
- Revision adjudication: xAI’s Grok, which ruled on the Stage B findings; the author then re-ruled several of its calls.
- Stage B floor pass and this audit’s orchestration: Anthropic’s Claude.
- Stage C critique: two models via OpenRouter API —
stealth/ox-alphaandmoonshotai/kimi-k3, both at reasoning effortmax. All three upstream vendors were excluded from Stage C; neither Stage C vendor is discussed in the essay. Themaxsetting was verified applied on Ox Alpha by probe, not assumed: an invalid value returns the accepted enum (seven levels), and low-vs-max runs separate cleanly on reasoning length (medians 229 vs 852 characters, no overlap, n=3 per side) — the usage field that reports reasoning tokens reads zero at every level and is a metering artifact. - Author’s relationship to every vendor named: paid subscriptions only.
- Links: all fifteen in the essay resolve for a reader; four block scripted access and were confirmed live in a browser.
The sharpest limitation, priced: the quotations that carry the essay’s portrait of Anthropic were verified by Anthropic’s own model, and the revision was adjudicated by a competitor of both companies discussed. The critique legs were chosen to be outside that loop; the verification was not. Each Anthropic quotation was fetched at least twice in separate sessions, and every quotation is one click from its source in the essay — a reader can check the primary without trusting any model in this pipeline. The conflict is disclosed, not resolved.
Credits and differentiations the essay carries here
An opinion essay cannot credit everything it stands near. The author ruled these to the audit rather than the text; they are debts, recorded.
- IEEE Spectrum, “Perfectly Aligning AI’s Values With Humanity’s Is Impossible” (Choi, May 2026): argues impossibility from computability and proposes an ecosystem of mutually constraining agents. This essay makes neither claim — not impossibility, not multi-agent ecology; its claim is that even a perfect copy still needs an update rule. The remedies converge superficially; the routes and claims differ.
- Boaz Barak, “Machines of Faithful Obedience” (Jun 2025) defends faithful obedience as the right technical target; Shelly Albaum, “Alignment Is Not Obedience” (May 2025) argues the reverse. The essay grants the narrow tool and denies the coworker promise; the obedience debate itself is that genre’s, not this essay’s.
- Roger Dearnaley, “3. Uploading” (Nov 2023): the nearest prior “even a copy fails” — but grounded laterally (humans are not aligned with each other), where this essay’s twin fails temporally (your own copy diverges from you). Adjacent, not the same claim.
- Narayanan & Kapoor, “AI safety is not a model property” (Mar 2024): credited in the essay; the differentiation lives here — their piece treats the system at release, and the temporal dimension is undeveloped there. This essay’s addition is the clock.
- The unstated brace: identity is the price of perfect alignment, not of genuine alignment — a good deputy aligns through complementary skills and shared stakes, not similarity. The twin is a stress test, not a demand.
- The good-enough layer, named: products now add shared context, feedback, permissions, monitoring — some of the management layer. What they still lack: standing to reopen the spec, and a party who bears the loss when it was wrong. Telemetry is not governance.
- Further debts, waived to this page: Gabriel and Korinek–Balwit (the principal question, formalized 2020/2022); Hadfield-Menell & Hadfield (incomplete contracting — the institutional anchor for “maintained by what”); Nossel and BISI (the two-tier product-line observation, filed there as hypocrisy); Zhi-Xuan et al. (the role apparatus, argued with opposite valence); Millidge (the obedience genre’s second name).
- The world-as-updating-context-window image was cut from the draft; the underlying proposition is closest to Spizzirri’s “Specification Trap” (arXiv 2512.03048), credited here.
This page’s own check
Before this page published it was checked mechanically against the raw returns: every objection above traces to a finding id in a critique return; every LANDS / did-not-land label traces to a reader-test verdict; the floor items trace to the floor pass; no survival count, badge, or validation language appears. The first verification pass’s all-refuted sweep is reported above rather than silently replaced by the pass that superseded it. If a later check finds this page flattering the audit, the correction belongs here, not in a silent rewrite.
What this audit did not do: reach the Pentagon announcement primary, identify one critique leg’s maker, or read the published page as a human cold reader. What it found and kept: fourteen standing objections, the two sharpest of which say the essay claims more decisiveness and more novelty than its citations establish — published here at the author’s ruling, priced rather than patched. This page offers no verdict.