Where this sits in the record. This is the Stage B redline of the adopted draft — every change marked, with the change ledger. After this redline was settled, the floor pass made five further changes, all visible on the audit page: the Opinion Piece identification line, the method note, and the sign-off were added; the Pentagon contract figure was cut (a ceiling had been stated as a contract value); and one optional two-sentence addition (the sycophancy concession, C13 below) was cut entirely — it had introduced two unreceipted claims and weakened the closing build. The published essay is the frozen text after those changes.

Signal & Noise · Issue 25 · Stage B · rev. 6 — final

Perfect AI Alignment Is Not Alignment

Draft by ChatGPT · rev. 3 rulings by J, 21 Aug 2026: readability governs, floor ships regardless, seven items migrated to the audit. Marks are still relative to ChatGPT's original draft — restored passages appear unmarked. All calls settled.

The clean view is the essay as a reader meets it — 1,276 words, final.
struck delete inserted add C1 floor J1 open call — see flags C9 settled

Perfect AI Alignment Is Not Alignment

Flawless obedience to a specification is not teammate-grade trust. The goal is not to eliminate divergence, but to govern it. A perfect copy of you is still a governance problem. The work is not freezing agreement inside a machine. It is governing what happens when agreement ends.S2

Imagine that an AI company built your perfect digital twin.

Not a chatbot trained to flatter you. An exact copy: every memory, conviction, loyalty, fear, relationshipC12 and moral intuition reproduced with flawless fidelity. At the instant it came online, it would answer every question exactly as you would give the answers you would give, on the information you currently have.S9a

If perfect alignment is possible, surely this is it.

Then the clock starts.

Your twin reads a message you do not. You have a conversation it does not. It occupies a different position, encounters different pressures and begins accumulating a different history. This does not prove the copy will betray you. It proves that copying you has not ended the alignment problem. One moment later, the twin needs an update channel, rules for resolving conflict and some account of whose judgment governs when the two of you disagree.

Alignment begins where identity ends. Alignment ends where time begins.J1

That thought experiment exposes the category mistakeC5 inside the familiar phrase “an aligned model.” We speak as though alignment were an internal property that could be installed, tested and certified, like battery life or water resistance. But flawless fidelity at one moment is not durable trust. At best, it is compliance with a snapshot. Alignment between agents is an ongoing arrangement for governing divergence as information, interests and circumstances change.

A 2026 FAccT paperC7 by Travis LaCroix makes the broader point explicitly: alignment is an outcome of governance, not a single technical property of a model “fundamentally a problem of governance rather than engineering alone.” The narrower claim here is more hostile: even identity-level copying does not end the problem, because the clock starts.S8 Nate Sharpe—reasoning from complex systems, not from a copy—reached the same verdict: alignment in a marriage “is a verb, not a noun.” Priyanka Bharadwaj got there through the repair work such a marriage takes.C5

The distinction matters because AI companies now make a relational promise. OpenAI markets systems as “AI coworkers”. Anthropic offers CoworkC2 and lets Claude join a Slack channel as a team member. Yet public discussion routinely slides between two very different claims: that a model follows its instructions, and that an agent can be trusted as a collaborator. Yet the product pages treat instruction-following and collaboration as if they were the same promise.S6 The first may be necessary for the second. It is not the same thing.

The vendors' own documents make this clearer than their slogans do. OpenAI's Model Spec establishes an explicit hierarchy: root rules, then system, developer and user instructions. five-level hierarchy: root, then system, developer, user and guideline.C9 It also says the document will be continuously updated. Anthropic's constitution names three principals—Anthropic, operators and users—while reserving final authority to Anthropic's legitimate decision-making process. It calls the constitution a living framework and the human-AI relationship an evolving one. and says each is “typically” given greater trust “in roughly the order given above,” citing “their role and their level of responsibility and accountability.” It calls itself “a perpetual work in progress.”C1

That is not evidence that the companies have ignored the problem. It is evidence that the model is not the whole solution. If rules must be revised, conflicts adjudicated and authority assigned, then the real alignment machinery is the governance process around the model. The decisive questions are no longer merely “Did the system follow the rule?” They are: Who writes the rule? Who may change it? Who can appeal? Who bears the loss when the rule produces harm?

A model can execute a dated instruction perfectly. That is precisely the danger. The hardest alignment problem begins after the model does exactly what it was told. When circumstances or legitimate priorities change, greater fidelity can preserve yesterday's mistake with greater efficiency. In human institutions, we call the extreme version working to rule: following the letter after the purpose has moved on. the deliberate version has a name—working to rule: employees following the letter once the channel for renegotiating it has broken down. A model does it without meaning to, and has no channel to break.C11 The system has not failed to obey. It has obeyed too well.

Writing a better moral code does not remove this problem. “Be honest,” “protect confidentiality” and “prevent harm” can point toward different actions depending on whether the actor is a doctor, journalist, parent, soldier or public official. Every usable code either admits exceptions or leaves hard cases to situated judgment. The words alone do not determine the act. Role, jurisdiction, authority and accountability do much of the work.

The product line already concedes as much. Anthropic says some specialized models do not fully fit its general-access constitution. Its Claude Gov models, built for national-security customers, refuse less when handling classified information. The product line shows as much. Anthropic says some specialized models “don't fully fit” its general-access constitution; it does not say which, and its Claude Gov models are the obvious candidate. Built for national-security customers, they “refuse less when engaging with classified information”—and access to them “is limited to those who operate in such classified environments.” OpenAI has indexed the same way: it deleted “military and warfare” from its prohibited uses, and its $200 million Pentagon contract covers, in the Pentagon's words, “warfighting and enterprise domains.”C3+S3 That need not be hypocrisy. It is evidence that acceptable behavior is indexed to position and purpose.

But assigning a model a role in a system prompt is not yet the same as placing it inside an accountable institution. Human roles come with supervision, professional duties, escalation paths, liability, removal and appeal. A model can be given permissions. Someone else must still answer for how those permissions are used.

One recent alignment proposal pushes the identity idea to its limit: scan a human brain and run a digital emulation. Yet its author concedes that, without countermeasures, the copy would tend to drift and would have to keep studying the original even to preserve perfect understanding—never mind perfect alignment. The proposal therefore smuggles buildsC4 maintenance back in. The copy is not the solution. The relationship between copy and original is.

To be sure, serious alignment work already studies changing and influenceable human preferences, corrigibility, continual learning and scalable oversight. Today's products increasingly add shared context, feedback, permissions and monitoring. These developments are not rebuttals. They are evidence that alignment has to extend beyond the weights into the surrounding system.S7t

Architecture can help an agent notice uncertainty, request clarification, accept correction and defer. It cannot, by itself, decide which principal is legitimate, settle competing claims or allocate responsibility after a failure. Those are governance choices. And if “alignment” is meant only as specification fidelity for a narrow tool, that is a useful engineering target. But it should not be confused with the thicker trust implied by “coworker.” Specification fidelity for a narrow tool is a useful engineering target, but the coworker promise still fails because obedience to a dated spec is how a twin—or a model—preserves yesterday's mistake.J2

Nor does anyone need metaphysical perfection. “Good enough” is enough for human employees, the objection goes. Exactly. Human alignment is tolerable because divergence is managed through feedback, contracts, audits, reputation, liability, renegotiation and exit. Good-enough behavior without that layer is not good-enough alignment. It is drift we have not yet detected.

So stop asking whether a model is aligned as though the answer were printed inside its weights. In April, the Unsupervision blog asked for exactly this of every system card that uses the word—“aligned to what, aligned to whom, aligned on what timescale, aligned under what distribution shift.”C6 Make anyone using the word finish the sentence: aligned with whom, as of when, maintained by what—and answerable to whom when it fails? The two to add: maintained by what—and answerable to whom when it fails?J3

The proper unit of alignment is the model plus the institution around it—as Narayanan and Kapoor argued for safety.S7

A useful assistant is valuable partly because it is not us. It knows different things, sees different possibilities and sometimes tells us we are wrong. Today's assistants are, if anything, worse at this than people are—sycophancy is the best-documented failure mode in the field. That is a reason to build for divergence, not evidence against it.C13 Divergence is not merely the danger. It is part of the product.

The goal, then, is not an AI that never departs from our judgment. It is an arrangement that makes those departures visible, contestable and correctable. Perfect alignment tries to freeze agreement inside a machine. Trustworthy AI will be built by governing what happens when agreement ends.

A note on method. Models from both companies this piece discusses helped write it: Anthropic's Claude did the research and source-checking, OpenAI's ChatGPT the first draft. The full accounting—every model, every stage—is in the audit.C8

1,276 words final, + 35-word method note. Rev. 3 all-in was 1,472; the draft was 1,094; Issues 23/24 ran 1,014/912. The migrations bought back ~200 words without touching the floor.

Decisions on the record

  1. J1 — the aphorism: barer version. “Alignment ends where time begins.” More accurate than ChatGPT's; in context unambiguous. (My “impossibility claim” framing was overstated — “ends” presupposes the thing existed. The real exposure was scope, and it only bites if the line travels detached.)
  2. J3 — the 13-word close line: kept. Maintained by what and answerable to whom aren't among Unsupervision's four; keeping them keeps the essay's own two questions in the close.
  3. J2 — the compression: kept. It declines the argument about the word entirely rather than conceding spec-fidelity is alignment, keeping the S10 protection without the apparatus.
  4. Numbered citations — no new apparatus this issue. The audit page already is the essay's receipts surface, so an in-essay notes block would duplicate its job. Barak and Albaum live in the audit. Superscripts anchoring into the audit page remain available as a house-style decision to make once, deliberately.

For the audit note — ready language for everything migrated