Issue 20 · Reference

Who Checks the AI in Your Medical Record?

What the sources changed, narrowed, or refused to support — and which parts of the landing remain argument.

CrosswalkStructure ↔ Reference

Reference

The late source check did not merely decorate the receipts. It corrected the essay.

The short version

The reference layer groups twenty-three public sources into eight dependencies: adoption and entry channels; calibration and evidence maturity; the Epic sepsis chronology; ambient-note and automation effects; alert behavior; patient attitudes and doctrine; contracts and liability; and current outside governance.

The final source-fidelity pass produced nine recorded corrections. Among them: the study-maturity categories were separated correctly (88.2% preclinical; 97.6% preclinical or early clinical; 2.4% randomized trials); alert-override ranges were corrected and scoped; the studied systems and hospital universe were stated accurately; settings and reader experience were narrowed; and legal doctrine was no longer summarized more broadly than its source allowed.

Those changes do not certify the landing. The sources establish adoption patterns, failures, behaviors, instruments, and limits. Assigning first-line checking to hospitals remains the essay’s argument.

What the references did

Changed

  • Separated the source’s study-maturity categories correctly: 88.2% preclinical; 97.6% preclinical or early clinical; 2.4% randomized trials.
  • Corrected alert-override ranges and kept them tied to the review designs that produced them.
  • Corrected the hospital and tool counts behind adoption claims rather than implying a broader universe.
  • Narrowed study-setting, reader-experience, and legal-doctrine language to what the original sources supported.

Narrowed

  • The Epic sequence is version- and site-specific; later local performance does not validate the earlier model.
  • Behavioral measures remain behavior: an override, signature, or unchanged draft does not reveal private trust or careful review.
  • Patient surveys show heterogeneous preferences, not consent in a specific encounter.
  • Federal rules and professional initiatives remain partial, status-specific backstops.

Killed or not imported

  • No universal claim that hospitals perform no local checking.
  • No universal error rate for ambient documentation or predictive tools.
  • No claim that existing doctrine cleanly assigns responsibility for every clinical-AI harm.
  • No inference that a complex development process proves the essay.

Source ledger

Open any entry to see the outside constraint and what it could change.

RL-020-01: Adoption and the EHR channel.

Why it matteredThe opening depends on clinical AI entering through institutional software and vendor channels, not on a fresh bedside choice.

Public sourcesASTP/ONC Data Brief No. 80; Yang and Graetz, AJMC (2026).

ScopeSupports broad hospital use of predictive AI in or through EHR systems and the rapid spread of ambient documentation in the studied Epic-hospital sample. Adoption does not establish local performance, safety, or earned reliance.

RL-020-02: Calibration drift and the maturity of clinical-AI evidence.

Why it matteredThe essay’s continuing-checking premise requires a documented reason performance can change after deployment and a sober account of how much evaluation has reached real clinical use.

Public sourcesGuo et al., Applied Clinical Informatics; npj Digital Medicine evidence-maturity review (2026).

ScopeSupports calibration as a distinct performance property that can degrade across settings or time, and the correctly separated maturity findings: 88.2% preclinical, 97.6% preclinical or early clinical, and 2.4% randomized trials. It does not imply that every deployed model drifts or fails.

RL-020-03: The Epic sepsis chronology: v1 validation, pause, and later local tuning.

Why it matteredThis is the essay’s concrete case for distinguishing installation from earned local reliance.

Public sourcesWong et al., JAMA Internal Medicine (2021); Wong et al., JAMA Network Open (2021); JAMA Network Open local evaluation of Epic sepsis model v2 (2026).

ScopeSupports poor external performance of the earlier model in one large health system, the Michigan pause described in the public record, and improved results after local tuning of a later version. Versions, hospitals, workflows, and endpoints differ; the sequence is not a single controlled comparison.

RL-020-04: Ambient-note errors and automation effects.

Why it matteredThe essay argues that a fluent interface can hide error and that observed use does not reveal how carefully a person reviewed the output.

Public sourcesMayo Clinic Proceedings: Digital Health (2025); Dratsch et al., Radiology; Chen et al., Lancet Digital Health (2024); medRxiv ambient-note preprint, version 2.

ScopeSupports documented factual errors in generated notes, automation effects in the studied imaging tasks, and the bounded unchanged-note figure reported in the preprint. The source check corrected study settings and reader experience; none of these studies supplies a universal error or review rate.

RL-020-05: Alert behavior and the TREWS counterexample.

Why it matteredThe essay needed both the long record of high alert overrides and a serious example where clinicians engaged with a sepsis alert in a supported workflow.

Public sourcesvan der Sijs et al., JAMIA; Poly et al., Journal of Medical Internet Research; Henry et al., Nature Medicine (TREWS).

ScopeSupports wide, context-dependent override rates in the legacy alert literature and the TREWS engagement counterexample: clinicians entered an evaluation for 89% of alerts and confirmed 38% of those evaluated. Overrides are not a direct measure of trust, and the engagement figures are not an outcomes claim.

RL-020-06: Patient attitudes, notice, and doctrinal limits.

Why it matteredThe patient section had to distinguish exposure from choice while accurately representing what surveys and legal analysis can show.

Public sourcesNong et al., JAMA Network Open (2024); Platt et al., JAMA Network Open (2024); Cavalier et al., JAMA Network Open (2025); Cohen, Georgetown Law Journal (2020).

ScopeSupports heterogeneous public preferences about disclosure and the narrow 2020 doctrinal reading that, in general, liability will not lie for failing to inform a patient about medical AI. Survey responses do not show what a specific patient knew or chose; the legal reading is dated, jurisdiction-sensitive, and not a universal no-notice rule.

RL-020-07: Contracts, disclosure, and the liability asymmetry.

Why it matteredThe remedy depends on recognizing that hospitals may control local use while vendors can control information and contractual allocation.

Public sourcesKoppel and Kreda, JAMA (2009); ONC, EHR Contracts Untangled.

ScopeSupports the documented history of contract terms that can restrict disclosure or shift risk and the public guidance for negotiating health-IT contracts. It does not establish the terms of any unnamed hospital’s current contract or settle liability in a particular case.

RL-020-08: External governance: final rules, proposed rules, and announced initiatives.

Why it matteredThe essay treats outside governance as a partial backstop and therefore must keep the legal and institutional status of each item exact.

Public sourcesJoint Commission and CHAI announcement (June 2026); HTI-1 final rule; HTI-5 proposed rule.

ScopeSupports federal algorithm-transparency requirements for certified health IT, a later proposal that remained proposed at the issue’s as-of date, and an announced professional/accreditation initiative. None by itself verifies local performance after deployment.

What the sources do not do

The sources bind the factual claims; they do not prove the assignment of responsibility. “The hospital should own first-line checking” is a governance judgment built from control, local information, and exposure — not a sentence found in any one citation.