AI Scribe Errors in Healthcare: What Actually Goes Wrong in 2026
AI scribe errors in healthcare are real, measurable, and mostly boring: these tools leave things out far more often than they invent things. Research summarised by Forbes in March 2026 found omissions account for 71–83% of documentation errors in AI-generated notes, while reported hallucination rates sit near 1–3%. This guide is for clinicians, practice managers, and compliance leads who already use an ambient scribe or are about to buy one. Skip it if you want a vendor leaderboard — this covers failure modes, costs, and controls. Every price below comes from an official vendor page.
Key Takeaways
- Omissions, not hallucinations, are the dominant AI scribe error type — 71–83% of documented mistakes.
- Reported hallucination rates for large language model scribes are roughly 1–3%, versus 7–11% error rates for older dictation software.
- Independent testing found error rates ranging from about 12% to 25% across four ambient scribe platforms — vendor choice matters.
- A JAMA study of 1,800 clinicians (1 April 2026) found only 16 minutes of documentation time saved per eight hours of care.
- The clinician who signs the note owns the error, not the vendor. Every note needs a real read-through before signing.
What Are AI Scribe Errors in Healthcare?
An ambient AI scribe listens to a patient visit through a phone or laptop microphone, transcribes the conversation, and writes a structured clinical note. An AI scribe error is any gap between what happened in the room and what ends up in the chart.
A September 2025 commentary in npj Digital Medicine grouped these into four failure modes. The table below adds how often each appears and who is realistically able to catch it.
| Error type | What it means | How common | Who catches it |
|---|---|---|---|
| Omission | Something discussed is missing from the note | 71–83% of all errors; 45% of omissions carried moderate clinical importance | Only the clinician who was in the room |
| Hallucination | Fabricated content, such as an exam that never happened | Roughly 1–3% of notes | Careful line-by-line review |
| Misinterpretation | Context flipped — a dropped “not,” or a discussed-but-declined treatment recorded as prescribed | Not separately quantified in published studies | Clinician review; sometimes pharmacy |
| Speaker attribution | Patient statements recorded as clinician findings, or vice versa | Not separately quantified | Clinician review |
Real examples reported by the Australian medical defence organisation Avant include a note recording olmesartan when the patient takes irbesartan, a missing “not” that reversed a diagnosis, and a fabricated neurological examination.
One more pattern gets little coverage: performance is not equal across patients. The npj Digital Medicine authors flag reduced transcription accuracy for Black patients compared with White patients, and for speakers with non-standard accents. That makes error rates a fairness issue, not just a quality one.
The Real Cost of AI Scribe Errors in Healthcare
Vendor pricing is the small number. Freed publishes its plans openly: Starter at $39 per month for up to 40 notes, Core at $79 per month with unlimited note generation, and Premier at $104 per month billed annually or $119 monthly. Freed lists a 7-day free trial with no credit card, free access for residents, students, and trainees, and custom pricing for groups. Heidi publishes a free tier with unlimited transcription and a 14-day trial on its Clinician plan, but does not publish a dollar figure on its pricing page — treat any specific Heidi price you see in a roundup as unverified.
The costs most articles skip:
- Review time. Correcting a note takes minutes that come straight out of the time saved. A JAMA study of 1,800 clinicians across five academic centres found net savings of 16 documentation minutes and 13 EHR minutes per eight-hour shift — thin margins that vanish if error rates are high.
- Year-one return. Only 8% of adopters reached positive return on investment in their first year, according to analysis of that same five-centre dataset published in August 2026. Most organisations expect 24–30 months.
- Governance staffing. Someone has to sample notes for accuracy. Only 10% of organisations have a formal AI oversight board.
- Medico-legal exposure. A signed note is the clinician’s note. There is no published mechanism for shifting liability for AI scribe errors in healthcare back to a vendor.
- Integration and rework. EHR push is often a higher-tier feature, as it is on Freed’s Premier plan.
Where AI Scribes Break: Three Clusters
1. Hearing the room
Background noise, overlapping speech, family members, telehealth audio, and accented English all degrade transcription before any summarising happens. If the transcript is wrong, everything downstream inherits the mistake. Most AI scribe errors in healthcare start here, before the model writes a word.
2. Deciding what matters
This is the omission engine. The model compresses a 15-minute conversation into a paragraph and must judge relevance. It routinely drops the thing that mattered — a symptom mentioned in passing, a medication the patient stopped taking. In one hospital discharge analysis, AI-generated summaries averaged 1.75 omissions per note against 0.86 for physician-written ones. Negation handling belongs here too: “no chest pain” and “chest pain” are one word apart.
3. Fitting the workflow
Templates pull notes toward a shape. If a specialty template expects a full review of systems, some tools will produce a complete-looking one. Add copy-forward habits and a rushed sign-off queue, and one error propagates through every future note.
Alternatives
| Option | Typical cost | Error profile | Verdict |
|---|---|---|---|
| Ambient AI scribe | $39–$119/month per clinician (Freed, published) | ~1–3% hallucination; heavy omission | Best time-per-dollar, but only with disciplined review |
| Human scribe | Salaried or contracted; no public standard rate | Fewer omissions; notes rated accurate far more often than self-documentation | Highest quality, hardest to afford at small scale |
| Speech recognition dictation | Often bundled with the EHR | 7–11% error rates reported | Errors are obvious rather than plausible, which makes them easier to spot |
| Self-documentation with templates | No licence cost | Roughly 50% of verbally discussed problems go undocumented | Cheapest and slowest; the baseline everything else is measured against |
Who Should Use It
Best for
- Clinicians with high visit volume and consistent visit structure, where the time saved is real.
- Practices that can commit to reviewing every note before signing.
- Organisations with someone accountable for sampling note accuracy monthly.
Skip it if
- You see complex, multi-problem patients where omission risk is highest.
- Your patient population includes many accented or non-native English speakers and you cannot audit for the accuracy gap.
- You want a tool that reduces work without adding a review step. That tool does not exist yet.
Main Competitors
Compare on the job you are hiring the tool to do, not feature counts.
Freed is built for the solo or small-practice clinician who wants to stop charting at night. Published pricing, a 7-day trial, and free access for trainees make it easy to test with your own patients. Its ICD-10 and CPT coding and EHR push sit on the top plan.
Heidi is built for clinicians who want to trial ambient documentation at zero cost first. Its free tier includes unlimited transcription, which is unusual, but advanced templates, coding, and sharing are gated. It suits evaluation and low-volume use more than it suits a full switch.
Neither publishes an independent error rate for AI scribe errors in healthcare. Ask both for their internal accuracy testing before signing anything.
Recent Changes
- August 2026: Analysis of the five-centre dataset (Mass General Brigham, Emory, UCSF, UC Davis, Yale New Haven) put first-year positive ROI at 8% and highlighted that only 10% of organisations run a formal AI oversight board.
- June 2026: Coverage questioning the ambient scribe productivity narrative widened after the April JAMA results, shifting the conversation from time saved to note quality.
- April 2026: The JAMA study of 1,800 clinicians published its modest time-savings figures — the largest such dataset to date.
- January 2026: The FDA reissued its Clinical Decision Support guidance (6 January, reissued 29 January), keeping software outside device regulation only where clinicians can independently verify its logic. As scribes add order pre-population and care-gap detection, that boundary gets thinner.
- Also January 2026: ECRI’s Top 10 Health Technology Hazards for 2026 ranked misuse of AI chatbots first — and did not list ambient scribes at all. That is a fair signal that scribes are not the top patient-safety risk in health IT right now.
On expired offers: older roundups still quote longer free trials than vendors currently publish. Freed’s official page lists 7 days; Heidi’s lists 14. Check the vendor page, not the roundup.
How to Cancel or Downgrade
Most ambient scribes are month-to-month self-serve subscriptions, so the steps are similar:
- Sign in and open account or billing settings in the web app, not the mobile app.
- Choose Change plan to downgrade, or Cancel subscription to end billing. Downgrading from Freed’s Core to Starter reinstates the 40-notes-per-month cap.
- Export or push existing notes to your EHR first — feature gating can remove EHR push on lower tiers.
- Confirm the cancellation date. Annual plans generally run to the end of the paid term.
- Cancelling does not end your duty to retain the clinical record.
Enterprise and group contracts follow your signed agreement, not the self-serve flow.
FAQ
How accurate are AI medical scribes? Reported hallucination rates are roughly 1–3%, but overall error rates in independent testing ranged from about 12% to 25% across four platforms. Accuracy varies more by vendor and patient population than most marketing suggests.
Who is liable for AI scribe errors? The clinician who signs the note carries responsibility for its content. No published framework shifts liability to the scribe vendor. Your organisation’s policy and your medical defence cover should address AI-assisted documentation explicitly.
Are AI scribes regulated by the FDA? Documentation scribes are generally treated as administrative software, not medical devices. The FDA’s January 2026 Clinical Decision Support guidance keeps that exemption only where a clinician can independently verify the software’s reasoning.
Do AI scribes actually save time? A JAMA study of 1,800 clinicians found 16 fewer documentation minutes and 13 fewer EHR minutes per eight-hour shift. That equals roughly one extra patient every two weeks — real, but smaller than vendor claims.
What is the most common AI scribe mistake? Omission. Between 71% and 83% of documented errors are things that were said but never made it into the note, and 45% of those omissions carried moderate clinical importance.
Verdict
Ambient scribes are worth using, and the errors are manageable — if you read every note before signing. The evidence points to omissions as the real risk, not fabricated content, and to wide quality gaps between vendors. Time savings are genuine but modest. Next step: before you renew or expand, pull 20 signed notes from the last month and compare them against your own memory of those visits. That number, not a vendor benchmark, is your actual error rate.
Related reading on AI Era: how AI voice cloning scams work, which AI tools are worth paying for in 2026, and what OpenAI and Anthropic’s cheaper models really cost.
Sources: npj Digital Medicine on AI scribe risks, Freed official pricing, Heidi official pricing, ECRI Top 10 Health Technology Hazards for 2026.
About AI Era
AI Era reviews AI tools using current pricing from official vendor pages, official documentation, practical testing where available, and direct comparison with alternatives. When a vendor does not publish a number, we say so rather than estimate it.