On October 9, 2026, the Journal of Medical Systems published an evaluation of eight speech-to-text systems. One speaker read 100 published surgical case reports, producing 800 transcripts across the systems. Depending on the system, 30–68% of transcripts contained at least one error the reviewers judged clinically significant.
These were changes in meaning that were not obvious at first glance and could mislead a reader and potentially affect patient care. The study did not measure actual patient harm.
Testing took place from October 22 to November 1, 2025. It covered controlled transcription, without evaluating complete consultation-note workflows. The findings describe the versions and settings tested then and cannot establish a ranking of today’s products.
Our practical takeaway for a clinic: assign a clinician to check drafts before they enter the patient record. In a pilot, count checking and correction time alongside any time saved on typing. Start technical tests with fictional cases; agree on patient-data safeguards before using real recordings. Review alone does not establish that a system is safe.
Existing NHS England guidance also calls for checking AI-generated notes before adding them to the record. That is background guidance for England, not a new rule announced by this study.
October 9 is the journal-version date. King’s College London also lists an accepted manuscript; its first public-availability date has not been established.
Read also: where human review is needed.



