Skip to content
    Back to writing
    August 17, 2026 · 6 min read

    Ontario's own test of the AI scribes never made it into the exam room

    TL;DR [show]

    Ontario's Auditor General reported in May 2026 that all twenty AI-scribe vendors approved for provincial use produced inaccurate notes in the procurement test: nine of twenty fabricated additions to treatment plans, twelve of twenty recorded a different drug than the one prescribed, seventeen of twenty missed mental-health details. The piece argues the numbers are not the finding. The finding is that Supply Ontario's scorecard weighted domestic presence in Ontario at 30 percent of 530 points and accuracy of medical notes at 4 percent, with no minimum passing score anywhere in the stage, so a bidder could score zero on accuracy and still be approved. That asymmetry is the same one operating in the exam room where an Ottawa patient discovered her OBGYN's scribe had recorded her without being told: in both cases the party carrying the risk is the party with no information. Revises the November 2025 read that ambient scribes were the healthcare-AI category where evidence and adoption converged, on the grounds that the clean buyer-is-beneficiary alignment praised there is the mechanism that kept accuracy out of the buying criteria. Closes on the February 2024 routing-not-supply claim: the tool collects its consent from the shortage it claims to relieve.

    Ontario's own test of the AI scribes never made it into the exam room — by Thomas Jankowski, aided by AI
    The test results stayed in the file— TJ x AI

    An Ottawa woman worked out, after her appointment was already over, that her OBGYN's AI scribe had recorded the whole thing. Nobody had told her. She was not especially angry about it in the post she wrote, which is the part I keep returning to. She was trying to establish what had happened to her.

    The thread underneath her post did something better than outrage. One person praised an emergency doctor who had asked permission first. The next pointed out that by the time you are in front of that doctor you have probably been waiting for hours, so are you really going to stop the appointment and ask him to explain the software. A third put the mechanism in one line: a lot of people will consent to this AI stuff out of guilt.

    That is one end of this. The other end is a procurement file, and the two of them are the same story told from opposite sides.

    What was she agreeing to?

    Three months before that thread, on May 12, Ontario's Auditor General published a special report on how the provincial government uses AI. A section of it covers AI scribes, the ambient-listening tools that sit in on a clinical conversation and produce the note afterwards. The Ministry of Health initiated and part-funded the program. Supply Ontario, the provincial procurement agency, ran a bid process and stood up a vendor-of-record list so any health-care provider in the province could buy from a pre-qualified supplier. Twenty vendors were approved.

    To score the accuracy criterion, Supply Ontario gave every bidder the same two simulated recordings of a doctor talking to a patient, then had medical professionals from OntarioMD and Ontario Health read the notes that came back. All twenty of the approved vendors produced notes with at least one of three kinds of error.

    Nine of the twenty fabricated material. Not stylistic drift: the systems invented additions to treatment plans, referring the patient to therapy or ordering blood tests that nobody in the recording had mentioned. Evaluators flagged notes stating there were no masses found, and notes recording the presence of anxiety, when neither subject had come up at all. Twelve of the twenty wrote down a different drug than the one the doctor prescribed. Seventeen of the twenty missed key details about the patient's mental health in at least one of the two tests, and six of the twenty missed them across both.

    Those are results from a controlled procurement exercise with scripted recordings, not incidents in clinics. The Ministry says no patient harm has been reported in Ontario. Hold that thought carefully, because it is the whole reason the numbers are interesting: this is what the systems did on the test they knew they were taking.

    What was the province actually buying?

    Here the report stops being a story about software quality.

    Supply Ontario's second-stage scorecard ran to 530 points. "Domestic presence in Ontario" was worth 159 of them, thirty percent of the total. "Accuracy of medical notes generated" was worth 20 points, four percent. Where the vendor is registered outweighed whether the note says the right drug by roughly eight to one.

    And the four percent was not even a floor. The report notes there were no minimum passing scores anywhere in that stage, so a bidder could score zero on accuracy of medical notes, zero on system security controls and zero on bias controls, and still clear the 371 points needed to be approved. Accuracy, on that scorecard, was a preference rather than a requirement.

    The rest of the file reads the way you would now expect. Eleven of the twenty approved vendors submitted no third-party audit reports at all. Supply Ontario relied mainly on vendors attesting to their own compliance rather than obtaining the evidence and reading it. Nobody had to run a live demonstration: vendors were handed the recordings, generated the notes offline on their own time, and submitted them along with a signed statement that they had not edited them. The Auditor General observes, politely, that this created a risk vendors could run the recordings more than once.

    Meanwhile NHS England had issued guidance in April 2025 telling its organisations to use only scribes registered with the UK medicines regulator as Class I medical devices, with clinical accuracy requirements attached. Ontario ran a scorecard with a domestic-content weighting.

    Why did I call this the category that worked?

    In November 2025 I wrote that ambient scribes were the one healthcare-AI category where evidence and adoption moved together, unlike imaging tools that sit in trials nobody implements or clinical decision support that gets deployed with no trial at all. I gave a reason for it that I was pleased with. The clinician is the buyer of the procurement decision and the beneficiary of the deployment, I wrote, and I called that alignment _operationally clean_. I also filed the fact that scribes could be bought through ordinary EHR procurement cycles as a point in the category's favour.

    Both of those were the failure, described approvingly.

    I missed the third person in that alignment.

    When the buyer is the beneficiary, the purchase gets scored on the benefit the buyer can feel. That benefit is the minutes back, the charting that does not follow you home. It is real, and it is why I still think the category works on its own terms. But the risk from a wrong note lands on a third party who is not in the purchasing conversation, has no vote in it, and in the Ontario case was not told the conversation had occurred. Accuracy was never going to be the buying criterion, because nobody buying carries the cost of it being wrong. Four percent was an accurate statement of whose problem the accuracy was.

    The convenient procurement path did the rest. Routing a clinical-safety question through the cycle built for commodity software is how it ends up on a sheet where the vendor's head-office address is worth eight times more than the drug name. I was reading the trial literature, which measured documentation time and burnout, and the trials measured those things well. Nothing in that literature was measuring whether anyone downstream checks the output, and I did not notice that I had no evidence on the question I was implicitly answering.

    Who is supposed to catch the wrong drug?

    The answer the system gives is the doctor. OntarioMD issued guidelines telling physicians to review the generated notes manually before they go in the chart, which is the correct instruction.

    Doctors are not required to attest that they did it. There is no sign-off in the software, no record that a human read the paragraph before it became part of a medical history. The Auditor General has to spend a recommendation asking vendors to build that control, which tells you it does not exist. So the safeguard against a fabricated referral is an unverified habit, practised at the end of a day, by someone who acquired the tool specifically because they did not have enough time.

    That is the part that closes the circle. The review that catches the error costs exactly the minutes the tool exists to return. Any doctor who performs it perfectly gets less value from the product; any doctor who skips it is doing what the purchase was for. The province did not have to decide which of those to encourage, because it never asked.

    So why does anyone say yes?

    In February 2024 I argued that Ontario's primary-care failure is a routing problem, not a supply problem: roughly 2.3 million people unattached, real capacity sitting unused in practices below their roster limits, patients self-routing into walk-in clinics and emergency departments because no matching layer exists to send them anywhere better. I still think that is right, and I think it explains the woman in Ottawa better than anything in the audit does.

    The scribe is sold as a supply fix. It hands minutes back to a doctor the province cannot produce more of. That framing is also what makes its consent model survivable, because the patient asked to agree to it is agreeing under a scarcity she has already paid for in waiting. Declining means spending some of the appointment she waited hours for on an argument about software, with the person she needs something from. The guilt the commenter named is a correct reading of the incentives in the room. Scarcity turns permission into a courtesy offered to the system.

    _The tool collects its consent from the shortage it claims to relieve._

    That is not a fix arriving into a broken system. It is a broken system being spent, quietly, as the price of admission.

    The province ran the test, wrote down what it found, and weighted it at four percent. She was never told there was a test.

    —TJ