Cheap to Run, Expensive to Trust

Cheap to Run, Expensive to Trust

Share

1. The Case That Started It

A 53-year-old man presented with a tumour in the left retromolar trigone, where the cheek meets the last molar. The treatment involved removing part of the jawbone (segmental mandibulectomy), part of the upper gum ridge (partial upper alveolectomy), and the neck lymph nodes on that side (left MRND), followed by reconstruction using a chest muscle and skin flap (PMMC).

His adjuvant plan was already decided — radiation, based on the real pathology, was already underway. That was when I ran his outsourced report through a pipeline I am still piloting: a pressure test to see what I am walking into before I take it to an ethics committee.

The report was unambiguous: deep resection margin, 1.5 cm, uninvolved tissue. The model read it, flagged its own uncertainty, and still returned "ambiguous ... involved" — the opposite of what was written. Nothing changed for the patient; a human had already read their document. However, it provided me with the information I required.


2. Why This, Why Now

This failure mode matters because "agentic AI" is being promoted to hospital committees as the obvious next step beyond a plain language model: give it tools, let it browse, let it verify itself, and accuracy should improve. This paper tests that claim against hard benchmarks, not a vendor demo.

This is important to me because my problem is narrow: reading a structured pathology report against a known checklist and stating plainly what is missing, contradictory, or not assessable. That is not an open-ended diagnosis; I would call it a deterministic extraction task, and it deserves to be judged as one.


3. The Paper at a Glance

  • Citation: Liu Y, Carrero ZI, Jiang X, et al. Benchmarking large language model-based agent systems for clinical decision tasks. npj Digital Medicine. 2026;9:259. DOI: 10.1038/s41746-026-02443-6. Open Access.
  • Study design: Benchmarking of two agentic systems and baseline large language models across three benchmark families.
  • Benchmarks: AgentClinic, MedAgentsBench, and Humanity's Last Exam biology/medicine subset.
  • Primary outcome: Accuracy against the ground truth. Secondary outcomes included token use, latency, workflow complexity, hallucination frequency, and diagnostic impact.
  • Key result: Agentic systems modestly improved accuracy, but at the cost of far more tokens and longer latency.

4. What the Paper Really Shows

The main message is not that agentic AI is useless; rather, it is that a generalist agent, made to reason, call tools, and verify itself in one loop, can be expensive for only modest gains. This distinction is important. As I read it, this architecture is not the same as a lighter orchestrator-worker design that assigns narrow subtasks to narrower workers; a different study tested that design on different tasks, so this comparison is my own reading, not a finding either paper makes directly.

Another important point is hallucination. This study reports that even after strong filtering, a meaningful fraction of scenarios remained affected by hallucinations. This is the kind of "looks careful, still wrong" failure mode that should worry clinicians more than a simple wrong answer.

One result especially stays with me: MedGemma, despite being tuned for clinical use, performed poorly on this particular benchmark such that it failed to produce usable answers. It is a narrow result — one model, one benchmark — but a useful reminder that specialist branding does not automatically beat a stronger general backbone.

 


5. Why This Is Not the Wrong Lesson

This study should not be read as "agentic AI is too expensive for medicine." It should be read as "this particular architecture is expensive for this particular task." This is a very different claim.

For a structured pathology report, I do not need a model that thinks like a consultant. I need to extract, compare, and abstain when the evidence is incomplete. Therefore, the right validation question is not just accuracy; it is how often the system safely says "not assessable" instead of guessing.

The pathology literature supports this. CAP's oral cavity protocol and the ICCR oral cavity dataset require structured reporting of elements such as margin status and closest margin distance because, as Woolgar's review of histopathological prognosticators established, these elements carry real prognostic weight.


6. Does This Work Here?

Yes, but only in a bounded role.

Access: OpenManus is an open-source software, so the barrier is not vendor access. The real issue is whether we should deploy the wrong workflow simply because it is available.

Infrastructure: The API cost is low enough for batch retrospective screening. The hard cost is not money; it is time, governance, and the review bandwidth.

Latency: A minute per case may be acceptable for a retrospective audit but not for anything that needs to sit inside a live clinical workflow.

Data: None of the three benchmarks included Indian patients. One partial exception is that a portion of MedAgentsBench's hardest questions come from MedMCQA, which is built from the AIIMS and NEET-PG entrance exams. That is Indian exam content, not clinical data.

Regulation: India's CDSCO framework for medical device software has recently been finalised, moving to a risk-based classification approach. Under this framework, a retrospective, human-adjudicated screening pipeline should be a considerably safer first step than autonomous clinical use.

Workforce: Reading an agent's tool-call trace is a different skill from reading its answer. This skill is not yet part of anyone's training.

Equity: The real gap is not a token spend. It is who has the informatics time, governance support, and ethics oversight to validate the tool before trusting it.

🟡 Adoptable With Conditions — retrospective, human-supervised, abstention-enforced screening only.


7. The Real Experiment

If I wanted to validate this in Patan, Gujarat, I already have the outline.

A hundred consecutive outsourced oral cavity resection reports will be mapped against the CAP current oral cavity protocol. Two blinded expert reviewers, with third-reviewer adjudication. Each report run twice through a locked model version to test repeatability.

The primary outcome should not be the raw accuracy. It should be sensitive to missing or ambiguous treatment-critical elements because a confidently wrong answer is more dangerous than a correctly abstained one. The model should be required to say "not assessable" rather than make a guess.

This is the safety test that this class of system needs before it is allowed anywhere near a margin status unsupervised.


8. The Verdict

Running this agent will not bankrupt anyone. At today's prices, it costs cents a case. What it costs is trust — a system that flags its own uncertainty and still resolves it the wrong way has no business near a margin status unsupervised, no matter how cheap the tokens are.


9. Next Issue + Reader Question

Reader question: How many of your own outsourced reports would you trust an "ambiguous" flag on before pulling the primary document yourself? Next issue: a systematic review of 23 studies pitting handcrafted radiomics against deep learning for head and neck cancer prognosis — and a finding that should sound familiar. Across the entire field, the best-looking numbers come from the least rigorously reported studies.


10. Sources Cited in This Issue

  1. Liu Y, Carrero ZI, Jiang X, et al. Benchmarking large language model-based agent systems for clinical decision tasks. npj Digit Med. 2026;9:259. doi:10.1038/s41746-026-02443-6
  2. Klang E, Omar M, Raut G, Agbareia R, Timsina P, Freeman R, et al. Orchestrated multi agents sustain accuracy under clinical-scale workloads compared to a single agent. npj Health Syst. 2026;3:23. doi:10.1038/s44401-026-00077-0
  3. College of American Pathologists. Protocol for the Examination of Specimens from Patients with Cancers of the Lip and Oral Cavity. Version 4.3.0.0. April 2026.
  4. Müller S, Day TA, Griffith CC, et al. Carcinomas of the Oral Cavity Histopathology Reporting Guide. 2nd ed. International Collaboration on Cancer Reporting; 2024.
  5. Central Drugs Standard Control Organisation. Guidance Document on Medical Device Software. Doc No. CDSCO/MD/GD/MDSW/01/2026. 21 July 2026.
  6. Woolgar JA. Histopathological prognosticators in oral and oropharyngeal squamous cell carcinoma. Oral Oncol. 2006;42(3):229–239. doi:10.1016/j.oraloncology.2005.05.008

Read More Articles
Comments (0)
Your comments must be minimum 30 character.