Houston Methodist’s chief innovation officer called it the first independent, peer-reviewed proof that ambient AI increases surgical case volume. The permutation test inside her own paper returned an empirical p-value of 0.182.
That number never made the press release.
The Problem
Health systems spent 2024 and 2025 buying clinical AI off vendor case studies: a single pilot site, no control group, no statistician who didn’t also work for the vendor. Buyers got burned often enough that “peer-reviewed” became the new gate. Show me the published study, or don’t get the meeting.
Vendors heard that. Houston Methodist and Apella, the OR computer-vision company whose cameras now run in more than 200 of the system’s operating rooms, answered on August 5 with two studies: one in the Journal of Imaging, one in Perioperative Care and Operating Room Management [Journal of Imaging, 2026]. Peer review happened. Editors reviewed the methods. Reviewers checked the math. Both papers cleared the bar buyers said they wanted.
That bar measures the wrong thing. Peer review certifies that a study’s methods were disclosed and scrutinized before publication. It says nothing about whether the topline result clears statistical significance, and nothing about whether the paper that passed review is the paper anyone signing a purchase order will actually read past the abstract.
The system generating the underlying data is unglamorous by design. Four ceiling-mounted cameras per room feed a YOLO-based object detector paired with a transformer event model, trained on 137,517 surgeries across 315 operating rooms, that timestamps patient entry, draping, turnover, and wheels-out with F1 scores above 0.98 BMJ Health & Care Informatics, 2025. That is genuinely good computer vision, better than the manually entered EHR timestamps most ORs still run on. The question was never whether the cameras work. It is what a health system should conclude from the specific volume number a hospital that partly owns the vendor chose to publish first.
The Insight
The case-volume study is, on its own terms, unusually rigorous for hospital AI research, and more careful than the scorecards most pharmacy AI ROI claims run on. Houston Methodist tracked 5,417 surgeries over 16 months in a 15-room cardiothoracic suite and built a synthetic control from 11 sister sites, 116,098 comparison cases that had not deployed Apella, using difference-in-differences estimation borrowed from labor economics rather than the simple before-and-after math every vendor deck runs [Journal of Imaging, 2026]. Roberta Schwartz, the paper’s senior author and Houston Methodist’s chief innovation officer, said the goal was a methodology “designed to demonstrate causation, not just correlation” Schwartz, PR Newswire, 2026.
Then the paper does the part vendor decks skip: it discloses what the method actually found. The topline is a 7% increase in case volume, roughly 25 additional cases a month, about 300 a year, with a 95% confidence interval spanning 8.3 to 41.0 [Journal of Imaging, 2026]. That interval clears zero, which is the number that traveled. The permutation test built into the same design, the check against the result being chance, returned an empirical p-value of 0.182, well short of the 0.05 threshold “significant” normally means [Journal of Imaging, 2026]. Secondary outcomes fared worse: unplanned overtime showed no meaningful change, and total operative minutes did not move, meaning the gain came from fitting more cases into the same OR time rather than running longer days [Journal of Imaging, 2026]. The authors’ own language settles it: the deployment was “associated with, rather than caused by” the increase [Journal of Imaging, 2026].
“The permutation test built into the study’s own design returned a p-value of 0.182: real rigor, and a result that falls short of significant.”
Two disclosures sit in the same section. The intervention was never the camera system alone. It shipped bundled with turnover-time reviews, housekeeping and anesthesia workflow changes, and surgeon notification alerts, and the authors state plainly that the design “cannot separately attribute the observed gains to one or the other component” [Journal of Imaging, 2026]. And Houston Methodist Hospital holds a minority equity stake in Apella, acquired when it joined the company’s Series B round in January 2026 [Journal of Imaging, 2026]. An Apella employee ran the original analysis from aggregated data; a Houston Methodist analyst independently replicated it, a reasonable way to manage a real conflict. Managing it does not make the confidence interval any wider.
None of that is misconduct. It is closer to the opposite: a study that shows its work, published in a journal built for imaging methodology rather than surgical outcomes, reviewed by people qualified to check the computer vision and less positioned to weigh in on what a health system should consider clinically meaningful, a gap in kind with the ongoing-validation blind spot in how hospitals buy clinical AI generally. That is a fair venue for a technical validation. It is a strange venue for the sentence “AI increases surgical case volume” to enter a sales deck as settled fact.
In Practice
| The August 5 announcement led with | The study’s own disclosures say |
|---|---|
| First independent evidence AI increases case volume | Houston Methodist holds equity in Apella; a company employee ran the first analysis |
| A methodology built to show causation | Authors’ language: “associated with, rather than caused by” |
| Ambient AI drove the 7% gain | Gain traces to a bundled intervention the design cannot decompose |
The Bottom Line
Apella’s own announcement, naming the source plainly since the finding sits behind a paywall this newsroom could not independently open, puts the second study’s number at a 40% improvement in scheduling accuracy on flagged cases, mean error falling from 57.4 to 34.2 minutes, and a 46-percentage-point drop in late-ending OR days across six Houston Methodist hospitals, with no loss in case volume; rescheduled cases landed closer to their actual duration 72% of the time [Perioperative Care and Operating Room Management, 2026]. That is a cleaner, more defensible result than the volume claim, and it got a fraction of the attention, because “40% more accurate scheduling” does not sell like “AI makes hospitals more money.”
Every OR computer-vision vendor watching this rollout just learned the actual playbook: fund or co-author a peer-reviewed study, disclose the conflicts honestly in the limitations section, and let the abstract carry the number that closes deals. None of that requires bad faith from Houston Methodist, whose disclosures here are more thorough than most published clinical AI research manages, or from Apella, which built its case on real academic collaboration instead of a marketing white paper. It requires buyers who stop treating “peer-reviewed” as a finish line and start reading past the abstract the way Houston Methodist’s own statisticians did. The next OR pitch that opens with a published paper will move through procurement faster than one that doesn’t, regardless of what the limitations section says, unless somebody in the room gets that far first.