Six Features Health Systems Now Demand from Your AI Tools: A Recap of Our Recent AI Webinar
On July 30, AI Catalyst, THMA's AI research team, hosted a webinar for our industry members titled, "4 Hard Truths About AI in Healthcare." Below, AI Catalyst Executive Director Thomas Seay recaps six features health systems now expect from industry partners.
Even a year ago, many health systems chose their AI tools for essentially ad hoc reasons: Which had the shiniest demo? Who promised the highest ROI?
Not anymore. Health systems have professionalized their AI purchasing, and they've established formal bars that industry partners must clear.
To pin down these emerging standards, AI Catalyst — The Health Management Academy's collaborative of 53 leading health systems — has collected more than 100 real-world AI governance artifacts: risk-tiering rubrics, vendor questionnaires, review-gate checklists, and more. We've uncovered six new expectations that AI partners must meet.
1. A "silent mode": Before launch, your model must prove itself on each health system's data.
Health systems have been burned by AI hype before, and they want proof that your tool will deliver — not just in theory, but on their patients. Duke now requires a Silent Evaluation phase for every high-risk AI tool, running the model on live data "without results being shown in clinical practice." Cedars-Sinai, UF Health, and other systems have established similar requirements, although some accept validation on historical data.
Your action step:
Make silent or retrospective demonstrations of your tool's capability a standard offering. Not only will this help you clear governance review, but it's a powerful sales tactic: It's your chance to prove your results.
2. Reversibility: Health systems want proof they can shut your tool down if needed.
One reason why health systems want a "kill switch" is the fear that AI might act unpredictably or dangerously. But just as importantly, leaders worry their clinicians will grow dependent on a particular tool or their system's data will get locked in, forcing IT to maintain and support each tool perpetually. To avert this risk, systems now demand proof of wind-down capability. Mass General Brigham, for instance, vests a committee of technical experts with the authority to halt a deployment, and Duke requires sunset criteria before an AI tool's general deployment.
Your action step:
You might worry that advertising reversibility could invite your customers to walk, but our data says otherwise: Of 633 AI use cases we catalogued in a recent survey, only two had been fully sunset. Instead, consider the upsides of reversibility. If health systems know they can reverse course — perhaps because you've provided an emergency "kill switch" and a rollback playbook — they're more likely to give you the chance to deploy in the first place.
3. Risk scores: Human reviewers need to know which AI outputs deserve the greatest scrutiny.
Most healthcare AI implementations require a "human in the loop" to review every AI output, but health systems are increasingly experimenting with end-to-end automated AI solutions — and even when humans still review AI outputs, they're often overwhelmed and at risk of distraction. To help allocate scarce human attention, our members have told us they want risk scoring on every AI output. They need help differentiating where your model is 99.9% confident (and could bypass some human review) from where it's only 70% sure.
Your action step:
Include confidence scores whenever possible — and take care also to quantify the consequences if your output is wrong. A $5 error is much less consequential than a $5,000 error or, worse still, patient harm. Consequence scoring allows health systems to apply tighter or looser risk tolerances depending on how bad a mistake would be.
4. A ready-to-use model card: Health systems want to steal your documentation for their own governance library.
Every health system governance committee needs the same core facts, typically captured in a "model card." Too often, though, each health system must build its own card from scratch. If you do that work on their behalf, you might speed up implementation: UF Health's policy, for example, may accept vendor-supplied model cards that cover required elements.
Your action step:
A draft model card is increasingly a baseline expectation — if you fail to provide needed information, you risk looking behind the times. Be ready to address even sensitive topics, such as populations not represented in training, subgroup performance, contraindicated uses, whether customer data trains the model, and FDA status. Be sure to include version numbers; Duke requires re-documentation of any change that affects the regulatory pathway.
5. Straight answers: Health systems want answers about AI risks in writing, quickly — and they prefer a partial answer to none.
UVM Health Network uses a 45-question vendor evaluation form that other systems have since borrowed or adapted. Some of the questions are rather pointed (more on those below), but if you're tempted to skip them, be aware that your customers view missing answers as red flags. Cedars-Sinai, for instance, scores vendor responses 1 to 5, and the bottom of that scale reads "unacceptable or entirely missing."
Your action step:
Prepare templatized answers for the questions health systems ask most. Pay particular attention to four challenging questions that, in our experience, stall deals: Is the model yours, or built on a commercial or open-source foundation model? Does it train on health systems' data? Can health system users reach the audit trail? And do you run an adverse-event reporting channel for providers?
6. Credible ROI measurement: Health systems need your help measuring the numbers they can't measure on their own.
In a recent AI Catalyst survey of 52 health systems, the owners of 51% of all AI deployments reported an ROI of "unknown." Worse, health systems are less likely to measure ROI when they're working with an external vendor than when they're building AI in-house. But an unquantified AI investment is one that's at risk of cancellation if and when budgets tighten.
Your action step:
Create a clear playbook for measuring your tool's ROI. What KPIs must a health system measure? What should they measure before vs. after AI implementation? What's the math that turns that into an ROI? Texas Health Resources, for instance, measured ROI on its ambient documentation deployment (a notoriously hard-to-quantify use case) by benchmarking each clinician's wRVU productivity percentile before and after go-live. In THR's case, the work was labor-intensive, but there's no reason a vendor tool couldn't automate much of it.
Let's be candid: While some of these six features are relatively easy to provide, others require more hassle — and even vulnerability about your product's limitations. But if you meet these emerging expectations, your reward could be a speedy path through health systems' governance processes.
And if you fail to keep up? Even the most mind-blowing AI technology risks stalling in governance review.
Want to learn more about how you can get involved with AI Catalyst? Reach out to catalystmembership@hmacademy.com.
