An AI system renewing prescriptions without a doctor sounds like the risky experiment; an AI assistant answering pharmacy phone calls sounds like the safe one.
Early evidence from two deployments suggests the opposite may be true, or, at minimum, that health systems are using the wrong shorthand to judge AI risk in pharmacy care.
In Utah, a state-authorized pilot is testing whether AI can evaluate and recommend renewals of existing prescriptions for chronic conditions. The system operates within a defined formulary and escalates patients who need additional review. Five months of early data suggest it has been cautious: Doctronic’s AI recommended renewal in 72% of cases, and reviewing physicians agreed with that recommendation 91% of the time. In the remaining 28%, the AI escalated the case instead of recommending a renewal; physicians agreed that more information was needed 69% of the time. The lower agreement rate on escalations raises a different possibility: the AI may be erring on the side of escalation more often than reviewing physicians believe necessary. That could represent a deliberately cautious failure mode, though the early agreement data alone cannot show whether those escalations were appropriately conservative or unnecessary.
The pilot is operating through Utah’s AI regulatory sandbox, which gives the state a controlled way to test emergency AI applications under defined, temporary regulatory relief. State officials have framed the program as an experiment in improving medication access and affordability while reducing routine renewal burden on clinicians. The approach has also faced substantial clinical opposition. The Utah Medical Licensing Board called for the pilot’s suspension, and the Utah Medical Association has publicly opposed the model on patient-safety grounds.
The results are preliminary. The pilot remains in its first phase, meaning a licensed practitioner still reviews every renewal request. Utah regulators say there have been no reported serious safety incidents, but they also acknowledge that the available evidence is too limited to draw firm conclusions about performance or benefit.
Still, this is not the obvious disaster that “AI prescribing medication” might bring to mind.
Meanwhile, customers of Kinney Drugs, a regional pharmacy chain in Vermont and New York, have reported problems with a considerably less ambitious use case: an AI assistant introduced to handle patient communications and refill requests.
According to reporting from VTDigger, customers said the system contacted them about medications they did not need, struggled to identify existing accounts, communicated incorrect dosage information and failed to alert them when prescriptions were ready. One customer reported accumulating four bottles of a medication taken only twice a week after approving automated refill requests she believed were legitimate. Others described delays that required more—not less—intervention from already strained pharmacy staff.
These accounts do not establish that every reported pharmacy delay or error was caused by the AI. Kinney Drugs told VTDigger that it was aware some patients experienced frustration during the rollout and was reviewing their concerns. But the implementation exposes a larger problem: healthcare organizations often assess AI risk based on the apparent complexity of the task rather than the position the tool occupies between the patient and care.
The “simple” workflow was not simple
Automating a pharmacy phone line looks like a low-risk administrative application. The system is not diagnosing a condition, selecting a medication or formally authorizing treatment.
But it is communicating about sensitive medications, interpreting patient intent and initiating actions that affect whether, when and in what quantity a patient receives a drug. It may also become the patient’s primary route to a pharmacist.
Kinney Drugs’ previous system allowed callers to enter information through a telephone keypad and navigate toward a staff member. Customers told VTDigger that the new assistant required them to speak with the AI, leaving some to visit a store in person to avoid it. That alternative is inconvenient for many patients and unavailable to some rural residents with limited transportation.
The deployment therefore changed more than call handling: It changed access.
That distinction matters for health systems evaluating pharmacy automation, contact-center AI and patient-facing agents. A task can appear administrative on an organizational chart while still functioning as a clinical access point for the patient.
The same dynamic can show up well beyond pharmacy:
A call-center AI that appears to be handling scheduling may influence whether a patient with worsening symptoms reaches a nurse or waits for a routine appointment.
A patient-facing agent that helps with referrals or prior authorization may look administrative, but if it misroutes a request or makes escalation difficult, it can delay access to a specialist, diagnostic test or treatment.
In each case, the risk comes less from whether the AI is making the final clinical decision than from whether it controls the path to someone who can. This is particularly relevant for patients with transportation challenges, limited digital access, or few pharmacy alternatives.
The relevant question is not simply, “Does this tool make a clinical decision?” It is also, “What care becomes harder to reach when this tool fails?”
Guardrails matter more than the autonomy label
Utah’s pilot is more clinically consequential, but it also has unusually visible boundaries.
The AI can renew an existing prescription; it cannot initiate a medication or change its dose or frequency. Controlled substances are excluded. The system must escalate cases when information conflicts, clinical complexity crosses defined thresholds or a patient reports a potential complication. Pharmacists can request physician review, and a licensed physician remains attached to each prescription authorization.
Utah has also revised the program while it is underway. Regulators changed the threshold for moving medication groups into the next phase and removed two drugs from Doctronic’s formulary.
None of this proves the system is safe enough to operate autonomously. The published results come from a small, early sample, and the state has initiated a separate review of interactions rather than relying entirely on reports from physicians employed by Doctronic. The 91% agreement rate also should not be treated as a clinical outcome: it measures whether a reviewing physician considered the renewal appropriate based on the information collected, not whether patients avoided adverse events or achieved better disease control.
But the pilot shows what structured experimentation looks like. The use case is narrow. Escalation rules are explicit. Human review is built into the initial phase. The regulator can alter the formulary and prevent progression when evidence is insufficient.
By contrast, a patient communication tool can be deployed broadly while its clinical importance remains obscured by the word “administrative.”
Efficiency is not the outcome
Both implementations began with a familiar promise: remove repetitive work and give clinicians or pharmacists more time for higher-value care.
In Utah, the underlying burden is real. Prescription renewals consume clinician time, often without reimbursement, while lapses can interrupt treatment for patients managing chronic conditions.
Kinney Drugs similarly said its assistant was intended to reduce routine phone calls reaching pharmacy staff. Yet customers interviewed by VTDigger described errors and confusion that sent them back to pharmacists for help.
That is the operational trap. A tool can successfully deflect calls while making the overall process worse. It can shorten an employee’s task queue while increasing patient effort, repeat contacts, abandoned requests or downstream corrections.
Health systems should therefore be skeptical of AI business cases built primarily around transactions avoided: calls contained, messages answered or renewals processed. Those measures show that the tool did something. They do not show that the work disappeared.
A better test for evaluating the AI-assisted medication journey:
Did patients receive the correct medication on time, with fewer medication lapses?
Did pharmacists spend less time resolving preventable errors or correcting AI-generated work?
Did patient effort decrease, including repeat calls, abandoned requests or avoidable in-person visits?
Could patients quickly reach a person when the system was wrong, uncertain or unable to complete the request?
Questions to consider
Where has an administrative AI tool quietly become a clinical access point?
Are you measuring work removed—or work displaced?
Can patients reach a human before the error becomes consequential?
What evidence would justify greater autonomy?