Google tests AMIE for clinical video consultations

Google’s research medical AI system, AMIE (Video), conducted synchronous video consultations with professional patient actors and received clinical evaluator ratings on par with primary care physicians across several core measures.
Fifteen trained actors portrayed conditions across cardiopulmonary, abdominal, HEENT, neurological or psychiatric, and musculoskeletal presentations. Google says studies involving real patients and their own health conditions must follow before anyone can draw conclusions about clinical use.
AMIE divides a video consultation among three agents
AMIE uses an asynchronous multi-agent architecture rather than assigning dialogue, clinical reasoning, and perception to one model process. Google says a single agent cannot currently sustain natural conversational response times while also conducting detailed reasoning and continuously processing audio-visual input.
The talker agent handles the spoken interaction with the patient. It aims to maintain conversational flow, drawing on information from the other agents. The planner agent runs in the background, updating differential diagnoses and management plans as the consultation progresses. It also identifies missing information and reprioritises clinical goals.
A perception agent reviews video and audio streams continuously. It looks for non-verbal signs, physical findings, and auditory signals, then places those observations into the conversation’s clinical context.
Latency remains central. Deep clinical reasoning takes time, and long pauses can affect rapport during a consultation. Google’s architecture separates the patient-facing dialogue from the slower work of reasoning and perception, allowing the talker agent to respond without waiting for every background process to finish.
Google reports that automated evaluations found each agent improved clinical measures, including history-taking, clinical reasoning, and treatment recommendations. The evaluations also covered patient-centred communication and response latency.
The study compared video AMIE with text and physicians
Google structured the human evaluation as a multi-arm randomised study. AMIE completed real-time video consultations. A text-only AMIE version provided a modality baseline. Ten board-certified primary care physicians used the same video interface.
An independent panel of 20 experienced primary care physicians reviewed every consultation using established clinical rubrics. The panel assessed general clinical competence, then applied scenario-specific criteria tailored to the case.
The study covered five body systems. Each scenario followed a standardised consultation format with a trained patient actor.
Google says evaluators rated AMIE on par with the PCP group for history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality. AMIE (Video) also matched or exceeded AMIE (Text) across those measures.
Evaluators rated the video system higher on eliciting physical signs and proactively guiding actors through virtual examination manoeuvres than either the PCP group or text-only AMIE. Case-specific perception and examination scores reflected the same reported pattern.
Patient actors also preferred the synchronous video interface to text chat. Google says they rated video as easier to use and more effective for communicating health concerns. The actors rated AMIE favourably for empathy, rapport, and confidence in care when compared with both study alternatives.
Automated testing came before physician review
Google built an automated evaluation suite to develop the video system before the human study. Its framework draws on a taxonomy of telehealth competencies from medical literature, covering visual cues, auditory signals, and physical examination manoeuvres.
Single-turn assessments tested specific perception and reasoning tasks. Google gives anatomical laterality and signs of respiratory distress as examples. Multi-turn simulated audio consultations assessed the system’s conversational performance over a longer interaction.
Visual input entered some of these multi-turn simulations as text descriptions. In a Parkinson’s scenario, for example, an AI patient simulator could describe a patient holding paper to the camera showing cramped, tiny handwriting. This arrangement helped Google test dialogue behaviour alongside visual information, though it does not replicate an end-to-end live video feed.
The automated suite allowed rapid changes to the system design and exposed capability gaps before the actor-based study. The subsequent OSCE evaluation used a synchronous video consultation interface, although the patient presentations still followed prepared scenarios.
The split between these methods should shape procurement and governance discussions. Automated assessments can test defined perceptual tasks at scale. Simulated video consultations can assess interaction quality under controlled conditions. Neither method establishes performance with patients whose symptoms, behaviour, connectivity, environment, and medical history fall outside a prepared case.
Production evidence remains limited to text-based work
Google identifies several limits in the AMIE research. Professional actors cannot fully reproduce the variability of real patient encounters. The scenarios also excluded presentations that actors could not portray authentically, including cases where audio-visual perception may carry more diagnostic value.
Targeted automated evaluations found occasional perception and reasoning errors. Google also reports intermittent technical issues that can interrupt conversational naturalness. Project Astra remains a prototype, with system-level technical considerations outside this medical application. Google states that real patient research is the next stage.
The company has begun related work in clinical settings with the text-based AMIE. A feasibility study with Beth Israel Deaconess Medical Center provided initial evidence on safety and utility in clinical practice, Google says. An ongoing nationwide randomised study with Included Health is evaluating AI in real-world virtual care.
The Google study provides controlled evidence on video consultation behaviour, physical-examination guidance, and clinician scoring. It does not yet provide evidence that AMIE can safely diagnose or manage real patients in production.
See also: Novo Nordisk and AWS bring agentic AI into drug discovery
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.



