AMIE learned to look at you. That is the news.
This is the first credible demonstration that a general-purpose AI can conduct a synchronous telehealth consultation at parity with a primary-care physician on the core competencies — and above PCPs at the one skill telehealth has always been weakest at, physical observation. The clinical-practice implication is not "AI replaces GPs." It is that the ceiling on asynchronous, low-cost triage just moved up by a full modality.
TL;DR
- Google Research and Google DeepMind published "Towards expert-level medical AI for real-time video consultations" on 11 August 2026 (Google Research).
- AMIE (Video) — built on Gemini and Project Astra — ran a randomised, blinded OSCE-style study: 100 scenarios, 300 live consultations, 15 trained patient actors, 30 board-certified primary-care physicians, with a 20-PCP evaluation panel scoring against standard rubrics (Google Research).
- Across history-taking, diagnostic accuracy, management appropriateness and communication, AMIE (Video) was rated on par with PCPs. On eliciting physical signs and guiding examination manoeuvres, AMIE (Video) was rated significantly higher on average than both PCPs and the text-only AMIE baseline (Google Research).
- Patient actors preferred the video interface to text chat and rated AMIE (Video) favourably on empathy, rapport and "confidence in care" (Google Research).
- Study is simulated, not real-patient. Google explicitly says real-world validation is required before any clinical conclusions (Google Research). A nationwide randomised real-world study with Included Health is already running (Google Research).
What actually shipped
The paper describes AMIE (Video) as an asynchronous multi-agent system — three specialised agents running in parallel (Google Research):
- Talker agent — patient-facing, low-latency spoken interaction.
- Planner agent — background clinical reasoning, differential-diagnosis updating, information-gap identification.
- Perception agent — continuous review of the audio and visual streams, flagging clinically relevant non-verbal cues.
The architectural point is not that Google put a face on a chatbot. It is that they decoupled the conversational tempo problem from the deep reasoning problem. A single agent trying to do both either stalls (patient loses trust) or shallows out (diagnosis suffers). Splitting the labour is what makes it feel like a doctor and reason like one at the same time (Google Research).
The study covered five body systems — cardiopulmonary, abdominal, HEENT (head/eyes/ears/nose/throat), neurological/psychiatric, and musculoskeletal — with three arms: AMIE (Video), AMIE (Text) as an ablation baseline, and PCPs consulting via the same video interface (Google Research).
What it actually means
The interesting result is not that AMIE matched PCPs on diagnostic accuracy. Prior AMIE work already suggested that in text-based OSCEs (Nature, Towards conversational diagnostic AI, 2024–25 line of work referenced in the paper). The interesting result is that AMIE (Video) beat PCPs at eliciting physical signs and guiding examination manoeuvres over video (Google Research).
That inverts a decade of assumptions about telehealth. The standard critique of remote care has been: you cannot palpate an abdomen through a screen, you cannot really assess gait, so telehealth is inherently a lesser modality. The AMIE result says the bottleneck was never the video — it was clinician attention. A tireless perception agent, watching every frame for cramped handwriting, laboured breathing, subtle asymmetries, is doing something a distracted PCP on a 12-minute slot cannot.
That is a very different framing of the AI-in-medicine story than "AI diagnoses better than doctors." It is closer to: AI is better at paying attention than doctors have time to be.
Hype deconstruction
Three things this study is not.
One — it is not a real-world clinical result. Every consultation was with a professional patient actor. Google is unusually direct about this: "Patient actors, however skilled, cannot fully replicate the complexity and unpredictability of real clinical encounters, and the scenarios were limited to conditions that can be authentically portrayed through acting" (Google Research). Actor-mediated OSCEs systematically favour AI because actors present textbook symptom scripts. Real patients wander, minimise, forget medications and lie about drinking.
Two — it is not a regulatory-clearance moment. No FDA, TGA, EMA or MHRA authorisation is being sought on the basis of this study. AMIE (Video) is a research configuration of Project Astra. It is not a product (Google Research).
Three — it is not evidence AMIE can manage care over time. The study measured a single consultation. Longitudinal disease management — dose titration, follow-up, watching for adverse events — was the subject of a separate AMIE paper (Nature-line work on multi-visit management, June 2026) and remains an open problem for the video modality.
Stakeholder landscape
- Primary-care physicians — not immediately displaced. But the "GPs are irreplaceable because of physical exam" defence just got weaker. Expect uncomfortable conversations about scope-of-practice and reimbursement bundling for AI-assisted virtual visits.
- Telehealth vendors (Teladoc, Amwell, Doctor Anywhere, Included Health) — this is a moat problem. Included Health is Google's clinical study partner (Google Research); everyone else is now watching a research demo they cannot ship.
- Health systems and payers — the economics of triage change if a Gemini-derived agent can perform pre-visit assessment at PCP quality. That is the volume layer where GP shortages bite hardest.
- Regulators (FDA, TGA, MHRA, EU AI Office) — AMIE (Video) sits squarely in what the EU AI Act classifies as high-risk AI in healthcare. The moment this leaves the lab, the regulatory pathway becomes the story.
- Patients — the patient-actor preference finding is the most quietly important result. If patients prefer talking to an AI over a text chat and an AI is warmer on empathy scores than a rushed PCP, the adoption curve is not going to be gated on clinical performance. It will be gated on trust, indemnity and payment.
Cross-layer implications
- Model architecture. The multi-agent asynchronous pattern — a fast Talker, a slow Planner, a continuous Perception stream — is going to generalise well beyond medicine. Any domain that needs conversational latency and deliberate reasoning (legal intake, complex customer service, tutoring) now has a reference architecture to copy.
- Project Astra. This is the first serious clinical application built on Astra's real-time multimodal stack. It moves Astra from "impressive demo" to "load-bearing platform."
- Competitive positioning. OpenAI's medical work (HealthBench, o1 clinical benchmarks) and Anthropic's clinical reasoning results have been strong on text. Google now holds the audio-visual clinical benchmark and the real-world study partnership. That is a hard combination to catch.
- Health-economics. The story of the next 24 months is not "AI diagnoses cancer." It is "AI takes the pre-visit intake" — the highest-volume, lowest-margin, worst-experience part of primary care.
What this means for you
For the general reader: nothing changes at your next GP appointment. Nothing changes at your next telehealth visit either. But the model of care you experience in 2028–2029 — especially for triage, minor complaints, and post-discharge follow-up — is likely to include an AI video agent doing the first pass. Ask, when it happens, three questions: Who reviews its recommendations? What happens if it misses something? Is my data being used to train future systems?
For clinicians and health-service leaders: the question worth putting on the agenda now is not "should we adopt AI video triage." It is "what does our clinical governance framework look like for an AI system that observes patients?" AMIE (Video) is a research prototype; the pattern it demonstrates will ship, from Google or someone else, inside 18 months. RACGP, ACRRM, RCGP and AAFP position statements will lag the technology; internal policy should not.
For health-tech builders: the multi-agent decoupling is the transferable insight. If you are building a clinical assistant on GPT-4o, Gemini 2.5 or Claude Sonnet 4.5 and running everything through one model, you are going to lose on latency, reasoning depth, or both. Split the loop.
Uncertainty ledger
- Real-world patient performance — unknown. Included Health study results are the next major data point.
- Failure modes on non-white, non-English-language, or disabled patients — under-reported. The 30 PCPs and 15 patient actors are not characterised by demographic breakdown in the released materials (Google Research).
- Adversarial-input behaviour — malingering, factitious disorder, coached symptom presentation — untested.
- Regulatory pathway — no clear jurisdiction has yet articulated how a general-purpose multimodal AI performing clinical assessment gets authorised.
- Data-governance question — where does the video go, how long is it retained, and under what consent framework? Not addressed in the research post.
Bottom Line
AMIE (Video) is the strongest published evidence yet that a foundation-model AI can conduct a clinical consultation at the same level as a primary-care physician — and outperform them on physical observation. It is a simulated study, not a clinical trial, and Google has been unusually careful to say so. But the direction of travel is now legible: within two to three years, the first pass of a telehealth visit will not be conducted by a human. The regulatory, indemnity and workforce arguments about that will be the actual story. The technical argument is effectively over.
Sources
- Google Research, Advancing AMIE towards expert-level audio-visual clinical consultations (11 Aug 2026) — Tier 1 (primary)
- Google Research, Towards expert-level medical AI for real-time video consultations (paper, 11 Aug 2026) — Tier 1 (primary)
- Google Research, Collaborating on a nationwide randomized study of AI in primary care (Included Health partnership) — Tier 1
- Nature Medicine coverage of AMIE multimodal ECG/imaging study (May 2026) — Tier 1 (background)
- News-Medical, AI beats primary care doctors in simulated diagnosis study using images and ECGs (May 2026) — Tier 2 (context)