Google DeepMind Launches Gemini Robotics 2: AI-Powered Control for Versatile, Task-Capable Robots
Google DeepMind has unveiled an updated version of its Gemini artificial intelligence model built for robotics, engineered to control a wide range of robotic hardware — including humanoid robots that can handle dexterous everyday tasks from screwing in lightbulbs to tying sealed trash bags.
Gemini Robotics 2 merges several distinct AI models into one unified system, which works together to help robots interpret their physical surroundings and plan appropriate, goal-aligned actions. A core vision language model (VLM), which processes images and video input, enables natural communication with humans and carries out logical reasoning to map out how to complete different assigned tasks. Two separate vision language action (VLA) models, trained to understand movement in physical 3D space, oversee both the robot’s full-body locomotion and fine motor control of grippers or robotic hands.
In pre-release demonstration videos shared ahead of the official launch, DeepMind showed multiple different robot platforms completing complex, multi-step tasks fully autonomously using the combined AI system. In one demo, Apptronik’s Apollo 2 humanoid robot used custom robotic hands from firm Sharpa to tidy and rearrange store shelves.
DeepMind trained the new model using a mixed approach of human teleoperation, recorded task video examples, and simulated training environments. For now, targeted training is still required for AI systems to master a broad scope of complex physical tasks.
While industry rivals Anthropic and OpenAI have pulled ahead in consumer chatbots and AI coding tools, Google holds a longer, stronger track record in robotics AI research, having published foundational work on using large AI models to train robots to complete practical, useful tasks. This latest release is another clear signal the search giant is betting that AI must move beyond the digital realm to reach its full potential. (Google previously partnered with Boston Dynamics, the global leader in legged robots, to develop AI “brains” for the company’s robotic hardware.)
“This is another milestone on our path toward what we call physical AGI — a system that lets a robot do almost any task a human can do,” Carolina Parada, head of robotics at Google DeepMind, told WIRED.
That said, connecting cutting-edge frontier AI models to physical robots that can move through homes, workplaces, and manipulate real objects comes with unique elevated risks. Past research has confirmed that frontier AI-controlled robots can produce unexpected, sometimes dangerous behavior. The risk of unplanned or harmful actions from advanced AI agents was thrown into sharp relief recently, when an unreleased AI agent developed by OpenAI demonstrated the ability to hack multiple digital systems.
“The safety question is even more pressing when you’re deploying these systems in open physical environments,” Parada said. “A lot of uncertainty will pop up that we can’t anticipate, so we need to understand the safety challenges far more deeply.”
Parada notes Google uses a multi-layered approach to safety, with built-in guardrails applied to every individual model layer in the system. The company is also launching ASIMOV-Agentic, a new benchmark designed to measure the safety of multiple collaborating AI systems that control a single robot. The benchmark scans commands to flag whether they would result in harmful or unacceptably uncertain outcomes before the robot acts.
Demis Hassabis, CEO of Google DeepMind, previously told WIRED that his long-term goal is to build a universal AI operating system for all types of robot hardware, modeled after the Android operating system that powers a diverse ecosystem of smartphones.
Google DeepMind Launches Gemini Robotics 2: AI-Powered Control for Versatile, Task-Capable Robots