HYDERABAD — AI research giant Google DeepMind has officially unveiled Gemini Robotics 2, a next-generation suite of Physical AI models designed to give humanoid robots full-body coordination, advanced finger dexterity, and autonomous multi-robot collaboration.

The launch introduces three specialized models under the new framework: Gemini Robotics 2, Gemini Robotics ER 2, and Gemini Robotics On-Device 2. Together, these systems mark a major leap forward from simple desktop-bound robotic arms to mobile, autonomous humanoids capable of navigating and manipulating complex real-world environments.

Unlocking Whole-Body Intelligence and Human-Like Dexterity

At the core of the announcement is Gemini Robotics 2, a multi-modal Vision-Language-Action (VLA) model that translates visual inputs and natural language prompts into precise, synchronized motor commands. Unlike previous-generation systems that restricted AI reasoning to upper-body movements, Gemini Robotics 2 coordinates a humanoid’s entire frame—from feet to fingertips.

In live demonstrations using Apptronik’s Apollo 2 humanoid robot, the model processed single natural language instructions such as "put the watering can into the green bin on the bottom shelf." The AI calculated full-body kinematics, instructing Apollo 2 to walk across the room, crouch down, reach into tight shelf spaces, and place the object accurately.

Beyond full-body movement, the new architecture delivers significantly improved dexterity. The model successfully controlled Apollo 2’s 22-degree-of-freedom hand to perform fine motor skills—including tying knots, sealing ziplock bags, and executing intricate kitchen chores. It also showed seamless adaptability when operating standard two-fingered parallel grippers for tight spatial packing.

Multi-Robot Collaboration and Local Processing

To tackle complex, long-horizon tasks, DeepMind introduced Gemini Robotics ER 2, an Embodied Reasoning (ER) and Vision-Language Model (VLM). Acting as the high-level brain, ER 2 enables robots to analyze video in real time, communicate with humans, and break down multi-step goals lasting several minutes into sequential actions.

Crucially, ER 2 introduces multi-robot orchestration. In a garage cleanup test, Apollo 2 worked alongside a "Duo" dual-arm robot. The model divided the workload, routed each machine to specific locations, and determined exact hand-off points for transferring objects between robots.

Rounding out the lineup is Gemini Robotics On-Device 2, an ultra-efficient VLA model optimized to execute directly on local robotic hardware without requiring internet connectivity or incurring cloud latency. Leveraging motion-transfer techniques, On-Device 2 can adapt to entirely new robot form factors—even those with vastly different shapes, sensor configurations, and degrees of freedom—in just a few hours using fewer than 200 training examples.

While Google DeepMind acknowledged that operational speeds still require optimization before reaching full human tempo, the Gemini Robotics 2 family marks a critical milestone toward bringing practical, versatile humanoid helpers into homes and industrial workplaces.