Give your robot a supervisor
Robots have gotten good at moving, but they still get stuck. The next wave of robot intelligence is an AI supervisor, a large vision-language model that watches the robot, figures out what went wrong, and gets it moving again.
Robots learned to move
Over the last few years robots got genuinely good at moving. Fast control models and drive stacks run on the robot itself and handle the gripping, balancing, and steering. They work like the robot's reflexes. They react in milliseconds, and they have to live onboard because motion can't wait on a network.
But robots still get stuck
Moving and judging are different skills, though. A chair blocks a doorway, a person steps into the path, a cup slips out of the gripper, or the elevator shows up full twice in a row. The reflexes are fine. The robot just doesn't know what should happen next, so it stops and waits for help. In robotics this is called an impasse.
Today, people watch the robots
Right now the industry solves this with people. Fleet operators hire staff to sit at screens, watch the robots, and click "go around" whenever one freezes. It works, but one person can only watch a few machines, so the watching bill grows with every robot you deploy.
The better fix is an AI supervisor
An AI supervisor is a vision-language model doing that same job. It looks at the robot's cameras, reads the scene, and picks the next move from a short list of skills the robot is allowed to use, things like retry, reroute, wait, slow down, or call a human. It can also point at objects, answer questions, and walk the robot through a tricky spot, which is really just teleoperation by a model instead of a person.
The way we keep this safe is a dual-brain architecture. The low-level control runs on the robot, in its local control policy or drive stack, while the supervisor only advises. It picks what to do, and it doesn't force how the robot moves. Balance, collision stops, and emergency behavior stay onboard where they can react instantly, while the supervisor's answer can afford to take half a second, because it arrives over an API.
This wave has already started
- Anthropic just published Claude Plays Robotics, where they let language models drive robot bodies at every level, from raw motor commands up to high-level steering. Models that had to drive the joints themselves "mostly fail," in their words, while the same models supervising a pretrained controller completed real navigation and manipulation tasks. A frontier lab reached the same conclusion we build on.
- Waddle calls itself "Claude Code for robots." Its agents watch the camera feeds, break a goal into steps, and write programs that call skills and local policies, so the language model plans and checks the work while the motor loop stays with the specialists.
- Pigey puts a vision-language model in the loop as the decision maker, thinking every few seconds while frozen skills handle the motion. There's a demo where the robot's target is hidden under a container, and the model notices, reasons it through, clears the container, and carries on, which is exactly the job a supervisor is meant to do.
We built one, so we know it works
Robomart runs delivery vehicles on RoboMind, our full self-driving system. It drives well on its own, but streets are full of stuck moments, so we trained Fathom, a supervisor for RoboMind. Fathom stays quiet while RoboMind drives, and when the vehicle hits an impasse it wakes up, reads the scene, and picks the safe next move. It decides what to do, and the steering stays with RoboMind. We wrote the approach up in our technical paper, Dream Reckoning. Fathom is why RSI exists. We needed a supervisor for our own robots, and we think every robot company is about to need one too.
Why this is the path to RSI
Good judgment takes a big model, and big models don't fit on a robot, which is why the supervisor lives in the cloud. The cloud is also where it learns. Every stuck moment a supervisor clears becomes training data, so you can tune the model on your own robots' episodes, promote the new revision, and the whole fleet gets smarter at once. Round after round, the supervisor improves on the work it just did.
That loop is our mission. Supervisory intelligence is the first step, and followed far enough it becomes robot superintelligence. Both spell RSI.
RSI serves supervisors today. Pick a vision-language model that fits your robot, deploy it as a Policy with a stable name, tune it when the weights are open, and call one route. Open and frontier supervisors. One RSI API.