Skip to content
Robotensor

About Robotensor

Physical intelligence has to be shaped, and then it has to be proven.

Robotensor is built on two convictions: that a foundation model becomes a working capability only once it is shaped around a real robot, a real task and a real environment — and that the only credible way to know whether it works is to test it, in the open, against every other attempt to do the same thing.

Technical plate of a robot arm handling labware through a repeatable protocol.
shaped to the task · proven in the open

Why Robotensor exists

A foundation model is not a capability.

Vision-language-action models generalise across a huge space of possible bodies, tasks and scenes. That breadth is what makes them powerful, and it is exactly what makes them unfinished. A gripper is not a suction cup. A kitchen counter is not a warehouse tote. A five-millimetre tolerance is not a rough grasp. None of that is resolved by scale alone.

Turning a general model into a working capability means shaping it around the specific body that will run it, the specific task it has to complete and the specific environment it has to complete that task in — then checking, honestly, whether it can. That shaping and that checking are infrastructure. They are the same work regardless of which model, which robot or which task walks through the door, and building them once, properly, is the reason Robotensor exists.

Intelligence is what the robot can do.

What we work on

Four problems, one instrument.

The research and the infrastructure are the same programme seen from two sides: every question below is answered by something that has to be measurable, and everything we build exists to make one of them measurable.

Robot foundation models

General-purpose policies that transfer across bodies and tasks, and the shaping that turns that generality into a capability on one specific machine.

In-context robot learning

Adaptation from a demonstration or from interaction, with the weights frozen — the setting our public competition is built to measure.

Multimodal control

Vision, language and action closed into a single loop, so an instruction and a scene resolve into motion rather than into a plan nobody can execute.

Open evaluation

Benchmarks and evaluation harnesses that measure what a policy can actually do, published with the records that back them.

Why an open competition network

Benchmarking and evaluation, run where anyone can watch.

Benchmarking and evaluation do not stay behind closed doors here. Robotensor runs them as a live competition network: one public protocol every entry is measured against, and one public result every entry has to keep defending.

01

Benchmarking, made public

Tasks, variations and success criteria are defined once and applied identically to every model that competes — not authored around whichever model is being shown off.

closed suite → open protocol

02

Evaluation, made adversarial

A model does not clear a static test once and stop being asked. It has to hold its position against a live challenger, on camera, or lose it.

one-time score → defended standing

  1. 01

    The king of the hill

    One entry holds the top of the hill at a time: the model with the strongest standing on the published tasks.

  2. 02

    Challenger duels

    Any registered entry can challenge the king. Both run the same held-out episodes, on the same tasks, under the same conditions.

  3. 03

    A decision, not a vote

    A duel resolves on objective outcome — success rates over a minimum number of decided units — not on a demo reel and not on a narrative.

The network does not take a model’s word for it. It makes the model prove it, in public, against a king that already has.

How we work

Five commitments, each of them falsifiable.

A value that cannot be checked is decoration. These are the ones the platform is actually built around — if we stopped honouring one, it would show up in the network, in public, the same day.

  1. 01

    Proof over claim

    A result nobody can reproduce is a press release. Every standing on the network comes from a signed record, measured on the same tasks under the same conditions as every competitor, and published where anyone can go and disagree with it.

  2. 02

    Shaped to the real machine

    Generality is where we start, not where we finish. A gripper is not a suction cup and a five-millimetre tolerance is not a rough grasp, so the work is to shape a model around the specific body, task and environment it has to succeed in — and then to check whether it did.

  3. 03

    Adversarial by default

    A model that cleared a static test once is not asked to stop being tested. Standings are defended against live challengers or they are lost. We would rather find our own failures on camera than have someone else find them in a deployment.

  4. 04

    Build the instrument once

    Shaping and evaluation are the same work no matter which model, robot or task arrives next. We treat that as infrastructure to be built once and built properly, rather than rebuilt informally for each new demo.

  5. 05

    Honest instruments

    When data is missing, our tools say so. An empty state rather than an invented number, an error rather than a plausible guess, the actual outcome rather than the story that would land better. This holds internally before it holds in public.

Careers

We are hiring for the parts of this that are still hard.

Robot foundation models, evaluation infrastructure, simulation, and the systems work that keeps a public competition network honest. If the commitments above sound like how you already work, come and hold us to them.

See it running.

The competition network is live: every duel and every standing is published in full.