Careers
Research Engineer, Benchmarking, Robotics
- Engineering
- Global (China preferred) · Remote
- Full time
About this role
About Robotensor
Robotensor is building the intelligence and evaluation infrastructure for general-purpose robots.
We work on robot foundation models, embodied AI, multimodal learning, simulation, and real-world robot evaluation.
As robot models become more general, measuring their capabilities becomes increasingly difficult. A benchmark that works for a language model is not enough for a physical system: robot performance depends on embodiment, environment, initial conditions, perception, control, latency, object variation, and many other factors.
We want to build rigorous ways to measure what robot models can actually do.
About the Role
We are looking for a Research Engineer, Benchmarking, Robotics to design the benchmarks and evaluation systems we use to measure robot intelligence.
You will define tasks, evaluation protocols, metrics, and experimental methodology for robot foundation models.
Your job is not simply to run tests.
You will decide what should be tested, how it should be tested, and what constitutes meaningful progress.
You will work with researchers and robotics engineers to turn questions about model capability into reproducible experiments across simulation and real robots.
The output of your work should make it possible to answer questions such as:
- Is the new model actually better?
- What capabilities improved?
- What capabilities regressed?
- Does performance generalize to new objects and environments?
- Can the model recover when something goes wrong?
- How sensitive is performance to initial conditions?
- Does a benchmark score correspond to useful real-world robot capability?
- Are we measuring the model, or accidentally measuring the setup?
What You’ll Do
- Design benchmarks for robot foundation models and general-purpose robot policies.
- Define task suites that measure manipulation, perception, reasoning, adaptation, recovery, and long-horizon behavior.
- Design evaluation protocols for both simulation and physical robots.
- Develop clear success criteria and quantitative metrics for robot tasks.
- Build benchmark environments with controlled variations in objects, poses, scenes, instructions, and initial conditions.
- Design tests for generalization across unseen objects, environments, tasks, and embodiments.
- Develop stress tests that expose model limitations and edge cases.
- Build evaluation frameworks for comparing different robot models and checkpoints.
- Create reproducible procedures for real-world robot evaluation.
- Develop automated and semi-automated evaluation infrastructure.
- Build tools for recording trajectories, videos, observations, actions, failures, and evaluation metadata.
- Design failure taxonomies and diagnostic tools for understanding why policies fail.
- Measure variance and uncertainty in physical robot evaluations.
- Determine how many trials are required to make meaningful comparisons between models.
- Build dashboards and visualization tools for tracking model performance and regressions.
- Establish evaluation gates for model releases.
- Work with researchers to design targeted evaluations around new model capabilities.
- Compare simulation, offline, and real-world evaluation results and understand where they disagree.
- Help develop public benchmarks and evaluation protocols that can be used by the broader robotics community.
Benchmark Areas
You may design evaluations around areas such as:
Manipulation
Pick-and-place, rearrangement, articulated objects, tool use, deformable objects, contact-rich tasks, and bimanual manipulation.
Generalization
Changes in object identity, position, appearance, lighting, backgrounds, camera viewpoints, robot embodiments, and environments.
Language Following
Whether the robot correctly understands and executes natural-language instructions, including compositional and ambiguous commands.
Robustness
Performance under perturbations, imperfect perception, calibration changes, distractors, latency, and unexpected physical events.
Recovery
Whether the robot can recognize failure, modify its behavior, and recover instead of simply terminating the episode.
Adaptation
How quickly a model can acquire a new capability from demonstrations, context, interaction, or small amounts of additional data.
Long-Horizon Tasks
Whether a robot can successfully execute sequences of actions while maintaining task state over extended periods.
Cross-Embodiment Performance
How capabilities transfer between robot arms, grippers, mobile manipulators, humanoids, and other embodiments.
What We’re Looking For
- Strong software engineering skills, particularly Python.
- Experience with robotics, machine learning, computer vision, or embodied AI.
- Experience designing experiments and interpreting experimental results.
- Strong understanding of evaluation methodology and statistics.
- Ability to turn ambiguous concepts such as “generalization,” “robustness,” or “reasoning” into measurable experiments.
- Experience building reproducible data and evaluation pipelines.
- Strong debugging and failure-analysis skills.
- Ability to work with large amounts of experimental data.
- Comfort working across software, machine learning, simulation, and physical robot systems.
- Strong attention to experimental confounders and measurement quality.
Nice to Have
- Experience evaluating robot-learning or Vision-Language-Action models.
- Experience designing robotics benchmarks.
- Experience with manipulation benchmarks or robot task suites.
- Experience with ROS / ROS2.
- Experience with simulation environments such as MuJoCo, Isaac Sim, ManiSkill, RLBench, or similar platforms.
- Experience with hardware-in-the-loop testing.
- Experience running large-scale robot experiments.
- Experience developing benchmark datasets.
- Experience with statistical analysis of noisy physical experiments.
- Experience with automated evaluation pipelines.
- Publications or open-source work related to robotics evaluation, benchmarks, robot learning, or embodied AI.
What Success Looks Like
Researchers at Robotensor can answer whether a model is better without relying on cherry-picked demonstrations.
A model improvement can be broken down by capability rather than represented by a single opaque success rate.
Real-world evaluations are repeatable enough that different model versions can be compared meaningfully.
Failures are categorized and measurable rather than stored only as videos and anecdotes.
Benchmark tasks are difficult enough to reveal meaningful differences between frontier models.
Evaluation results influence model architecture, training data, and research priorities.
And ultimately, Robotensor develops a rigorous picture of what current robot foundation models can—and cannot—actually do.
Apply
