Careers
Research Engineer – Robot Foundation Models
- Engineering
- Global (China preferred) · Remote
- Full time
About this role
About Robotensor
Robotensor is building the intelligence layer for general-purpose robots.
We work on robot foundation models, multimodal learning, in-context adaptation, simulation, and real-world evaluation. Our goal is to build models that can transfer across tasks, environments, and robot embodiments—and turn broad learned intelligence into reliable physical capability.
About the Role
We are looking for a Research Engineer – Robot Foundation Models to develop the models that power Robotensor’s physical intelligence stack.
You will work on multimodal models that connect vision, language, robot state, action, and temporal context. You will train and evaluate models on large-scale robot data, develop new learning methods, and test whether improvements translate into better behavior on real robots.
This role sits between research and engineering. You should be comfortable reading recent papers, developing new ideas, implementing them from scratch, running large-scale experiments, and debugging failures in real robotic systems.
You will not just integrate existing models. You will help determine what robot foundation models should look like.
What You’ll Do
-
Design, implement, and train robot foundation models and Vision-Language-Action models.
-
Develop multimodal architectures combining vision, language, proprioception, robot state, actions, and temporal context.
-
Explore architectures for general-purpose manipulation and embodied intelligence.
-
Develop methods for transferring capabilities across tasks, environments, and robot embodiments.
-
Work on in-context learning and fast adaptation from demonstrations, trajectories, and interaction.
-
Investigate action representations, tokenization strategies, prediction horizons, memory, and temporal modeling.
-
Develop training strategies using imitation learning, offline learning, reinforcement learning, and large-scale pre-training.
-
Build data mixtures and training objectives for heterogeneous robot datasets.
-
Explore diffusion policies, flow-based models, autoregressive policies, world models, latent-action models, and other emerging architectures.
-
Train models across large GPU clusters and improve training efficiency, stability, and reproducibility.
-
Build evaluation pipelines for measuring generalization, adaptation, robustness, and long-horizon task performance.
-
Run controlled experiments and ablations to understand what actually improves model capability.
-
Analyze model failures and determine whether limitations come from architecture, training, data, perception, control, or embodiment.
-
Work closely with robotics, simulation, and infrastructure engineers to deploy learned policies on physical robots.
-
Translate promising research ideas into systems that work outside controlled benchmarks.
-
Contribute to the technical direction of Robotensor’s foundation-model research.
Research Problems You May Work On
Cross-Embodiment Learning
How can one model learn reusable physical knowledge across different arms, grippers, sensors, action spaces, and robot platforms?
Multimodal Robot Learning
How should vision, language, proprioception, actions, and history be represented and combined inside a general robot model?
In-Context Adaptation
Can a robot learn a new task from a few demonstrations or interactions without retraining the entire model?
Action Generation
What is the right representation for robot actions? Continuous actions, action chunks, latent actions, tokens, trajectories, or learned abstractions?
Long-Horizon Behavior
How can models maintain context, recover from mistakes, and complete tasks that require many coordinated actions?
Generalization
How can a policy trained on one distribution continue working when objects, viewpoints, environments, or embodiments change?
Learning From Real-World Data
How should robot trajectories, failures, demonstrations, teleoperation data, simulation data, and internet-scale visual data be combined?
What We’re Looking For
-
Strong experience with modern deep learning and neural network architectures.
-
Strong Python and PyTorch skills.
-
Experience training models rather than only using pretrained APIs or checkpoints.
-
Strong understanding of transformers, attention-based architectures, and multimodal learning.
-
Experience designing and running machine learning experiments.
-
Ability to independently implement ideas from research papers.
-
Strong debugging skills across models, datasets, training pipelines, and evaluation systems.
-
Comfort working in an open-ended research environment where the correct architecture or training method may not yet exist.
-
Ability to move quickly between theory, implementation, experimentation, and real-world testing.
Experience in one or more of the following is particularly relevant:
-
Robot learning
-
Vision-Language-Action models
-
Multimodal foundation models
-
Imitation learning
-
Reinforcement learning
-
Offline RL
-
Generative modeling
-
Diffusion or flow-based models
-
World models
-
Representation learning
-
Large-scale sequence modeling
-
Distributed model training
Nice to Have
-
Experience training robot foundation models or VLA models.
-
Experience with manipulation or mobile manipulation.
-
Experience working with large robot trajectory datasets.
-
Experience training models across multiple robot embodiments.
-
Experience with diffusion policies or action-chunking architectures.
-
Experience adapting vision-language models for robotics.
-
Experience with test-time or in-context adaptation.
-
Experience with simulation-to-real transfer.
-
Experience deploying learned policies on physical robots.
-
Experience with distributed GPU training and large-scale ML infrastructure.
-
Research publications or strong open-source contributions in robotics, machine learning, multimodal learning, or embodied AI.
Academic credentials are useful, but demonstrated ability to build and train difficult systems matters more.
What Success Looks Like
You build models that produce measurable improvements in real robot capability.
Models trained on one task or embodiment transfer useful knowledge to others.
New tasks require less task-specific data and less retraining.
Evaluation becomes systematic enough that we can distinguish real capability improvements from benchmark noise.
Failures on physical robots turn into better architectures, datasets, objectives, and training methods.
Over time, Robotensor develops a foundation-model stack that becomes increasingly general rather than a collection of isolated task-specific policies.
Who This Role Is For
You may currently be a Research Engineer, Machine Learning Engineer, Robotics Researcher, or Research Scientist.
The title matters less than your ability to build.
You should enjoy reading a new paper in the morning, implementing the important idea that afternoon, training it at scale, and then discovering on a real robot why the idea does—or does not—work.
We are looking for engineers who want to help define how general-purpose robot intelligence is built.
Apply
