Careers
Data Engineer, Embodied AI
- Engineering
- Global(prefer China) · Remote
- Full time
Data Engineer, Embodied AI
About this role
About RoboTensor
RoboTensor is a Physical AI research and development company building the infrastructure that turns robot foundation models into specialized physical capabilities.
For embodied AI, model performance is inseparable from training data. Physical intelligence depends on what experiences a model sees, how those experiences are represented, and how quickly failures in the real world can become better training data.
About the Role
We are looking for a Data Engineer, Embodied AI to build the data infrastructure behind policy learning.
You will own systems for generating, ingesting, processing, storing, curating, and serving large-scale embodied datasets—including video, images, robot state, actions, trajectories, language, simulation data, and evaluation results.
This is not a traditional analytics data-engineering role. Your customers are the models and the engineers training them.
You will work closely with embodied AI, simulation, and robotics engineers to continuously improve the quality and usefulness of the data entering our training pipelines.
What You’ll Do
- Build end-to-end data pipelines for embodied AI training.
- Develop systems for ingesting data from physical robots, simulation, demonstrations, existing datasets, and synthetic generation pipelines.
- Design representations and schemas for multimodal robot trajectories.
- Process and synchronize video, images, actions, robot state, sensor information, language, and metadata.
- Build systems for filtering, cleaning, deduplicating, balancing, and curating datasets.
- Develop tools for automatically identifying useful, difficult, failed, or unusual episodes.
- Build pipelines that convert policy failures into candidate training data.
- Create dataset versioning, lineage, reproducibility, and quality-control systems.
- Build efficient storage and loading systems for large-scale model training.
- Develop sampling and mixture strategies for training across tasks, robots, environments, and data sources.
- Work with ML engineers to understand how dataset composition affects policy performance.
- Build tools for inspecting and debugging training datasets.
- Improve the speed with which newly generated real-world or simulated experience becomes usable training data.
- Help establish quantitative measures of dataset quality and coverage.
What We’re Looking For
- Strong software and data engineering fundamentals.
- Experience building reliable, high-throughput data pipelines.
- Strong Python skills.
- Experience working with large datasets and distributed storage or compute systems.
- Ability to design clean schemas and abstractions for complex multimodal data.
- Strong understanding of data quality, reproducibility, and observability.
- Ability to reason about data from the perspective of machine learning rather than analytics alone.
- Comfortable working closely with researchers and changing systems rapidly as training requirements evolve.
Nice to Have
- Experience building datasets for deep learning or foundation-model training.
- Experience with video, image, multimodal, sequential, or trajectory data.
- Experience with embodied AI, robotics, autonomous systems, or reinforcement learning.
- Experience with large-scale dataset generation or synthetic-data pipelines.
- Experience with distributed training input pipelines.
- Experience with data labeling, active learning, or automated data selection.
- Familiarity with imitation learning or learning from demonstrations.
- Experience working with simulation-generated datasets.
What Success Looks Like
Training engineers spend their time improving models rather than fighting data infrastructure.
They can quickly answer questions such as: What data do we have? Where did it come from? Which tasks and conditions are underrepresented? Which failures should we add to training? Did changing the data mixture improve the policy?
Most importantly, RoboTensor develops a tight data flywheel: every simulation run, physical experiment, evaluation, and policy failure can become better training data for the next model.
Apply
