Axis Robotics Open-Sources One of the Largest | Crypto News
Axis Robotics has launched Axis Sim Dataset V1, one of the largest open-source simulation datasets for Franka arm manipulation, with the full dataset, training code, and benchmarks publicly out there. V1 is constructed from more than 50,000 human-teleoperated simulation trajectories across 207 manipulation duties and 60,000+ scene variants on a simulated Franka Research 3 arm.
This dataset drew over 160,000 downloads, making it the most downloaded open-source simulation Franka manipulation dataset on Hugging Face. In benchmarks, continuous pretraining on V1 lifted π0.5 and beat a volume-matched RoboCasa baseline, with every consequence open and verifiable.
Axis Robotics is building the final compounding data engine for Physical AI, a vertically built-in system spanning large-scale simulation, selfish real-world seize, humanoid loco-manipulation, and human-gated DAgger post-training. The company raised $12 million in seed funding led by Hack VC, with participation from Nomad Capital, Pi Network Ventures, 10K Ventures, and angel traders.
A Bet Against “Clean Data Only”
A common assumption in robotics is that demonstrations must be near-optimal to start with — filter down to knowledgeable trajectories, standardize the setup, and discard something noisy before it’s protected to imitate. Axis’s thesis runs the other method: data high quality lives at the distribution stage, not the single trajectory. When a large and various enough crowd produces noisy, suboptimal trajectories and their errors are uncorrelated, the noise averages out and a working coverage survives during training.
Axis Sim Dataset V1 places that thesis to a public check. Its trajectories span pick-and-place, stacking, pouring, articulated-object manipulation, and instrument use, all collected through Axis’s browser-based teleoperation platform, Axis Hub, by a distributed crowd quite than a single knowledgeable group. The dataset was constructed with researchers from UC Berkeley, Johns Hopkins, the University of Michigan, and other establishments.
Results That Scale
On LIBERO-Plus, continuous pretraining on V1 lifts π0.5 from 83.9% to 88.8% success and outperforms a volume-matched RoboCasa365 baseline by 37.3%. Performance improves persistently as pretraining data scales from 25% to 100% of the dataset, with no saturation in sight, evidence that the beneficial properties come from range and coverage quite than a one-off bump. The largest enhancements seem under digicam, sensor-noise, and format perturbations, the precise axes Axis randomizes during era.
The group says V2 is already underway, scaling to 1.2 million trajectories across 1,200 duties, with cross-embodiment generalization and outcomes across a number of VLA fashions exhibiting that suboptimal simulation data trains sturdy insurance policies.
The Engine Behind the Dataset
The dataset is one output of a bigger, actively compounding data engine. Where a conventional data vendor collects to a fixed spec and stops, Axis makes use of model efficiency and failure instances to decide what needs to be collected next, so every training spherical informs the next. That engine runs on a hybrid strategy across 4 data strains, and all 4 now run at scale:
- Simulation: over 200,000 distributed contributors on Axis Hub, a top-3 dApp on Base, producing 4.7M+ trajectories across 13 embodiments.
- Egocentric: a managed community of 1,000+ full-time, QC-trained collectors capturing first-person exercise in real houses and companies across 14 industries: 200,000+ hours already banked and growing by 4,000+ hours every day, with Vicon-verified hand pose.
- Loco-manipulation: 500+ hours combining mobility and dexterity on real humanoids (Unitree G1, Booster T2) through hardware-agnostic teleoperation.
- Human-gated DAgger post-training: 500+ hours of human-in-the-loop correction focused at deployment edge instances.
Every activity and trajectory is recorded on-chain on Base for provenance, and contributors are rewarded for verified work high quality.
From Open Data to Commercial Deployment
Beyond open-sourcing simulation data, Axis works instantly with robot embodiment firms to construct custom-made, embodiment-specific data pipelines and model priors.
As Booster Robotics’ first sim-data companion, Axis rebuilt Booster’s real workspace as a task-aligned digital twin, had distributed contributors accumulate 42,000+ simulation episodes on it, and distilled them into a Booster-specific model prior. With just 30 real-robot demos, that prior reached 87.5% success versus 37.5% for an out-of-the-box Ï€0.5, matching Ï€0.5 utilizing half the real-world demonstrations.
Other companions span embodiment firms (Feagine Robotics), model firms (Manycore Tech, Dexmal) and industrial automation (Lotus Cars, Geely Auto). Axis also provides on-chain robotics networks: BitRobot on Solana and OpenRoboto on Bittensor.
Redefining Physical AI’s Data Foundation
“The future of Physical AI isn’t a static dataset you download once,” said Chris Feng, founder of Axis Robotics. “It’s an engine that keeps producing the data the model needs next. Scale gets you broad coverage. Diversity keeps the noise unbiased. The closed loop turns every failure into progress. That’s what compounds.”
Axis was based by researchers from UC Berkeley, CMU, Georgia Tech, and SJTU, alongside serial founders who have scaled shopper platforms to over 30 million customers. Its research is suggested by Jiachen Li, Assistant Professor at Georgia Tech.
Paper Link: https://arxiv.org/abs/2607.21588
Project Page: https://axisaiorg.github.io/AXIS-V1/
Dataset Link: https://huggingface.co/datasets/axisrobotics/Franka-Dataset
Github Codebase: https://github.com/AxisAIOrg/Axis-V1-Training
Stay up to date with the latest trending crypto news! Visit our web site daily for the freshest Crypto news and content, fastidiously curated to keep you informed.



