Teehoo Embodied · Robot data CRO

Task-specific robot data — ready to simulate, train, and deploy.

Every pilot ships a sim-ready scene, a 10× augmented LeRobot v2 trajectory dataset, and a fine-tuned LoRA adapter — drop-in for OpenVLA, π0, RT-X, Octo, and RDT. Four-to-six week engagement. PIPL-audited end-to-end. Delivered onshore by design.

Onshore capture
Anhui / Jiangxi
10× augmentation
data flywheel
Cross-gripper
3+ embodiments
Audit-ready
SHA-256 chain

The problem

Your robot's specific task data is the part open datasets don't solve.

AgiBot World, Open X-Embodiment, and LeRobot have made generic robot data almost free. Backbone policies pretrained on them already cover a huge surface. None of them, however, cover your specific task on your specific robot at your specific factory.

Open datasets
Cover generic skills, not yours

1M+ household trajectories on someone else's humanoid. Useful backbone. Does not contain a single demo of your SO-100 pouring your customer's reagent into your customer's vial.

In-house collection
Consumes a quarter of engineering capacity

Capture rigs, teleop operators, scene reconstruction, augmentation, cross-gripper retargeting, adapter training, and a compliance audit chain. Eight engineering disciplines, one quarter minimum.

Cross-border vendors
Trigger PIPL cross-border data review

Sending Chinese-collected video to an overseas annotation desk invokes formal cross-border data review under PIPL. Slower, costlier, and the auditable lineage becomes somebody else's responsibility.

What you get

A three-piece bundle that lands in your training stack.

Every pilot produces the same three artifacts. Each is a standalone deliverable; together they let your team simulate, train, and ship a new policy on day one of the bundle landing.

Isometric workcell · table-mounted robot arm with tripod camera observing a target object
Piece 1 · Sim-ready scene

A digital twin of your workcell.

Run simulation, evaluation, and policy testing before touching a real robot. Geometry, mass, and friction tuned to your actual environment so the sim and the real workcell agree.

Formats: URDF + MJCF · drop-in for Isaac Sim, Genesis, SAPIEN, MuJoCo

Sequence of six robot arm poses across time with a velocity profile sparkline below
Piece 2 · Trajectory dataset

Demonstrations multiplied 10× across grippers.

Real on-site demonstrations grown via jitter, time-stretch, pose-offset, and Blender re-rendering. Cross-gripper retargeted across Franka Panda, SO-100, and Allegro Hand so the dataset isn't locked to one embodiment.

Format: LeRobot v2 + Open X-Embodiment compatible · HuggingFace dataset card included

Flow diagram · stacked base model layers, plus a small adapter delta box, equals a task-specific deployable skill
Piece 3 · LoRA adapter

A drop-in task skill for your policy stack.

Load into your robotics stack and start evaluation immediately. Fine-tuned on your choice of backbone — OpenVLA-7B, π0, RT-X, Octo, or RDT — trained onshore and delivered onshore-only.

Format: HuggingFace PEFT · backbone choice at pilot kickoff

How it's made

Six steps. Same shape, every pilot.

Every Teehoo pilot runs the same six-stage pipeline. Each stage emits an audit event into the bundle's hash chain. You can verify the chain after delivery with one CLI command.

Six-stage Teehoo pipeline from left to right: Capture · Reconstruct · Scale · Retarget · Train · Deliver
  1. 01

    Collect

    Onshore capture — video, teleop, hand demos.

  2. 02

    Reconstruct

    Real workcell → URDF + MJCF digital twin.

  3. 03

    Scale

    10× via jitter, time-stretch, pose-offset, Blender.

  4. 04

    Retarget

    Remapped across Franka, SO-100, Allegro.

  5. 05

    Train

    LoRA on OpenVLA / π0 / RT-X / Octo / RDT.

  6. 06

    Deliver

    tar.gz · dataset card · SHA-256 audit chain.

Stage outputs and per-stage backend modes are visible inside every delivery's Honesty Matrix tab — REAL, MOCK (deterministic stub), or FIXTURE (bundled sample) per stage, with no hand-waving.

PIPL audit chain · onshore by design

Every byte you receive can be traced back to the camera that captured it.

Teehoo runs the data ops inside the PRC and delivers the trained adapter onshore-only. Every data-handling event is SHA-256 hash-chained against the previous one — verifiable from your laptop with the bundled teehoo-pipl verify command.

The audit chain is structural verification, not a third-party PIPL compliance certificate. External attestation is a Stage 5 roadmap item.

Onshore by design

Source video, intermediate artifacts, and trained adapter weights never leave PRC borders.

Hash-chain audit

Every event references the SHA-256 of the previous one. Tamper with one, the chain breaks visibly.

Lawful basis logged

Every event references the pilot's lawful basis identifier so your counsel can inspect the full ledger.

See it on a real delivery

Explore the Water Pouring pilot — AgiBot · SO-100 · onshore Anhui.

A four-episode demonstration of an actual delivery: the three-piece bundle, the HuggingFace dataset card, the PIPL audit chain, and the single-command reproduction recipe. No signup; we'll put you in as the demo customer.

Open the demo delivery