Robots do not learn from the internet.
Language models inherited a corpus. Embodied models never got one. There is no Common Crawl for a hand closing around a door handle, catching a slipping cup, or finding a seam by feel. That data does not exist on the web because it was never a document. It has to be captured, on purpose, by people, at a scale no research lab can film for itself.
Surge built the expert-data layer for language models. Behalf is building it for robots.
Request samples and the catalog ↗
Ten thousand hours of human hands.
Egocentric manipulation, captured on professional multi-sensor rigs by trained operators, and delivered in MCAP with intrinsics and extrinsics attached, not as a folder of MP4s.
- BinocularHead-mounted stereo RGB, 1920×1080 at 30 fps, H.265, IMU at 200 Hz, 3-DoF orientation. Raw or annotated.
- AnnotatedVIO/SLAM tracking, 2328×1748 at 30 fps, 6-DoF head pose, 2D and 3D hand keypoints per frame.
- Ego + gloveDexterous glove: 14-DoF finger joint angles, 460-cell tactile pressure, 16-bit depth stream.
- Ego + wristFirst-person RGB plus dual wrist-mounted cameras, for close-range contact and occlusion recovery.
- Ego-exoFirst-person camera plus four-corner exocentric positioning cameras, time-synchronized.
- Four / six cameraQuad-fisheye tracking cameras with binocular RGB, full calibration, 9-DoF IMU.
Two capture modes. Professional-device capture in controlled environments: stable lighting, full hand and forearm visibility, natural two-hand manipulation, synchronized streams. Mobile-device capture in real physical environments: kitchens, workshops, stores, homes, complete two-hand interaction sequences shot single-take. Both privacy-screened before delivery.
World models need frames with actions attached.
Video alone teaches a model what the world looks like. It does not teach what an action does to it. Our world-model corpora ship the control signal alongside the pixels: frame-level action and camera JSON, EXR depth maps, player and camera pose, keyboard and controller inputs synchronized to under ten milliseconds.
Multi-view 3D trajectories at twenty-five viewpoints per scene. SfM depth and camera-pose sets across first-person, third-person, objective, and vehicle viewpoints, indoor and outdoor. Top-down and bird’s-eye families for planning. Interaction data spanning roughly a hundred synthetic environments, HUD removed, ambient audio intact.
And real-world motion priors at volume: seven hundred thousand sports-motion clips of continuous human action, plus in-store, human-object, and try-on video, single-shot, unoccluded, stable-camera. The kind of footage humanoid policies need and consumer video dumps do not contain.
Finished inventory, not a roadmap.
- 10,000 hrsegocentric manipulation capture
- 300 hrscustom capture capacity, per day
- 6+rig configurations, from stereo to six-camera
- 1.5M+human-motion and interaction clips
- 70,000+ hrslicensed and produced video
Robotics and world models are where we go deepest, and they are not all of it. The same operation produces expert task data in the formats labs evaluate on: thirty thousand expert-authored tasks and rubrics, eight thousand GUI agent trajectories on real desktops, verifier-backed terminal and coding tasks, and a reinforcement-learning environment built on the fully de-identified operating data of a real four-hundred-person retail chain under documented rights authorization. The volumes above are on the shelf. The full catalog, with per-dataset formats and specifications, is available on request.
Tell us the task. We will go film it.
Most of what an embodied team actually needs has never been recorded by anyone. Name the task family, the environments, the rig, and the annotation schema, and we run it as a production: operator recruiting, capture, QA, and delivery in your format. Capacity runs to three hundred capture-hours a day, anchored in Asia, where the crews, the rigs, and the physical environments are cheap enough to iterate on.
That geography is also our second edge. Asia’s professional and physical world is the least-represented slice of every frontier training set, and we recruit inside it: in-country channels, local languages, credentials verified against local registries, paid test work graded by senior experts before anyone touches production data.
Every datapoint has a paper trail.
Who captured it, under what consent, what license it transfers under, and which export mechanism applies. We keep the answer per sample, not per contract. When your general counsel asks where the data came from, the answer is a lookup, not a meeting.
Where a dataset touches enterprise or personal data, it ships de-identified under documented data-rights authorization. Media corpora clear rights confirmation before commercial delivery, not after.
How it works.
Ask for the catalog. Name the datasets or the task family that interests you. We send free samples with full specifications: format, volume, capture method, and pricing. Approve the sample and we deliver at volume, or produce to your exact spec.
Practicing professionals and experienced capture operators: write to experts@usebehalf.com.
See the data before you believe a word of this. Request samples.