Vallum Labs
White paper · v0 · September 2026
Vallum Labs collects consent-cleared, action-labeled, first-person video of real outdoor manual work and licenses it to robotics and world-model labs. Now collecting in the Western Cape, South Africa. September 2026 to January 2027.
Team
Founded in 2026 by Jaiyen Shetty. Grew up farming in Fresno, California. Computer science and business at Berkeley. Previously at Meta, the UN, the Gates Foundation and Amplitude. Based in Cape Town through January 2027, then the next region.
Hardware
Head and chest mounted iPhone rigs, four in the field. Monocular RGB at 1920x1080, 30 fps, HEVC at capture, camera settings locked for the whole shift. Not in v0: IMU, depth, wrist cameras, per-phone intrinsics. Each of those is a hardware decision that waits for a buyer to ask for it.
Operations
Workers wear rigs during their normal paid shifts on commercial farms. Nobody is filmed doing anything they would not be doing anyway. Current tasks, Southern Hemisphere spring and summer:
- Valencia orange picking, Citrusdal and Piketberg, through October
- Blueberry picking, George, Paarl and Wellington, October to November
- Strawberry picking, Stellenbosch, September to December
- Vine suckering and shoot thinning, from October
Next verticals: construction, mining, forestry.
Consent and provenance
Every worker signs a written consent form in English, Afrikaans or isiXhosa before the rig goes on, and is paid for wearing it. Every clip carries a consent ID and a SHA-256 hash of the raw file. Faces are blurred and audio is stripped before any file leaves the farm. Farms are coded; deliverables carry a region, never coordinates. Built to POPIA, South Africa's data protection law.
Annotation
Schema v0 follows Ego4D and Ego-Exo4D conventions so any lab can read it on sight: timestamped narrations at 8 to 13 per minute, verb and noun action segments per hand, and pre, point-of-no-return and post critical frames with hand and object boxes propagated with SAM 2. Each segment also carries the Gemini Robotics 2 ER benchmark fields (progress bucket, critical-moment timestamp, success or failure with a fault code). Delivery is one Ego4D-shaped JSON per clip, plus a datasheet, verb and noun lists, consent summary and lineage manifest per batch. Send us your schema before we film and we map it onto this instead of rebuilding.
Evaluation
Every seventh clip of every session is held out, annotated to the same schema, and never delivered or shown. Labs can score policies against it. Outdoor manual work has no sim-to-real benchmark yet. This is the start of one.
Datasets
- Outdoor-20
- Valencia citrus harvest, Western Cape, 20 curated hours
- Oct 2026
- Outdoor-200
- Citrus, blueberries, strawberries, vines, 200 hours
- Feb 2027
- Outdoor-2K
- Agriculture and construction, 2,000 hours
- 2027
Datasets are licensed. Research groups get gated samples first.
Farms
We pay your crew, blur every face, and send you a picking report and a three-minute first-person training film within a week of the shoot. The footage never carries your farm's name.
Researchers and labs
If you train on human video and your corpus has no outdoor work in it, we fill that gap to your spec. Tell us the fields you need before we film. Footage cannot be reshot to a different schema. Paid pilots run 20 to 40 curated hours, delivered with a datasheet and held-out scores.
Write to jaiyen_shetty@berkeley.edu.
Reading
- Ego4D: Around the World in 3,000 Hours of Egocentric Video 2021
- Ego-Exo4D: Understanding Skilled Human Activity 2023
- EgoMimic: Scaling Imitation Learning via Egocentric Video 2024
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data NVIDIA GEAR, 2026
- HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining 2026