Robots that fold laundry and load dishwashers are trained on recordings of humans folding laundry and loading dishwashers. Nobody has enough of those recordings. That shortage is what this work is: companies building physical AI models pay people to generate demonstration data, either by driving a real robot through a task or by recording themselves doing the task.
The reason to take it seriously as an income path is that the rates are published on public job boards, so you can check them before you commit an evening. The reason to be careful is that it is contract work rather than a business. You are selling hours, someone else owns the asset those hours build, and the project ends when the dataset is full. Everything below is written with that distinction held in view, because it changes what the work is good for.
Language models had the internet. Robots have no equivalent. There is no vast pre-existing archive of a hand rotating a doorknob, recorded from the right angle with synchronised force and position data. It has to be manufactured, one demonstration at a time, by people.
The industry describes this openly. The Robot Report's August 2026 survey of physical AI infrastructure platforms characterises one major provider's approach as combining "centralized data factories, distributed human collectors, real robot systems, and robotless egocentric collection, with multimodal annotation and internal policy fine-tuning". Read that phrase slowly, because it is a job description hiding inside industry vocabulary. "Distributed human collectors" means people at home with a phone. "Robotless egocentric collection" means recording a task from your own point of view without any robot present. "Data factories" means shift work in a building full of robot arms.
Those three phrases map almost exactly onto the three tiers of work available, and the pay gap between them is roughly fivefold.
The shape of this job is a direct consequence of how the last fifteen years of machine learning went, and knowing that helps you predict where it goes.
Image recognition was solved, in commercial terms, by assembling a very large labelled dataset and then scaling models against it. Language followed the same pattern with a decisive advantage: the dataset already existed, because humans had spent thirty years typing the internet. Nobody had to commission it. The scaling recipe worked because the data was free and abundant, and a generation of the field internalised the lesson that capability follows data volume.
Robotics inherited the recipe and none of the data. There is no web-scale corpus of physical manipulation. What existed instead was thousands of small academic datasets, each recorded on a different robot, in a different lab, with different cameras and different conventions, which could not be pooled. A model trained on one lab's robot did not transfer to another's.
Two things changed around the middle of the decade. Shared data formats and open tooling made pooling across robot types practical, so a demonstration recorded on one arm became useful to someone with a different arm. And the field established that recordings of humans, captured from a first-person view, carry enough signal to be worth training on even with no robot involved. That second finding is what created the remote tier of this work. It converted robot data collection from something only a lab could do into something a person with a phone and a specification could do.
The consequence is that the bottleneck moved from hardware to logistics. The scarce resource is no longer robots, it is organised human hours producing consistent recordings across enough environments and enough variation. That is a staffing problem, and it is why the listings exist.
Every figure in this section comes from a live listing on a company's own job board rather than from a salary aggregator's estimate, which matters because the aggregate figures for this role are pulled from a mix of very different jobs.
Two structural facts fall out of that table. Pay tracks the hardware you are trusted with, not your experience. And the entry tier is genuinely low-paid, which nobody advertising this work says clearly, so it needs saying: at $10 an hour, phone capture is worth doing to get a track record with a platform, or as flexible fill-in income, and it is not worth reorganising your life around.
Headline rates attached to famous companies spread quickly in this field. A widely repeated figure puts teleoperation work at one large automaker at $48 an hour, and another claim has a data platform paying $50 an hour for people to record themselves. Both appear in secondary write-ups and social posts. I could not verify either at the hiring company's own careers page, so treat them the way you would treat any income claim with no primary source behind it: possible, unconfirmed, and not a basis for planning.
Salary aggregators report an average around $28 an hour for "robot teleoperation" in the United States as of mid-August 2026, with most of the range between roughly $22 and the mid thirties. That number is directionally consistent with the on-site listings above, but aggregates for a job title this new pool genuinely different roles, so the specific listing in front of you is better evidence than any average.
The practical rule: apply to the posting, read the stated range, and disregard the number you saw on social media.
A data buyer is not collecting videos of chores. They are collecting a dataset with consistent properties: a fixed camera position, a stated frame rate, particular lighting conditions, a defined start and end state for the task, a required number of repetitions, no faces or identifying documents in frame, no other people, sometimes a specified hand to use. A recording that captures a beautiful, natural, useful-looking demonstration and violates one of those constraints is worth nothing, because it cannot be pooled with the rest.
This is why the first week is where most people quietly wash out. The task looks trivial, the spec looks like bureaucracy, and then the submissions get rejected and the effective hourly rate collapses well below the posted one. The operators who make the posted rate are the ones who treat the spec as the deliverable and the chore as incidental.
Three habits carry most of the benefit. Read the spec end to end before the first recording, then record one clip and submit it alone for feedback before batching twenty. Set up a fixed physical rig, a mount position, a marked spot on the floor, the same lamp, so compliance is a property of your setup instead of something you have to remember. And keep a short log of what was rejected and why, because rejection reasons repeat and the second occurrence is avoidable.
Listings use a small vocabulary for the categories of data they want, and recognising the words tells you what a role involves before you apply. A posting from one humanoid robotics company for a Data Collection Operator describes operating humanoid robots "during data collection activities, including locomotion, manipulation, teleoperation, and human-robot interaction scenarios". Those four terms cover most of the field.
Long-horizon and multi-step tasks are the current frontier: clear the table, then load the dishwasher, then wipe the surface. Individually easy steps, collectively hard, because the model has to track a goal across minutes rather than seconds. Expect these specs to be the fussiest and the most likely to be revised mid-project.
Cutting across all four is what buyers call coverage, and it is the thing most collectors misunderstand. A dataset of a thousand flawless mug pickups in one kitchen is close to worthless. Variation is the product: different lighting, different surfaces, different object positions, cluttered scenes, unfamiliar objects, left hand as well as right. When a spec asks you to move a lamp between takes or record the same task at three times of day, that instruction is the entire point and not an inconvenience.
The corollary is genuinely counterintuitive. Failure is often requested. Specs increasingly ask for the fumble and the correction, the grasp that slips and the recovery, because a model trained only on success has no idea what to do when something goes wrong. New collectors instinctively delete those takes. Read the spec before you do.
The Kit for Remote Capture
The on-site tier supplies its own hardware. For remote work the equipment burden is small but specific, and getting it wrong is the most common cause of rejected batches.
A hands-free head or chest mount. The listings ask for "hands-free, first-person" video, which rules out holding the phone. A chest harness is steadier and cheaper; a head mount gives a viewpoint closer to where a robot's cameras sit and is usually what specs want. The requirement is repeatability: the same mount at the same angle every session, so your captures pool with your own earlier work.
A phone that records long clips without thermal throttling. Extended high-frame-rate capture heats phones, and a phone that drops frames or stops recording at minute eleven will fail a spec silently. Test a full-length capture before your first paid batch.
Storage and upload bandwidth. This is the cost nobody budgets. High-frame-rate first-person video is large, batches are many gigabytes, and buyers want the original file rather than anything re-encoded. If your upload speed is slow, that time is unpaid, and it can quietly halve an effective hourly rate at the entry tier.
Controlled, boring lighting. Consistent artificial light beats good natural light, because daylight changes between takes and the spec cares about consistency more than beauty. A cheap lamp in a fixed position solves more rejections than any camera upgrade.
A marked setup. Tape on the floor for where you stand, tape on the counter for object start positions. This is what turns compliance into a property of your kitchen rather than something you have to concentrate on.
Where a role involves an instrumented rig, such as the UMI-style gripper in the Argentina listing, the buyer ships it. Read the terms on that hardware carefully: you are usually responsible for it, and returning it at the end of the contract is a condition of final payment.
Who Is Actually Hiring
The buyers divide into three groups and it helps to know which one you are talking to.
Data platforms and collection companies. These are the businesses whose entire product is datasets and annotation services for robotics teams. They run the labs, recruit the distributed collectors, and hold the contracts with model developers. Most of the listings you will find come from here, and OpenTrain AI is one example of the type. Their work is steady, their specs are strict, and they are the most likely to have something for a beginner.
Robot developers hiring directly. Humanoid and manipulation companies post their own data collection and teleoperation roles, sometimes titled Data Collection Operator or Robot Operator. These tend to be on-site, better paid, and more competitive, and they often expect you to be comfortable near expensive hardware.
Annotation and infrastructure vendors. Companies providing the tooling layer also staff human work around it. Roles here drift toward reviewing, labelling and quality-checking other people's captures rather than generating your own, which is a different and often more sustainable job.
For finding real listings, the company job boards are better than the aggregators, because aggregators lag and duplicate. Searching the specific vocabulary works better than searching "robot jobs": teleoperator, data collection operator, egocentric video, demonstration data, physical AI, UMI gripper.
Spotting the Scams, Which Are Already Here
Any income category with a real published rate and a low skill floor attracts fraud within months, and this one has arrived at that point.
The reliable signals are old ones. Nobody legitimate asks you to pay for training, certification, equipment access or a "collector kit" before you can start earning. Nobody legitimate needs your bank details before an offer exists. A real contract states a rate, and a posting that describes earnings only as a potential monthly total is avoiding the hourly number for a reason.
Two signals are specific to this field. First, real collection work comes with a written specification, and a lot of it; an offer with no spec attached is not a data job. Second, legitimate work has narrow geographic eligibility, because of tax status, data protection law and shipping. A listing that will hire anyone anywhere with no eligibility conditions is either not real or not paying what it says.
One more, less obvious: be wary of anything asking you to record inside a workplace, a school, a hospital or anywhere with other people who have not consented. That request is a sign the buyer is careless, and the consequences of carelessness there land on you rather than on them.
Getting Through the Application
The hiring process is lighter than the pay range suggests, which is the main reason this work is worth an application even if you are sceptical.
Screening for the remote tiers usually consists of a short form, an eligibility check on your location, and a sample capture. The sample is the real filter. You are sent a specification and asked to produce one compliant clip, and the reviewer is not judging how well you fold a towel. They are checking whether you can follow a written spec exactly, because that predicts everything about whether you are cheap or expensive to work with.
Treat the sample accordingly. Follow the spec literally, including the parts that seem pointless. If something is ambiguous, ask before recording rather than guessing, and say what you interpreted it to mean. A candidate who asks one precise clarifying question reads as low-risk; a candidate who submits a confidently wrong interpretation reads as someone whose batches will need re-reviewing.
For the on-site tier, expect the additional questions to be practical rather than technical: shift availability, whether you can commit to the stated minimum hours, whether you can stand and move through a facility for a full shift, and how you handle repetitive work. Nobody is testing robotics knowledge. They are testing whether you will still be there in week six, because turnover is their expensive problem.
Two things help disproportionately in an application. Any evidence of precision work in your history, whether that is machining, lab work, surgery, music, sewing or competitive gaming, is worth naming. And a willingness to take an unpopular shift is real leverage at a facility that has to staff a midnight crew.
How It Compares to the Alternatives
Worth situating against the other things you might do with the same hours, because the honest answer depends on which tier you can access.
Against delivery and rideshare work, on-site teleoperation compares well: comparable or better hourly rates, no vehicle costs, no mileage depreciation, no weather, and a fixed schedule rather than surge chasing. The phone-only tier at $10 to $15 an hour compares badly once you account for unpaid upload time, and it wins only on flexibility and on not needing a car.
Against data annotation and labelling work, this is the same industry one layer down the stack. Annotation pays less per hour at the entry level and is more automatable, since a model can increasingly pre-label and leave a human to confirm. Physical data collection resists that, because there is no way to synthesise the original recording of a hand doing a thing. If you are choosing between them for durability, collection is the safer side of the same bet.
Against skilled remote freelancing, this loses on ceiling and wins on entry. Nobody needs a portfolio, a niche, a proposal or a client relationship to start, and that is genuinely rare. But there is no version of this work where the rate climbs with reputation the way a specialist's does, so it functions best as a bridge rather than a destination.
Against a conventional part-time job, the differences that matter are classification and stability. You carry your own tax, you get no sick pay, and the contract ends when the dataset fills. The compensation for that is schedule control at the remote tier and, at the on-site tier, a rate above most hourly work available without credentials.
The clearest way to hold it: this is well-paid unskilled work in a field that currently needs more hands than it has, which is a temporary condition. Use it for what temporary conditions are good for, which is cash now and a foot in a growing industry, rather than treating it as a plan for 2030.
Rookie Mistakes
Treating the posted rate as your rate. The posted rate applies to accepted work. Rejections, setup time between takes and unpaid reading of the spec all sit between the two numbers. Track your real hours against your real earnings for the first two weeks and you will know what tier you are actually in.
Batching before validating. Recording forty clips against a spec you have misread converts a day into nothing. One clip, submitted, confirmed, then batch.
Improvising for the camera. New collectors perform the task, making it smooth and legible and unnaturally tidy. Some specs want exactly that, and some explicitly want natural execution including fumbles and corrections, because a model that has only seen flawless demonstrations cannot recover from error. Do what the spec says rather than what looks good.
Ignoring the physical demands of the on-site tier. An eight-hour shift standing and moving through a facility, hitting productivity targets, four possible shift slots including a midnight to 6 AM crew, is industrial work. The listing says so plainly. People apply imagining a desk and a joystick.
Letting a project end without a next one. The Argentina listing is roughly four weeks. That is normal: datasets fill and the work stops. Line up the next platform while the current contract runs, and keep two or three relationships alive rather than one.
Skipping the classification question. These are contractor roles. Nobody is withholding tax for you, and depending on where you live you may owe quarterly payments and self-employment contributions on this income. Budget for it in month one rather than discovering it at year end.
Gotchas Worth Knowing Before You Apply
This is a job, not an asset. You trade hours for money and the dataset belongs to the buyer. There is no compounding, no equity in what you built, and no residual when the model trained on your data ships. That is a legitimate reason to do it, and a bad reason to expect it to become something else.
Geography gates the good tier. The $30 to $55 range is attached to on-site work in a US lab. If you do not live within commuting distance of one, that tier is closed to you regardless of skill, and the remote tiers pay a third as much. This is the single biggest determinant of what this work can be worth to you, and it is decided entirely by your postcode.
The night shift is real. A 24/7 facility staffs a 11:45 PM to 6:00 AM crew. Those hours are often easier to get and they carry a genuine health cost that a slightly higher rate does not offset.
Your home becomes the set. Phone-based capture means recording inside your kitchen and living space, repeatedly, to spec. Household members appear in frame and have to be kept out of it. Some people find this fine and some find it quietly intolerable after two weeks.
Nondisclosure is standard and enforced. You will likely sign an agreement covering the client, the hardware, the specs and the tasks. This restricts what you can say publicly, which in turn makes it hard to research the field, which is part of why so much of the available information about rates is secondhand.
The spec can change mid-project. A buyer refining what their model needs will revise requirements, sometimes invalidating a capture approach you had optimised. Ask how revisions are handled and whether already-accepted work stays accepted.
What You Are Signing Away
The paperwork on this work deserves more attention than a $12 an hour job would normally get, because what you hand over is unusual.
You are recording the inside of your home, your hands, your movement patterns, and often your voice. Those recordings go into a training set, and training sets are copied, retained, resold between companies and used to build models that persist long after your contract ends. There is no practical mechanism for retrieving your footage later. Assume anything you record is permanent.
Four questions are worth asking before you sign, and a legitimate buyer will answer all four.
What is the scope of the licence you are granting? Most agreements take a broad, perpetual, worldwide licence to use the captured data for model training. That is normal in this field. What varies is whether it extends to your likeness for other purposes, and whether they can sublicense to third parties. Those two clauses are the ones to read twice.
Who else appears in your footage, and did they agree? Housemates, family, children, and anything visible through a window. Most specs prohibit other people in frame, partly for data quality and partly because the buyer does not want consent problems. That prohibition protects you too, and violating it can void a batch after you have been paid for it.
What is visible that you did not intend? Post, screens, prescription labels, keys, documents, a laptop showing an email. A first-person camera at counter height sees a lot. Clear the set the way a photographer would, every session.
How is payment tied to acceptance? There is a real difference between being paid for hours worked and being paid for accepted captures. The OpenTrain teleoperation listing states compensation is "hourly and paid ... for the work you perform during scheduled shifts", which is the favourable structure. Piece-rate acceptance models put the cost of an ambiguous spec on you. Ask which one applies and get it in writing.
Turning It Into Something You Own
The honest limitation of this work is that it is hours for money. There are three ways people convert it into something with more leverage, in increasing order of difficulty.
Move to quality control. Reviewing and validating other people's captures pays more than producing them, requires the spec fluency you build in the first months, and scales without your hands being in frame. This is the most reliable step up and the one buyers most often need filled, because their bottleneck is reviewers rather than collectors.
Become a collection coordinator. Buyers who need coverage across many environments have a logistics problem: recruiting, briefing, spec compliance and quality across a distributed group. Someone who has done the work, can read a spec and can manage people is a natural fit, and this is where the work starts resembling a business with margin rather than a shift.
Run a collection site. The capital-intensive version: a space, several rigs, a small trained crew, and a contract with a buyer for volume. This is a real business with real risk, it depends on a small number of customers, and the customers may bring it in house at any point. Worth understanding as the ceiling of this path rather than as a plan.
What all three have in common is that the transferable asset is not your dexterity. It is your fluency with specifications and your reliability, both of which are legible to a buyer only if you have a record with them. That is the strongest argument for treating even the $10 an hour tier seriously while you are in it: the work is worth little, and the record is worth something.
Behind the Scenes: What a Shift Actually Involves
The on-site version is closer to precision manufacturing than to gaming.
You arrive for a shift that overlaps the outgoing crew, because the facility does not stop. You are assigned a cell: a robot arm or a humanoid, a set of objects, a task list. The task is narrow and repetitive by design. Pick up the mug, place it on the rack, in this orientation, thirty times, with variation in starting position because a model trained on one starting position learns one starting position.
You drive the robot through it, often through a control interface with imperfect force feedback, which is why the work is tiring in a specific way: your hands know what the task should feel like and the robot does not transmit it. Depth perception through a camera is worse than you expect. The first day involves knocking things over.
Between takes there is reset work. Objects returned to start, the cell tidied, occasionally a snag escalated to a technician. Reset time is a real fraction of the shift, and how briskly you reset is much of what a productivity target measures.
Periodically a batch gets reviewed and some of it comes back. A gripper occluded the object at the moment of contact. The demonstration drifted outside the workspace. The lighting changed when someone opened a door.
The remote version has the same shape at smaller scale. Rig the phone, mark the floor, run the task, check the clip, reset the kitchen, run it again. The hard part is not the chore. It is doing the chore identically twenty times while a phone watches, and staying accurate on the twentieth.
Where This Goes Next
Some reasoning about the direction, offered as reasoning rather than prediction.
The entry tier gets squeezed first. Phone-only egocentric capture is the least differentiated thing a human can supply, and it is the tier most exposed to synthetic data and simulation improving. Expect $10 an hour to stay flat in nominal terms and drift down in real terms, while the wearable and on-site tiers hold better because the hardware in the loop cannot be faked as easily.
Specification skill becomes the differentiator. The collectors who last will be the ones who can read a demanding spec, hit it consistently, and flag ambiguities before wasting a batch. That is a quality-control skill, and quality control roles pay more than collection roles. The realistic career path out of this work runs through reviewing other people's captures, not through collecting faster.
Geography spreads, then concentrates. Remote tiers will keep widening to more countries, because the buyers want diverse environments and cheaper hours. On-site labs will concentrate near the robot developers. The gap between the two tiers is more likely to widen than to close.
Contracts stay short. Dataset-shaped demand produces project-shaped work. Plan for a portfolio of platforms rather than a job, and treat every contract as ending on schedule even when the manager implies otherwise.
The information problem persists. Nondisclosure agreements plus fast-moving rates mean the public record on what this work pays will stay thin and unreliable, which is exactly why the discipline of reading the posting and ignoring the rumour will keep being worth more than any rate guide, including this section.