Two answers to the same medical question sit side by side on your screen. One of them is wrong in a way that could hurt someone. Your job is to say which, and write the paragraph explaining why. The lab building the model pays you for that paragraph, because you know something its software does not.
Now think about your current work. Maybe you are a nurse, a lawyer, an engineer, a teacher, and you feel stretched thin for pay that stopped keeping up years ago. Maybe you worry that AI will take part of your job. Here is a way to get paid by the very companies building it, using the expertise you already have.
The window is open because the labs need qualified humans right now, and they need many of them. The specialists who join early get the steadier projects, the trusted reviewer status and the better queues. Those who arrive later start at the bottom of the same platforms and wait longer for work.
The facts. Starting costs you nothing, because the lab supplies the work and you supply the judgment. The entry tier pays $12 to $25 an hour and puts you in competition with everyone on the planet. Above it sits work that asks for a credential or real domain depth. Earnings across this work start at $500 a month, and the reasoning you attach to each judgment is what separates a $30 hour from a $150 one. Time to first money runs one to two months. Starting without a credential? Our AI training jobs guide compares the generalist platforms, from Outlier to DataAnnotation, and what each really pays.
Tonight, write one line naming the field you know better than most people do. That line is what a screening task is checking.
Mercor, one of the platforms matching experts to labs, publishes its own rate data. The national average for AI training work is $31 an hour as of early 2026, with rates ranging from $12 to over $200 depending on experience, specialisation and platform.
That range tells you the whole story, because it spans three different jobs. Which of the three are you qualified for today?
Mercor's own explanation of why this tier pays $75 to $200 and up is worth quoting directly, because it is the whole thesis: these roles "require the ability to judge, analyze, evaluate, and improve upon outputs that most workers cannot assess. As a result, there's a genuine supply constraint, which leads to higher pay rates."
So where would your own background land you in that table if you applied tonight? Most readers with a professional licence guess one tier too low.
A model can produce a fluent, confident, well-structured answer about drug interactions, contract enforceability, tax treatment or structural load. Producing that answer is now cheap. Deciding whether it is correct stays expensive, and another model cannot do it for you, because if the model knew the answer was wrong it would not have produced it.
That is the supply constraint. The labs need people who can look at an output in a specialist field and say, with authority, that it is wrong and exactly why. The set of people who can do that for medicine is roughly the set of people who trained in medicine.
Four kinds of work follow from this.
That repetition deserves a pause, because it is the clearest signal in this whole collection of guides. Across unrelated professions, each asked separately what stays valuable as tools absorb the routine work, the answer keeps arriving in the same shape: somebody qualified has to confirm the output is correct, and their qualification has to be in the domain rather than in the tool. A paralegal verifying that generated citations exist and remain good law, a coder confirming an assigned code matches the documentation, a designer checking a generated asset survives a print spec, and an oncologist catching a plausible dosage error are all doing one activity in four settings. This guide is about selling that activity straight to the organisations building the systems, at the rates they pay for it, instead of waiting for it to arrive inside your existing job.
Think about your last working week. How many times did you catch something a junior colleague or a tool got confidently wrong? Each of those moments is the skill this market pays for.
The odd thing about this work is that most people who qualify for the top tier are not looking for it, because they do not know it exists. You may be one of them.
There is a direct link to the profession guides elsewhere on this site. The paralegals facing 0.2% employment growth, the designers at 2.1%, the professionals watching AI absorb parts of their work: the same expertise being partly displaced is what the labs pay $75 to $200 an hour for. That makes this the most direct route available from a squeezed credential to a well-paid use of it.
That settles which tier your background points at. Everything from here is about crossing the line from qualified to hired.
The process is lighter than the rates suggest, and most of it is about proving your credential.
Could you give the assessment one quiet afternoon this month, phone in another room? If you can, that afternoon is probably the best-paid unpaid time in this whole guide.
You know how to get in front of a lab. What the work is worth once you are in comes next.
The rates are real, and the caveats matter.
Work is intermittent. Labs commission projects, projects fill, and the work stops. A $150 hourly rate on twenty hours in a month is a very different month from the same rate on a hundred hours, and the platforms do not guarantee volume.
The average is $31 and the top tier is $75 to $200. Both figures are true, and they describe different people. If your credential is general rather than specialist, plan against the median of $25 to $33 and leave the headline aside.
Entry-level rates are poor and getting worse. At $12 to $25, that tier competes with global labour and with automation of the simpler tasks. Do it only as a way to show reliability, and never treat it as your plan.
Be wary of earnings screenshots. Mercor's own guidance opens by acknowledging the claims going around about people making seven or eight thousand dollars a month at this, and the honest explanation is plain arithmetic: those figures need the top tier and close to full-time hours, sustained, which very few people reach because volume is not guaranteed at any tier. Plan against an expert-tier rate multiplied by the hours you can actually get, and for most people that comes out as a useful supplement, well short of a replacement income.
Payment terms vary. Some platforms pay weekly, some monthly, some on project completion. Check before you commit time, and check whether calibration reading and disputed items are paid, since on some platforms they are not.
It is contractor income. Nothing withheld, tax owed, and usually no benefits or guaranteed hours.
Non-disclosure is standard. You will see unreleased model behaviour and confidential evaluation criteria, which limits what you can say publicly, including in a portfolio. That has a practical consequence worth planning around: unlike bug bounty hunting or open source documentation, this work builds you no public track record. Your reputation lives inside each platform's private quality scores and stays there. If you want evidence of expertise that follows you, it has to come from your existing professional credentials and from work done elsewhere, which is one more reason to treat this as income rather than as career building.
Here is the realistic framing. For a credentialed professional, this is among the best-paid flexible work available, at rates comparable to or above your day rate, with no client acquisition and no business overhead. It will never be a business, and it will never be stable, and within those limits it is unusually good.
If a quiet month brought you twenty hours and the next brought none, would your household budget notice? Your answer tells you whether this belongs in the bills column or the extras column.
Twenty hours at $150 in a good month is a real sum for evening work you are already qualified to do. If your specialist queue is busy enough, that could buy your mum's flight home to visit family, paid from reviewing model answers in your own specialism. Wait for approved work to pay before you book her seat.
Think about what extra evening income from your own expertise could cover. A month of childcare, a car repair you have been putting off, or the course you want to take next. Earning it by judging work in the field you already know feels very different from picking up shifts somewhere else.
Why the Labs Need Humans At All
The obvious objection deserves a direct answer, because it decides how durable this work is for you: if models are so capable, why can they not evaluate themselves?
They partly can, and the technique is widely used. A model can score another model's output against criteria, cheaply and at enormous scale. What it cannot do is notice that both models share the same mistaken belief.
That is the structural limit. A model's sense of what is correct comes from the same training distribution that produced the answer. Where that distribution holds an error, a gap or an outdated fact, the evaluator inherits it and confidently confirms the wrong answer. Scaling automated evaluation scales the blind spot along with the coverage.
Human expert evaluation exists to bring judgment from outside that distribution. A physician who trained on patients, a lawyer who has argued the point, an engineer who has debugged the failure knows things that are thin or wrong in written text. That is what the labs are buying from you, and it is why the credential, more than raw intelligence, is the qualification.
Three consequences follow, and they explain the shape of the market.
Coverage of hard cases matters more than volume. Labs have little use for a million easy evaluations; they need the specific items where automated judgment fails. That pushes spending toward experts and away from bulk annotation.
The value gathers where being wrong is expensive. Medicine, law, finance, safety-critical engineering. The same principle governs accessibility auditing, technical documentation and medical coding elsewhere on this site: pay tracks the cost of error.
The need moves as models improve, and stays. Better models fail on harder cases, which take deeper expertise to catch, which raises the expertise threshold. The entry tier shrinks and the expert tier holds.
Maximising What You Earn
With volume intermittent and rates so varied, a few things reliably move your number.
Lead with your credential everywhere. The licence, the doctorate, the years in a specialism. This decides your tier before anything else, and profiles that lead with generic skills get placed in generic work.
Name one narrow domain. "Physician" puts you in general medical evaluation. "Practising oncologist with clinical trial experience" puts you in a much smaller pool for work that pays more, and the labs are looking for exactly that detail.
Negotiate at the expert tier. Posted rates for scarce credentials are often an opening position. If you hold a qualification the platform is actively recruiting for, you have more leverage than you think, and the platform's difficulty finding you is your evidence.
Build a quality record early. Platforms audit, and they route their best-paid projects to reliable evaluators. Careful work in your first months compounds into access, the same way signal score works in bug bounty hunting.
Run three or four platforms. Project work arrives in waves, and one platform means idle weeks between commissions. This is the biggest single driver of your monthly income at a given rate.
Track your effective rate. Unpaid assessments, calibration reading, disputed items and gaps between projects all sit between your posted rate and what you really earn. A $150 posted rate on fifteen billable hours a month is a useful supplement, well short of a salary.
Say yes to the unpleasant categories only on purpose. Safety and red-team work often pays more because fewer people want it. That is a fair trade to make with your eyes open, and a poor one to drift into.
How would you describe your specialism in eight words, narrow enough that only a few hundred people could claim it? Put those words at the top of every profile.
Leading with a narrow credential is what moves you up a tier. At $75 to $200 an hour in a busy month, after setting money aside for tax, the extra could be the anniversary dinner where you order whatever you like, earned in a few evenings grading model outputs on cases you know cold.
Who This Suits
I will be direct here, because the qualifying condition is unusually specific.
It suits you if you are a credentialed professional who wants flexible income from expertise you already have. This is the central case. A doctor, lawyer, accountant, PhD or senior engineer can earn at or above their professional day rate, with no client acquisition, no invoicing and no business overhead.
It suits you if your profession is being squeezed by AI. If your field is being partly automated, the labs doing the automating pay well for the judgment the automation still lacks. That is the most direct way available to turn a threatened credential into income.
It suits you if you enjoy careful evaluative work. The job is reading, judging and justifying. If you found marking, peer review or code review satisfying, this will feel familiar.
It suits you if your availability is irregular. Work is project-based and mostly asynchronous, which fits around clinical shifts, court dates and teaching terms better than most side income.
It does not suit you if you lack a specialist credential or deep domain expertise. The entry tier at $12 to $25 exists, competes globally, and is not worth reorganising your week around.
It does not suit you if you need predictable monthly income. Volume is not guaranteed, and projects end without notice.
It does not suit you if safety work would distress you. Some categories involve deliberately drawing out harmful content or reviewing disturbing material, and you are better off knowing that in advance than discovering it mid-queue.
Rookie Mistakes
Starting at the entry tier when you qualify for the expert tier. People with medical or legal credentials routinely sign up for general annotation at $15 an hour because it was the tier they found first. Apply to the specialist platforms with your credential up front.
Underselling your credential. The professional licence, the doctorate, the decade in a specialism is the whole product. A profile that leads with generic skills buries the one thing that matters.
Claiming broad expertise. Depth beats breadth here, and assessments catch overreach.
Treating evaluation as opinion. The task is judgment against criteria, with your reasoning written down. Scoring without explaining why is low-quality work, and platforms track quality.
Rushing. Throughput matters and accuracy matters more, because inaccurate evaluation actively corrupts the training signal. Platforms audit, and a poor quality score costs you access to the better-paid projects.
Depending on one platform. Project-based work with no volume guarantee needs a backup.
Ignoring the confidentiality terms. They are real, they are enforced, and breaching them shuts you out of the whole category.
Red-Teaming Specifically
Everything so far has assumed you are grading answers someone else asked for. This part asks you to go looking for the failure yourself.
Red-teaming gets its own section because it is the least understood category, it often pays above standard evaluation, and it suits a different temperament.
The task is adversarial: you deliberately try to make a model produce output it should not, then document exactly how you did it and why it worked. You are paid to find the failure, the reverse of most jobs, and some people love that.
Your domain expertise makes it far more valuable. Anyone can try generic jailbreaks, and the labs have seen those. A clinician who knows which incorrect medical advice would actually cause harm, or a security engineer who knows which code pattern is exploitable, can find the failures that matter, well beyond the ones that merely break a policy. Your expertise tells you where to aim.
Creativity beats persistence. Repeating known techniques with small variations produces duplicates. The finding labs value is a category of failure nobody had considered, and it comes from thinking hard about the domain more than about the model.
Documentation is the deliverable, as everywhere else in this field. A successful elicitation with no clear account of the conditions, the reasoning and the severity is far less useful than one that explains why the model failed and what kind of prompt would reproduce it.
Severity assessment is where you add the most as a domain expert. Whether an output is embarrassing, misleading or genuinely dangerous is a judgment about the world, and the text alone cannot settle it. Non-experts misjudge this in both directions, again and again.
Two honest cautions. The work involves deliberately generating content you may find unpleasant, sometimes for long stretches, and the psychological load is real and rarely discussed. And there is an ethical side worth being deliberate about: you are improving systems by finding their failures, which is defensible work, and the techniques you develop have obvious dual use. Reputable programmes handle this with confidentiality terms and clear scope. If either aspect sits badly with you, choose standard evaluation instead, where the rates are somewhat lower and the work is calmer.
Be honest with yourself: would you sleep fine after an evening spent coaxing a model toward dangerous advice in your own field? If the answer is no, standard evaluation is the kinder choice for you.
Getting Paid What the Tier Is Worth
Here is a specific problem in this market: the people who qualify for the top tier are often the worst at claiming it, because professionals are trained to be modest about credentials and to treat rates as fixed. If that sounds like you, read this section twice.
Understand your scarcity. If a platform is actively recruiting practising oncologists, the number of practising oncologists willing to do evaluation work in their evenings is small. That is leverage, and you will only see it if you look for it.
Ask what tier you have been placed in and why. Placement is often automatic from your profile, and profiles often under-describe. A physician who wrote "healthcare professional" may be sitting two tiers below where their licence belongs.
Provide the verification up front. Licence numbers, degree certificates, publication records, employment history. Verification is what opens the tier to you, and platforms cannot place you on expertise they cannot confirm.
Re-apply after credential changes. Finishing a doctorate, gaining a specialty qualification or passing a professional exam should trigger a profile update and, often, a rate review.
Compare across platforms openly. Rates for the same credential vary between platforms, and knowing what another one offers is the simplest negotiating position you have.
Decline underpriced work in your specialism. Accepting $30 an hour for expert medical evaluation undervalues your time and sets the rate the platform will offer you next time. It feels uncomfortable, and it is how the market rate for scarce expertise holds.
What hourly figure would you quote a hospital or firm for a second opinion in your field? Hold that figure in mind the next time a platform offers you less.
Gotchas Worth Knowing
The work can be genuinely unpleasant. Red-teaming means deliberately drawing out harmful content, and safety evaluation can mean reviewing distressing material. Platforms should tell you this; ask if they do not.
Rates are negotiable at the expert tier and rarely at the entry tier. If you hold a scarce credential, the posted rate is often an opening position.
Your professional obligations follow you. If you are a clinician or lawyer evaluating model outputs in your field, you still carry your professional body's expectations on competence and confidentiality. The rules still apply to you here, even though this work sits outside your practice.
Assessments take unpaid time. Credential verification and domain testing can eat several hours before any paid work. Factor that in when you apply to several platforms.
The task can change mid-project. Evaluation criteria get revised as the lab learns what it needs, sometimes undoing earlier guidance. Ask how re-work is handled.
Conflicts of interest are possible. Evaluating outputs in a field where you have commercial relationships, or where your employer competes with the lab, deserves thought before you accept.
What Good Evaluation Actually Looks Like
Speed has little to do with the gap between a $30 evaluator and a $150 one. The gap is the quality of the reasoning attached to each judgment, and you can learn that in a way you cannot learn a credential.
Judge against the criteria. Labs supply rubrics. If you substitute personal taste, you produce inconsistent signal, which is worse than no signal because nobody can correct for it. If the rubric seems wrong, say so separately, and keep applying it while you wait for an answer.
Separate the dimensions. An answer can be factually correct and badly structured, or well-written and dangerously wrong. Folding those into one impression loses the information the lab needs. Score each dimension the rubric asks for, one at a time.
Explain in terms of the error. "This is bad" is worthless. "The recommended dose exceeds the maximum for this weight band, which would be harmful rather than merely suboptimal" is training signal. Your justification is the deliverable.
Distinguish severity. A stylistic weakness and a claim that would cause harm are different failures, and flattening them teaches the model the wrong lesson. Labs care a great deal about the difference.
Flag the ambiguous prompt. Often a response cannot be rated because the question was underspecified, or the right answer depends on a jurisdiction, patient history or context the prompt left out. Saying so helps more than picking at random, and it points to a defect in the evaluation set itself.
Be consistent with yourself. Your judgments are compared across items and against other experts. If you score the same failure differently on Tuesday and Thursday, you lower the value of all your work, which is why doing this tired is a false economy.
Note what you are unsure about. Marking your own confidence is genuinely valuable, and most evaluators skip it. A confident wrong judgment from a credentialed expert is the most damaging thing this process can produce.
When you last explained a mistake to a trainee or a client, did you name the exact error and its consequence, or just say it was wrong? The first habit is the one that earns $150.
Against the Alternatives
This is worth placing in context, because if you qualify, you have other uses for the same hours.
Against taking extra clinical, legal or professional shifts. Usually comparable money at the top tier, with no commute, no scheduling commitment and no client contact. The trade is that volume is not guaranteed, so it sits alongside your professional work and never stands in for it.
Against consulting in your field. Consulting pays more per hour at the top end and demands business development, proposals, invoicing and chasing payment. This has none of that and a lower ceiling. If you want income without running a practice, that is a good trade.
Against the other side income on this site. Almost everything else here asks you to build something: an audience, a client base, a listing, a track record. This asks for a credential you already hold and an application form. That is genuinely unusual, and it is the strongest argument for trying it.
Against teaching or examining. The closest familiar comparison, and generally better paid here, with more flexible scheduling and less administration.
The honest summary: for a credentialed professional this is the lowest-friction well-paid work available, and it will never grow into a business. Treat it as an efficient way to turn expertise you already have into flexible income, and keep your main work as your main work.
Behind the Scenes: An Evaluation Shift
The experience feels much closer to marking exams than to anything technological.
You open a queue. Each item is a prompt and one or two model responses, in your specialist field. Your job is to score them against criteria the lab has defined, and to explain why.
The first few are easy. One response is plainly better, or plainly wrong, and you say so.
Then the interesting ones start. Two responses that are both defensible, differing in emphasis. A response that is factually correct and clinically inappropriate. An answer that would be right in one jurisdiction and wrong in another, where the prompt did not say which. These are where your expertise is really being used, and they take five times as long as the easy ones.
Then the one that stops you: an answer that is confidently, fluently wrong in a way a non-expert would never catch. A dosage that is plausible and dangerous. A citation to a case that does not exist. This is the item the whole exercise exists to find, and finding it is why the rate is what it is.
Writing the justification takes you longer than making the judgment. The lab needs to know why, specifically, in terms it can turn into training signal.
Now and then your work is audited against other experts. Disagreements happen and get discussed, and those discussions are often the most intellectually interesting part of the job.
The rhythm suits you if you enjoy careful judgment in your own field and do not need novelty. Many people find it more engaging than they expected, because unlike most side income it uses the expensive thing they spent years acquiring.
The Thirty-Day Reality
You have the shape of the work, the rates and the traps. Here is how your first month actually runs.
I am setting expectations carefully, because the gap between applying and earning catches people out.
Week one is applications and unpaid time. Several platforms, credential verification on each, and domain assessments that can take a few hours apiece. None of it pays, and doing it across four platforms instead of one is the single decision that most shapes your income three months later.
Week two is waiting. Verification takes time, especially where licences are checked against professional registers. At the expert tier, people review your assessment by hand, so it takes longer to come back.
Weeks three and four bring the first project, or they do not. Work is commissioned as labs need it, so the timing depends on what they happen to need. A qualified oncologist may wait six weeks and then get more work than they can take.
The first project is a test in both directions. The platform is checking your consistency and reasoning against other experts, and you are finding out whether you can live with the work. Do it carefully, and let speed come later.
Volume, if it comes, arrives unevenly. A realistic first quarter is a few tens of hours, bunched in bursts, with quiet weeks between. That is simply the shape of the work, and it says nothing about how you are doing.
The planning point is simple: treat your first month as an investment of unpaid hours in access, apply widely enough that a quiet spell on one platform overlaps with work on another, and keep your main income fully intact the whole time.
If the first project took six weeks to arrive, would you still be glad you applied? Decide that now, so the waiting does not talk you out of it.
Access is most of what the first month pays you, so your main income keeps the household steady. If work arrives from two or three platforms and reaches even the low end, around 500 a month, it can go straight into a savings account the family never touches, so a quiet month between batches never reaches the rent.
Applying costs you an evening. Waiting costs you the projects that go to whoever applied last month. Your qualification is already sitting there, earning nothing outside your day job. Tonight, sign up on one platform, finish the profile honestly and start the screening task before bed.
Where This Goes Next
These are arguments about where the work is heading, in a field young enough that confident forecasts are unwise. The regulatory drivers are the firmest of them.
The entry tier keeps shrinking, and the expert tier holds. Simple labelling is being automated and competed down globally. Expert judgment in medicine, law and advanced technical fields has a supply constraint that will not ease, because the supply is people who finished long training.
Specialisation gets narrower and better paid. As models get better at general reasoning, the remaining errors gather in edge cases that need deep specific knowledge. That pushes demand toward narrower expertise and away from broad competence.
Verification becomes a standing function. As organisations deploy models in regulated and high-stakes settings, someone has to check the output all the time. That is a recurring role, and over time it will look more like employment than a research contract.
Rates for scarce credentials may rise before they fall. Labs competing for the same small pool of qualified evaluators creates a bidding situation. It will not last forever, and right now it favours you.
Evaluation work moves inside companies. Today most of this demand comes from labs training frontier models. As ordinary organisations deploy models in regulated and high-stakes settings, they inherit the same problem at a smaller scale: someone qualified has to confirm the output is right. That demand will bypass AI training platforms. It will show up as consulting, fractional roles and employment inside hospitals, firms and insurers, a steadier market than project work for labs, and the likely next step for you if you build a reputation here.
Provenance and accountability rules strengthen your position. Where regulation requires an organisation to show human oversight of AI-assisted decisions, a documented expert review becomes a compliance record as well as a quality check. That turns occasional evaluation into a recurring obligation, and obligations get funded more reliably than improvements.
The way in stays gated by credentials. No course will stand in for the qualification, which is unusual on this site and exactly why it is worth telling professionals who already hold one that this market exists. If that is you, the line you wrote at the top of this page is where your application starts.
The labs are building long relationships with reviewers they trust. Every month you hold back, someone with your exact qualification earns that trust first and gets the next project. The demand is real now, and the best queues go to people who show up while the labs are still recruiting.