The AI labs have a shortage, and it is not compute. It is people who can tell whether a model's answer is correct in a field where being wrong matters.
Mercor, one of the platforms matching experts to labs, publishes its own rate data. The national average for AI training work is $31 an hour as of early 2026, with rates ranging from $12 to over $200 depending on experience, specialisation and platform.
That range is the entire story, because it is not a range within one job. It is three different jobs.
Mercor's own explanation of why this tier pays $75 to $200 and up is worth quoting directly, because it is the whole thesis: these roles "require the ability to judge, analyze, evaluate, and improve upon outputs that most workers cannot assess. As a result, there's a genuine supply constraint, which leads to higher pay rates."
A model can produce a fluent, confident, well-structured answer about drug interactions, contract enforceability, tax treatment or structural load. Producing that answer is now cheap. Determining whether it is correct is not, and it cannot be done by another model, because if the model knew the answer was wrong it would not have produced it.
That is the supply constraint. The labs need people who can look at an output in a specialist field and say, with authority, that it is wrong and specifically why. The set of people who can do that for medicine is roughly the set of people who trained in medicine.
Four kinds of work follow from this.
That recurrence is worth pausing on, because it is the clearest signal in this whole collection of guides. Across unrelated professions, asked separately what remains valuable as tools absorb the routine work, the answer keeps arriving in the same shape: somebody qualified has to confirm the output is correct, and that somebody has to be qualified in the domain rather than in the tool. A paralegal verifying that generated citations exist and remain good law, a coder confirming an assigned code matches the documentation, a designer checking a generated asset survives a print spec, and an oncologist catching a plausible dosage error are performing one activity in four settings. This guide is about selling that activity directly to the organisations building the systems, at the rates they pay for it, rather than waiting for it to arrive inside an existing job.
The unusual feature of this work is that most people who qualify for the top tier are not looking for it, because they do not know it exists.
There is a direct connection to the profession guides elsewhere on this site. The paralegals facing 0.2 percent employment growth, the designers at 2.1 percent, the professionals watching AI absorb parts of their work: the same expertise that is being partly displaced is what the labs are paying $75 to $200 an hour for. That is not a consolation. It is the most direct route available from a squeezed credential to a well-paid use of it.
The process is lighter than the rates suggest, and it is largely about proving the credential.
The rates are real and the caveats matter.
Beware the earnings screenshots. Mercor's own guidance opens by acknowledging the claims circulating about people making seven or eight thousand dollars a month at this, and the honest reconciliation is arithmetic rather than mystery: those figures require the top tier and close to full-time hours, sustained, which very few people achieve because volume is not guaranteed at any tier. An expert-tier rate multiplied by the hours you can actually get is the number to plan against, and for most people that is a useful supplement rather than a replacement income.
Payment terms vary. Some platforms pay weekly, some monthly, some on project completion. Check before committing time, and check whether time spent on calibration reading and disputed items is paid, since on some platforms it is not.
It is contractor income. Nothing withheld, tax owed, and typically no benefits or guaranteed hours.
Non-disclosure is standard. You will see unreleased model behaviour and confidential evaluation criteria, which limits what you can say publicly, including in a portfolio. That has a practical consequence worth planning around: unlike bug bounty hunting or open source documentation, this work builds no public track record. Your reputation lives inside each platform's private quality scores and does not travel. If you want evidence of expertise that follows you, it has to come from your existing professional credentials and from work done elsewhere, which is another reason to treat this as income rather than as career building.
The realistic framing: for a credentialed professional this is among the best-paid flexible work available, at rates comparable to or above their day rate, with no client acquisition and no business overhead. It is not a business, it is not stable, and it is unusually good at what it is.
Why the Labs Need Humans At All
The obvious objection is worth answering directly, because it determines how durable this work is: if models are so capable, why can they not evaluate themselves?
They partly can, and the technique is widely used. A model can score another model's output against criteria, cheaply and at enormous scale. What it cannot do is notice that both models share the same mistaken belief.
That is the structural limit. A model's judgment of correctness is drawn from the same training distribution that produced the answer. Where that distribution contains an error, a gap or an outdated fact, the evaluator inherits it and confidently confirms the wrong answer. Scaling automated evaluation scales the blind spot along with the coverage.
Human expert evaluation exists to supply judgment from outside that distribution. A physician who trained on patients, a lawyer who has argued the point, an engineer who has debugged the failure knows things that are underrepresented or wrong in text. That is the thing being bought, and it is why the credential rather than the intelligence is the qualification.
Three consequences follow, and they explain the shape of the market.
Volume is not the objective, coverage of hard cases is. Labs do not need a million easy evaluations, they need the specific items where automated judgment fails. That pushes spend toward experts and away from bulk annotation.
The value concentrates where being wrong is expensive. Medicine, law, finance, safety-critical engineering. This is the identical principle that governs accessibility auditing, technical documentation and medical coding elsewhere on this site: pay tracks the cost of error.
The need does not disappear as models improve. It moves. Better models fail on harder cases, which require deeper expertise to catch, which raises rather than lowers the expertise threshold. The entry tier shrinks and the expert tier does not.
Maximising What You Earn
Given intermittent volume and wide rate variation, a few things reliably move the number.
Lead with the credential everywhere. The licence, the doctorate, the years in a specialism. This determines your tier before anything else, and profiles that lead with generic skills get placed in generic work.
Name one narrow domain rather than several broad ones. "Physician" places you in general medical evaluation. "Practising oncologist with clinical trial experience" places you in a much smaller pool for work that pays more, and the labs are looking for exactly that granularity.
Negotiate at the expert tier. Posted rates for scarce credentials are frequently an opening position. Someone holding a qualification the platform is actively recruiting for has more leverage than they assume, and the platform's difficulty finding you is your evidence.
Build a quality record early. Platforms audit and route their best-paid projects to reliable evaluators. Careful work in the first months compounds into access, in the same way signal score works in bug bounty hunting.
Run three or four platforms. Project work arrives in waves, and one platform means idle weeks between commissions. This is the single biggest determinant of monthly income at a given rate.
Track your effective rate, not your posted rate. Unpaid assessments, calibration reading, disputed items and gaps between projects all sit between the two. A $150 posted rate on fifteen billable hours a month is a useful supplement rather than a salary.
Say yes to the unpleasant categories only deliberately. Safety and red-team work frequently pays more because fewer people want it. That is a legitimate trade to make consciously, and a poor one to drift into.
Who This Suits
Direct, because the qualifying condition is unusually specific.
It suits credentialed professionals who want flexible income from expertise they already have. This is the central case. A doctor, lawyer, accountant, PhD or senior engineer can earn at or above their professional day rate, with no client acquisition, no invoicing and no business overhead.
It suits people in AI-squeezed professions specifically. If your field is being partly automated, the labs doing the automating are paying well for the judgment that automation still lacks. That is the most direct available conversion of a threatened credential into income.
It suits people who enjoy careful evaluative work. The job is reading, judging and justifying. Anyone who found marking, peer review or code review satisfying will find this familiar.
It suits people with irregular availability. Work is project-based and largely asynchronous, which fits around clinical shifts, court dates and teaching terms better than most side income.
It does not suit people without a specialist credential or deep domain expertise. The entry tier at $12 to $25 exists, competes globally, and is not worth reorganising your week around.
It does not suit anyone needing predictable monthly income. Volume is not guaranteed and projects end without notice.
It does not suit anyone who will be distressed by safety work. Some categories involve deliberately eliciting harmful content or reviewing disturbing material, and knowing that in advance is better than discovering it.
Rookie Mistakes
Starting at the entry tier when you qualify for the expert tier. People with medical or legal credentials routinely sign up for general annotation work at $15 an hour because it is the tier they found first. Apply to the specialist platforms with your credential foregrounded.
Underselling the credential. The professional licence, the doctorate, the decade in a specialism is the entire product. A profile that leads with generic skills buries the only thing that matters.
Claiming broad expertise. Depth beats breadth here, and assessments catch overreach.
Treating evaluation as opinion. The task is judgment against criteria, with reasoning documented. Scoring without explaining why is low-quality work and platforms track quality.
Rushing. Throughput matters and accuracy matters more, because inaccurate evaluation actively corrupts the training signal rather than merely failing to improve it. Platforms audit, and a poor quality score costs access to the better-paid projects.
Depending on one platform. Project-based work with no volume guarantee needs redundancy.
Ignoring the confidentiality terms. They are real, they are enforced, and breaching them ends access to the whole category rather than one project.
Red-Teaming Specifically
Red-teaming deserves separate treatment because it is the least understood category, it frequently pays above standard evaluation, and it suits a different temperament.
The task is adversarial: deliberately attempt to make a model produce output it should not, then document exactly how you did it and why it worked. You are being paid to find the failure, which is the opposite of most work and appeals strongly to some people.
Domain expertise makes it far more valuable. Anyone can attempt generic jailbreaks, and the labs have seen those. A clinician who knows which incorrect medical advice would actually cause harm, or a security engineer who knows which code pattern is exploitable, can identify failures that matter rather than failures that merely violate a policy. The expertise tells you where to aim.
Creativity beats persistence. Repeating known techniques with variations produces duplicates. The valued finding is a category of failure nobody had considered, which comes from thinking about the domain rather than about the model.
Documentation is the deliverable, as everywhere else in this field. A successful elicitation with no clear account of the conditions, the reasoning and the severity is much less useful than one that explains why the model failed and what class of prompt would reproduce it.
Severity assessment is where domain experts add the most. Whether an output is embarrassing, misleading or genuinely dangerous is a judgment about the world, not about the text. Non-experts systematically misjudge this in both directions.
Two honest cautions. The work involves deliberately generating content you may find unpleasant, sometimes for extended periods, and the psychological load is real and under-discussed. And there is an ethical dimension worth being deliberate about: you are improving systems by finding their failures, which is defensible work, and the techniques you develop have obvious dual use. Reputable programmes handle this with confidentiality terms and clear scope. Anyone uncomfortable with either aspect should choose standard evaluation instead, where the rates are somewhat lower and the work is calmer.
Getting Paid What the Tier Is Worth
A specific problem with this market: the people who qualify for the top tier are frequently the worst at claiming it, because professionals are trained to be modest about credentials and to treat rates as fixed.
Understand your scarcity. If a platform is actively recruiting practising oncologists, the number of practising oncologists willing to do evaluation work in their evenings is small. That is leverage, and it is invisible unless you look for it.
Ask what tier you have been placed in and why. Placement is often automatic from your profile, and profiles frequently under-describe. A physician who wrote "healthcare professional" may be sitting two tiers below where their licence entitles them.
Provide the verification proactively. Licence numbers, degree certificates, publication records, employment history. Verification is what unlocks the tier, and platforms cannot place you on expertise they cannot confirm.
Re-apply after credential changes. Completing a doctorate, gaining a specialty qualification or passing a professional exam should trigger a profile update and, frequently, a rate review.
Compare across platforms openly. Rates for the same credential vary between platforms, and knowing what another is offering is the most straightforward negotiating position available.
Decline underpriced work in your specialism. Accepting $30 an hour for expert medical evaluation both undervalues your time and establishes the rate the platform will offer next time. This is uncomfortable and it is how the market rate for scarce expertise is maintained.
Gotchas Worth Knowing
The work can be genuinely unpleasant. Red-teaming means deliberately eliciting harmful content, and safety evaluation can involve reviewing distressing material. Platforms should disclose this; ask if they do not.
Rates are negotiable at the expert tier and rarely at the entry tier. If you hold a scarce credential, the posted rate is frequently an opening position.
Your professional obligations follow you. A clinician or lawyer evaluating model outputs in their field still carries their professional body's expectations around competence and confidentiality. Do not assume that because it is not practice, the rules do not apply.
Assessments take unpaid time. Credential verification and domain testing can consume several hours before any paid work. Factor that in when applying to multiple platforms.
The task can change mid-project. Evaluation criteria get revised as the lab learns what it needs, sometimes invalidating earlier guidance. Ask how re-work is handled.
Conflicts of interest are possible. Evaluating outputs in a field where you have commercial relationships, or where your employer is a competitor to the lab, deserves thought before you accept.
What Good Evaluation Actually Looks Like
The difference between a $30 evaluator and a $150 one is not speed. It is the quality of the reasoning attached to each judgment, and that is learnable in a way the credential is not.
Judge against the criteria, not against your preference. Labs supply rubrics. An evaluator who substitutes personal taste produces inconsistent signal, which is worse than no signal because it cannot be corrected for. If the rubric seems wrong, say so separately rather than quietly ignoring it.
Separate the dimensions. An answer can be factually correct and badly structured, or well-written and dangerously wrong. Collapsing those into a single impression loses the information the lab needs. Score each dimension the rubric asks for, distinctly.
Explain in terms of the error, not the feeling. "This is bad" is worthless. "The recommended dose exceeds the maximum for this weight band, which would be harmful rather than merely suboptimal" is training signal. The justification is the deliverable.
Distinguish severity. A stylistic weakness and a claim that would cause harm are not the same failure, and flattening them teaches the model the wrong lesson. Labs care disproportionately about the difference.
Flag the ambiguous prompt. Frequently the response is unratable because the question was underspecified, or the correct answer depends on a jurisdiction, patient history or context the prompt omitted. Saying so is more useful than picking arbitrarily, and it identifies a defect in the evaluation set itself.
Be consistent with yourself. Your judgments are compared across items and against other experts. An evaluator who scores the same failure differently on Tuesday and Thursday reduces the value of all their work, which is why doing this while tired is a false economy.
Note what you are uncertain about. Marking your own confidence is genuinely valuable and most evaluators omit it. A confident wrong judgment from a credentialed expert is the most damaging output this process can produce.
Against the Alternatives
Worth situating, since the people who qualify for this have other options for the same hours.
Against taking extra clinical, legal or professional shifts. Usually comparable money at the top tier, with no commute, no scheduling commitment and no client interaction. The trade is that volume is not guaranteed, so it complements rather than replaces professional work.
Against consulting in your field. Consulting pays more per hour at the top end and requires business development, proposals, invoicing and chasing payment. This has none of that and a lower ceiling. For someone who wants income without running a practice, that is a good trade.
Against the other side income on this site. Almost everything else here requires building something: an audience, a client base, a listing, a track record. This requires a credential you already have and an application form. That is a genuinely unusual property and it is the strongest argument for trying it.
Against teaching or examining. The closest familiar comparison, and generally better paid here, with more flexible scheduling and less administration.
The honest summary: for a credentialed professional this is the lowest-friction well-paid work available, and it is not a business and will not become one. Treat it as an efficient way to convert existing expertise into flexible income, and keep the primary thing primary.
Behind the Scenes: An Evaluation Shift
The experience is closer to marking exams than to anything technological.
You open a queue. Each item is a prompt and one or two model responses, in your specialist field. Your job is to score them against criteria the lab has defined, and to explain why.
The first few are easy. One response is obviously better, or obviously wrong, and you say so.
Then the interesting ones start. Two responses that are both defensible, differing in emphasis. A response that is factually correct and clinically inappropriate. An answer that would be right in one jurisdiction and wrong in another, where the prompt did not specify. These are where your expertise is actually being used, and they take five times as long as the easy ones.
Then the one that stops you: an answer that is confidently, fluently wrong in a way that a non-expert would never catch. A dosage that is plausible and dangerous. A citation to a case that does not exist. This is the item the entire exercise exists to find, and finding it is why the rate is what it is.
Writing the justification takes longer than the judgment. The lab needs to know why, specifically, in terms that can be turned into training signal.
Periodically your work is audited against other experts. Disagreement happens and is discussed, and the discussion is often the most intellectually interesting part of the job.
The rhythm suits people who enjoy careful judgment in their own field and do not need novelty. Many find it more engaging than expected, because unlike most side income it uses the expensive thing they spent years acquiring.
The Thirty-Day Reality
Setting expectations properly, because the gap between applying and earning catches people out.
Week one is applications and unpaid time. Several platforms, credential verification on each, and domain assessments that can take a few hours apiece. None of this pays, and doing it across four platforms rather than one is the single decision that most affects your income three months later.
Week two is waiting. Verification takes time, particularly where licences are checked against professional registers. Assessments are reviewed rather than auto-scored at the expert tier.
Weeks three and four bring the first project, or they do not. Work is commissioned rather than continuously available, so the timing depends on what the labs happen to need. A qualified oncologist may wait six weeks and then receive more work than they can take.
The first project is a test in both directions. The platform is assessing your consistency and reasoning quality against other experts, and you are finding out whether you can tolerate the work. Do it carefully rather than quickly.
Volume, if it comes, arrives unevenly. The realistic first quarter is a few tens of hours, concentrated in bursts, with quiet weeks between. That is the shape of the work rather than a sign of failure.
The planning implication is straightforward: treat the first month as an investment of unpaid hours in access, apply broadly enough that quiet periods on one platform overlap with work on another, and keep your primary income entirely intact throughout.
Where This Goes Next
These are arguments about where the work is heading, in a field young enough that confident forecasts are unwise. The regulatory drivers are the firmest of them.
The entry tier keeps compressing and the expert tier holds. Simple labelling is being automated and competed down globally. Expert judgment in medicine, law and advanced technical fields has a supply constraint that does not resolve, because the supply is people who completed long training.
Specialisation gets narrower and better paid. As models improve at general reasoning, the remaining errors concentrate in edge cases requiring deep specific knowledge. That pushes demand toward narrower expertise rather than broad competence.
Verification becomes a standing function rather than a project. As organisations deploy models in regulated and high-stakes contexts, someone has to check the output continuously. That is a recurring role rather than a research contract, and it will look more like employment over time.
Rates for scarce credentials may rise before they fall. Competition among labs for the same small pool of qualified evaluators is a bidding situation. That is not permanent, and it is currently favourable.
Evaluation work migrates inside companies. Right now most of this demand comes from labs training frontier models. As ordinary organisations deploy models in regulated and high-stakes settings, they acquire the same problem at smaller scale: someone qualified has to confirm the output is right. That demand will not route through AI training platforms. It will appear as consulting, fractional roles and employment inside hospitals, firms and insurers, which is a more stable market than project work for labs and is the likely destination for anyone who builds a reputation here.
Provenance and accountability requirements strengthen the position. Where regulation obliges an organisation to demonstrate human oversight of AI-assisted decisions, a documented expert review stops being a quality measure and becomes a compliance artefact. That converts occasional evaluation into a recurring obligation, and obligations are funded more reliably than improvements.
The route in stays credential-gated. No course will substitute for the qualification, which is unusual on this site and is precisely why it is worth telling professionals who already hold one that this market exists.