AI evaluation and correction
AI companies need people to judge whether a model's output is good, rank competing answers, and correct mistakes. This is the fastest-growing category of paid human work anywhere.
What it involves
- Comparing two AI answers and saying which is better, and why.
- Marking output that is wrong, unsafe or unhelpful.
- Correcting a transcription or a summary the model got wrong.
What gets rejected
Stated up front, because a rejection reason you only discover afterwards is not a standard, it is a trap.
- Rankings with no stated reasoning.
- Using an AI to do the evaluation. It is detectable and it destroys the data.
- Failing the calibration items seeded through the batch.
What it pays
Among the best-paid work on the platform, and it rewards judgement and subject knowledge rather than speed. Specialist domains — medicine, law, a specific language — pay more because fewer people can do them well.
We do not publish an hourly figure for this or any other category. The honest answer depends on your country, your languages and what clients are asking for this week, and a number on a marketing page would be the best case dressed up as the typical one.