HiringDomain expertiseRemoteFlexible
AI Model Evaluator
Benchmark Collective · Remote (US) · Contract · Remote
$25 – $40 / hr
Why this role holds up next to AI
Deciding whether one model is genuinely better than another needs domain sense and honesty — the exact things you can't automate.
What the work is
Test AI models on real tasks and score how well they do. Bring expertise from a field you already know (law, medicine, finance, writing).
Day to day, you'd
- Run models through structured evaluation tasks
- Judge correctness and usefulness in your domain
- Write up findings that engineers can act on
You're more ready than you think
You already do this if you have deep knowledge in any professional field.
Apply for this role
Create your free account to apply — it takes under a minute.