AI Benchmark Quality Reviewer
Braintrust
Job description
About the role
Help review the quality and fairness of challenging tasks used to evaluate AI systems. You will inspect task instructions, model execution traces and grading behavior, then explain whether a result reflects genuine model performance or an issue with the task, grader or environment.
Key responsibilities
- Check task instructions, source materials, reference solutions and evaluation criteria for consistency and completeness.
- Review model execution traces, tool calls and deliverables to assess whether successes and failures are justified.
- Identify brittle grading checks, unsupported criteria and valid alternative solutions that may have been marked incorrect.
- Investigate discrepancies and distinguish model limitations from task, grader, tool or environment issues, then write concise, evidence‑backed findings.
Required profile
- At least five years of relevant technical or analytical experience.
- Strong written English, analytical judgment and attention to detail.
Required skills
- Python
- SQL
- Shell scripts
What we offer
- $23 per hour compensation.
- Remote contractor assignment for eight weeks, 40 hours per week with eight‑hour daily overlap with Pacific Time.
Questions fréquentes
Why are you reporting this job?
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 5 days ago
Expires 1 month from now
23 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Braintrust