Task Development Engineer at METR
METR is hiring a Task Development Engineer in Berkeley, CA, US. Remote. Pay: USD 261k-385k/yr.
About METR
METR is a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations. It focuses on threats related to AI R&D automation and misalignment.
Task Development Engineer job description
About METR
We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
About the role
Time Horizons is a central tool the world uses to understand AI progress. Our methodology has been included in system cards, called an "obsession" by the NYT, has wide reach online, and is used by governments to inform national policy. It is essential to our broader risk assessment work to have good capability evaluations.
Task Development Engineers contribute to METR’s expanding ambition of our evaluations with high quality tasks supporting the Time Horizons methodology. We expect our results to be seen by policymakers, frontier labs, national security stakeholders, and other key decisionmakers influencing society’s response to AI progress.
What this role looks like
(Primarily, and most importantly) Developing difficult, novel tasks for models. You will build well-scoped tasks that remain challenging as model time horizons grow, potentially to hundreds of hours.
Quality assurance for existing tasks. Once a task has been developed, you will verify that it's actually solvable as specified, and that the model is given (only) the information it needs.
Baselining and scoring tasks. Where helpful, you may be asked to baseline tasks within your domain of expertise, and/or score task completions from AIs or human baseliners.
Improving task development infrastructure. We're always improving our processes. Strong candidates will notice when existing workflows are inefficient or produce low-quality output, and take responsibility for improving them.
Skills we're looking for
Software engineering: You have several years of experience working on complex projects and codebases.
Evaluations: You have experience building hard (ideally agent-based) AI evaluations (e.g. RE-Bench, HCAST, SWE-bench Verified, Cybench, GPQA), ideally using the Inspect framework.
High attention to detail: You read closely, spot misspecifications and ambiguity, and pay attention to fiddly minutiae.
(Nice to have) Familiarity with METR infrastructure: Prior experience with Hawk, and familiarity with the methodology behind our Time Horizons work, is a plus.
- The office: Catered lunch and dinner daily; in-office gym and shower
- Relocation support: Stipend for moving to the Bay Area
- Time-off and leave: Unlimited PTO and 21-week parental leave for new parents
- Commuter benefit: Monthly transit/parking stipend and an annual Uber budget
- Professional development benefit: for training, courses, conferences, and AI safety education
- Mental health benefit: for therapy, medication, and other mental health expenses
- Wellness benefit: for gym memberships and other wellness expenses
- Work equipment benefit: for home office and workstation equipment expenses


