Task Development Engineer
METR (Model Evaluation and Threat Research)RemoteGlobal catastrophic risksAI safetyEngineeringResearch
This role is on our board because we think it is an unusually effective way to spend a career. How we reason about it
Free to leave the Netherlands? Look at 80,000 Hours, Probably Good and EA Opportunities first — the strongest opportunities are there. Why we say this
This board is still in beta. The reason below is our judgement — tell us if you think we have it wrong. Give feedback
Why this is on the board
METR's measurements — how long an AI model can work on a task on its own — are taken up in the system cards of AI labs and used by governments for policy. You build the tasks those figures rest on: one badly specified task quietly distorts the picture regulators have of AI progress.
About the role
We are a nonprofit research organisation that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment.
Requirements
| Language | English is enough |
|---|---|
| Work authorisation | Unclear |
| Screening | Not mentioned |
| Where you work | Remote · Remote |
| Level | Mid level |
| Salary | Not stated |
| Posted | 1 August 2026 |
| Closes | No closing date given |
This takes you to the employer’s own site, where you apply.
More about METR (Model Evaluation and Threat Research)Why this problemGlossary
What do you make of this?
The board is still in beta, so this is exactly the moment to say what you think. A vacancy we have missed, a judgement you do not share, something that does not work — all of it helps.
Curious?
If you want to know how we arrive at this list, we run a short intro course and a newsletter. Neither costs anything.