All roles
Agents
Applied AI engineer, agents
The part that decides when an agent speaks and which tool it reaches for.
- Where
- Remote, Europe (±2h CET)
- Type
- Full-time
- Band
- €80–110k + equity
An agent that answers everything is noise and an agent that answers nothing is furniture. The distance between those two is prompts, tool definitions and a loop, and that is the job.
You would also own the router: read the message, pick the cheapest model that will still get it right, and be honest in the evals about when that fails.
What you would be doing
- Own the prompt and the tool-calling loop, and the evals that keep them honest
- Make the router pick the cheapest model that still gets it right
- Turn “the agent was annoying in this thread” into a reproducible test
What we look for
- You have shipped an LLM feature that survived real users, not a demo
- You can read a transcript and say which instruction caused that sentence
- Suspicious of benchmarks, fond of examples
What to send
A transcript of a model doing something wrong and your write-up of why it happened.
No cover letter. One screen of writing is plenty, and we answer either way within a week.
Apply for this roleHow it goes from there
- 1A replyWe read every letter ourselves and answer, yes or no, within a week.Week 1
- 2A callThirty minutes with whoever you would work with. No panel, no riddles.Week 1
- 3A paid half-dayA real problem from our backlog, paid at your day rate. You keep the work either way.Week 2
- 4The offerA chat with the other two, then numbers in writing the same day.Week 2