What you'll do
- Write and iterate the prompts that run Aria and our client agents — qualification flows, multilingual replies, tool-use instructions.
- Build evaluation sets out of real conversation logs, then use them to decide whether a prompt change shipped or got reverted.
- Read transcripts. A lot of them. Find the ten conversations where the agent got it wrong and work out what they have in common.
- Write the documentation for how each agent is supposed to behave, so the next person changing a prompt knows what they might break.
- Get on a call with engineers when a prompt problem turns out to be a retrieval problem, and learn the difference.
This role is fully remote, anywhere in India. It's entry-level and remote at the same time, which means we'll pair with you on calls often in the first months — but you have to ask. Nobody will spot you looking lost from across a room.
What we're looking for
- 0-2 years experience. This is an entry-level role and we will teach you the craft — we're hiring for judgement and writing, not a CV.
- You write clearly in English. Prompting is writing instructions for something that takes them literally, and most people are worse at that than they think.
- Evidence you've actually used LLMs seriously — a project, a tool you built, a workflow you automated for yourself. Show us the thing.
- Patience for detail. The job is often a one-word change and a re-run of fifty test cases.
- You're comfortable being told your prompt was worse than the old one, because the eval said so.
Nice to have
- Any Python — enough to run a script over a batch of test cases.
- A second language used well, especially Hindi or Bengali; several of our agents answer in more than English.
- Having read enough model documentation to have opinions about where different models break.