What you'll do
- Build the AI layer in Aria and client systems: retrieval, tool-calling agents, prompt pipelines, and the services around them in Node.js, TypeScript and Python.
- Own features end to end — schema, API, model calls, deploy, and the 2am failure. We are five people; nobody hands your work to a separate ops team.
- Write and maintain evaluation sets from real conversations, and re-run them on every change. If a prompt edit drops accuracy, we want the number before the client finds it.
- Make the AI parts cheap and fast enough to run in production — caching, model selection, token budgets, fallbacks for when a provider degrades.
- Work directly with clients on scope. You'll be on the call where we say a feature isn't worth building.
This role is fully remote, anywhere in India. We are five people spread across a handful of cities, so the work happens over calls and written docs — you'll need to be good at writing and comfortable saying when you're stuck without anyone reading it off your face.
What we're looking for
- 3-5 years shipping backend software that real users depend on, with at least one system using LLMs in production — not a notebook demo.
- Strong in TypeScript/Node or Python, and comfortable in the other.
- You can debug a bad answer down to its cause: wrong chunk retrieved, wrong tool called, wrong context window, wrong prompt.
- You treat model output as something to be measured, not trusted. You've argued with someone about whether an eval was actually measuring the thing.
- Comfortable with Postgres, an API you designed yourself, and reading someone else's code without rewriting it first.
Nice to have
- Experience with agentic patterns — multi-step tool use, planning loops, and the failure modes they bring.
- Next.js, or having shipped the frontend of the thing you built the backend for.
- Having run an LLM feature that cost too much and having been the person who fixed it.