AI Product Development
Most AI projects stall between the demo that impressed everyone and the version that has to survive real users, real data and a real bill. We build for the second one from the start.
How the engagement runs
Agreed before we start, not discovered later. You own everything we produce, whether or not you continue with us.
- 01
Ground truth
Eval set + baseline score
- 02
Retrieval
Working retrieval + accuracy figure
- 03
Generation
Answering system + escalation path
- 04
Production
Live system + runbook
- Starts with
- Discovery
- You get
- Running system + runbook
- Shape
- Fixed scope or retainer
You are probably here because
If none of these sound familiar, this may not be the practice you need. Tell us what is actually happening.
- 01
A prototype that works in a demo but not on real data
- 02
No way to tell whether a change improved anything
- 03
Model spend that nobody can predict or cap
- 04
An AI feature nobody trusts enough to put in front of customers
What a production AI system needs
Not all four land in every engagement. We scope to what the problem actually needs.
Retrieval that stays grounded
RAG over your own content with citations and a confidence threshold, so the system escalates instead of inventing an answer.
Agents with a boundary
Tool-calling loops with a spend ceiling, an allow-list, and human approval before anything irreversible.
Evaluation before deploy
A fixed eval set that every prompt or model change is scored against, so regressions are caught before a user finds them.
Model routing and fallback
The right model per task rather than one for everything, with a fallback path when a provider degrades.
What we can point at
Including where we cannot. An unproven claim is worth less to you than a stated limit.
An AI assistant live for a client, checked before it speaks
Appooppanthaadi: Appu answers from the live catalogue, and a second pass verifies every figure against source data before a reply is shown
See itOne published client AI engagement, not a portfolio
Most of our production AI experience is in products we run ourselves. That is real experience, but it is not the same as a shelf of client references. Ask us what we have and have not done.
Questions this answers
If you cannot answer these about your current system, that is usually where we start.
Will this hold up when a thousand people use it at once?
What does it cost per request, and what stops that running away?
How do we know a prompt change made things better, not worse?
What happens when the model is wrong?
How the engagement runs
Indicative, not a template. The shape holds; the depth of each phase moves with the problem.
Ground truth
We collect the questions your system actually has to answer and turn them into a fixed evaluation set. Nothing gets built until we can measure it.
Eval set + baseline score
Retrieval
Your content is chunked, embedded and indexed, then tuned until the right passage comes back for the questions that matter. Citations from day one.
Working retrieval + accuracy figure
Generation
Prompting, tool access and the confidence threshold that decides when to escalate. Every change scored against the eval set before it lands.
Answering system + escalation path
Production
Spend caps, rate limits, observability and the runbook your team uses when something looks wrong at 2am.
Live system + runbook
What we reach for first
Not the only things we work with: the ones we default to unless there is a reason not to.
Models
- OpenAI
- Anthropic
- Open weights
Retrieval
- pgvector
- Pinecone
- Hybrid search
Runtime
- Next.js
- Node
- Streaming
Operations
- Evals
- Tracing
- Spend caps
Tell us what you’re trying to build
An idea, a stalled project, or a system that has outgrown its architecture. Send it over and we’ll tell you where we would start and whether we’re the right team for it.
- Response time
- Within one business day
- First call
- 30 minutes, no deck
Prefer to write first? Send the brief and we’ll come back with questions before we come back with a proposal.