
A practice for production AI
Engineeringfor the terrainof production AI.
Dunes helps teams design, harden, and operate the systems that actually run — inference, platforms, retrieval, evals, and the organizations around them.
Now booking Q4 2026 and Q1 2027 · Diagnostics from $22k
Typically engaged by
Applied AI orgs
who already shipped
Platform teams
who own the paved path
Infra leads
who feel the GPU bill
Engineering leads
who need a second chair
Services
The landscape under the models.
01
Production AI systems
LLM platforms, retrieval, agents, and the evaluation loop that keeps them honest once real users arrive.
Read more02
AI infrastructure
Inference, serving, GPU and cost envelopes, data planes — the landscape under the demos.
Read more03
Architecture & platforms
Architecture, technical strategy, and the internal platforms teams actually want to use.
Read more04
Readiness reviews
Time-boxed audits before you scale spend, go live, or put an agent in front of customers.
Read more
Approach
Map, name, move, leave a trail.
01
Map the terrain
Current architecture, constraints, spend, failure modes, and who actually owns what. We write it down until the picture is shared.
02
Name the ridge
Most programs have one or two moves that change the slope. We name them, sequence them, and kill the rest of the backlog theater.
03
Build or bind
We implement alongside your engineers, or we sit in the reviews while they do. Either way the system moves, and ownership stays with you.
04
Leave a trail
Evals, runbooks, decision records, cost envelopes, named owners. The engagement is done when the path is walkable without us.
Selected work
Systems that had to survive contact.
System map
- Users
- Surfaces
- Call paths
- Retrieval
- Models / tools
- Evals
- Serving
- Cost
- Owners
The ridge was serving and cost — one path, one bill. B2B data
An inference bill that was becoming the product
Realtime copilots had quietly become the largest line item. We redesigned serving, caching, and routing so quality held while unit cost fell.
41%lower inference cost
System map
- Users
- Surfaces
- Call paths
- Retrieval
- Models / tools
- Evals
- Serving
- Cost
- Owners
The ridge was quality — gold set, citations, refusal. Legal
Retrieval that could survive a partner review
A RAG assistant was impressive in the demo and unreliable in the wild. We rebuilt evaluation, citations, and refusal behavior around how lawyers actually work.
3.4×citation accuracy
Engagement diagnostic
Describe the terrain. We’ll name the engagement.
A few sentences about the system, the pressure, and what “working” would mean. You get a recommended engagement, first moves, and what to watch. Not a sales script.
Engage
If the terrain is already real, we should talk.
A conversation, then a written proposal for one package. Diagnostics from $22k. We will say so if we are not the right fit.