Skip to content

Work

Engagements that had to hold under load.

Anonymized because the systems are still running. Metrics are from the work, not a pitch deck.

System map

  • Users
  • Surfaces
  • Call paths
  • Retrieval
  • Models / tools
  • Evals
  • Serving
  • Cost
  • Owners
The ridge was serving and cost — one path, one bill.

Series C analytics platform · B2B data

An inference bill that was becoming the product

41%lower inference cost

p95 latency −28%

The team had shipped three model-backed features in a year. Each made sense alone. Together they shared no cache, no routing policy, and no cost owner.

We mapped every call path, introduced a shared serving layer with semantic cache and model routing, and put evals on the quality bar the product already promised customers.

Finance got attribution. Engineering got a paved path for the next feature. The copilots stayed.

  • Serving topology
  • Cost attribution
  • Quality evals
  • On-call playbook

System map

  • Users
  • Surfaces
  • Call paths
  • Retrieval
  • Models / tools
  • Evals
  • Serving
  • Cost
  • Owners
The ridge was quality — gold set, citations, refusal.

Legal-tech product team · Legal

Retrieval that could survive a partner review

3.4×citation accuracy

Refusal rate made explicit

Hallucinated citations are not a model problem; they are a system problem. The pipeline rewarded fluency and hid misses.

We introduced a gold set from real matters, citation verification, and a refusal path that was allowed to say 'not in the corpus.' Retrieval and chunking were redesigned against that bar.

Partners started using it. Support tickets about 'the AI made something up' dropped off the weekly review.

  • RAG architecture
  • Eval harness
  • Citation policy
  • Human review loop

System map

  • Users
  • Surfaces
  • Call paths
  • Retrieval
  • Models / tools
  • Evals
  • Serving
  • Cost
  • Owners
The ridge was the platform boundary — one runtime, named owners.

Internal platform org, public company · Enterprise

An agent platform instead of twenty agents

7 → 1runtime to maintain

Time-to-first-agent: days, not quarters

Shadow AI spend was showing up in six cost centers. Security had no inventory. Each team was solving tool use, auth, and tracing from scratch.

We defined a thin platform: identity, tool gateway, tracing, evals, and a starter runtime. Divisions kept product logic. The platform kept the dangerous parts.

An engineer on the inside now owns it. We left the RFC, the paved-path docs, and a working control plane.

  • Platform boundaries
  • Tool gateway
  • Tenancy
  • Enablement

Similar terrain? We’ll tell you honestly if it rhymes.

Start a conversation