Field notes · Playbooks

AI delivery, by the playbook.

The AI delivery playbooks behind every engagement: how we take LLM features and the platforms around them to production. Short, specific and tested across 150+ projects.

AI delivery playbooks we run every week.

Each one started as a lesson from a real engagement. See them at work in our prior-auth case study, and in how our QA and LLM evaluation team gates every release.

Onboarding

Four weeks to credible velocity.

Domain shadowing, an architecture and risk map, a first slice shipped behind a flag, then a roadmap and team handoff.

See the four weeks
Delivery

Written hypothesis first.

Before any code, we agree the problem, the success metric and the eval criteria in writing. The hypothesis is the contract.

Our operating rules
Release

Feature flags everywhere.

Every new path, model or prompt lives behind a flag for at least one sprint: 1% of users, then 10%, then everyone. No 3 AM rollbacks.

Performance

Performance is a budget.

LCP, p95 latency, model response time and bundle size are enforced in CI from the first commit, so speed never becomes a retrofit.

Cost

Cost guardrails in CI.

Infracost runs on every PR. If a change adds more than $200/month, the reviewer sees the number before it ships.

Operations

Don't disappear at launch.

Observability, on-call, eval reviews and post-launch tuning. We stay until the system runs smoothly in production, not just until it ships.

Put them to work

Want these on your project?

Every AI engagement ships with these practices from week one.