Services
AI Architecture Review
- Duration
- 1 week
- My role
- I review and advise. Your team implements, or we scope a sprint for me to do it.
You have an AI system that works in demos and mostly works in production, and you need someone outside the team to tell you where it will break before it does. The review reads the architecture and the code, exercises the system against real traffic or evaluation data, and returns a ranked list of risks with the fix for each.
Discuss this engagement
Signs this is the right engagement
- Output quality varies in ways the team cannot explain from the logs
- Human reviewers still check most outputs, so the automation saves little
- Cost or latency per request is higher than expected and nobody owns the number
- A launch, funding round, or compliance review is coming and the architecture has not been challenged
How it runs
- Kickoff call (45 minutes): goals, constraints, and what keeps you up at night
- Days 1 to 3: read the architecture and code, trace real requests end to end, run the evaluation data
- Day 4: write the risk register and recommendations
- Day 5: 60-minute walkthrough with your engineers, with the written report delivered the same day
What you get
- Written risk register, ranked by likelihood and blast radius
- Prioritized recommendations with effort estimates
- An evaluation plan you can run before and after changes
- 60-minute walkthrough call with your engineers
What I need from you
- Read access to the architecture docs and code
- A sample of real traffic or evaluation data
- One 45-minute kickoff call
When this is the wrong engagement
- You do not have a system yet and want help deciding what to build (start with a prototype sprint)
- You want a vendor or model recommendation without a system to evaluate it against
Afterwards
Most teams either implement the recommendations themselves or book a prototype sprint for the hardest item.
Relevant work
Fintech · In production
98.7% extraction accuracy
A dual-LLM extraction and judge pipeline reaching 98.7% accuracy and cutting manual review time by 78.3% across financial documents.
Manufacturing, industrial data · Delivered
18 verification steps before any answer, backed by 700+ automated tests
A plain-language query layer over a manufacturing data warehouse that runs 18 verification steps before answering, and says so when it cannot answer reliably.
Related writing
Essay · September 2, 2026 · 10 min read
One model doing the work and an independent mechanism checking it is the single most reliable pattern I know for production AI. Here is how to build it.
Essay · September 2, 2026 · 12 min read
Why the most important design decision in a production AI system is deciding when the model is not allowed to answer, and how to engineer that refusal.
Pricing
Every engagement is a fixed quote agreed in writing before work starts. The quote moves on four things:
- How many systems, data sources, and environments are in scope
- Whether the work happens inside regulated or restricted environments (data residency, audit trails, access approvals)
- How fast you need it: a compressed timeline costs more than a steady one
- How much evaluation data and system access already exists versus needs to be built first
Start with a short note
Describe the system and where it hurts. I reply within two business days; if there is a fit, we schedule a 30-minute call and I send a written scope within a week.