Tagged: evaluation

5 items · All tags

A Second Check, Always

Essay · September 2, 2026 · 10 min read

One model doing the work and an independent mechanism checking it is the single most reliable pattern I know for production AI. Here is how to build it.

essaysllm-as-judgeevaluationreliability

Kimi K3: The Architecture Was Decided By The Kernel

July 30, 2026 · 6 min read

Reading the Kimi K3 technical report from the seat of someone who builds agentic systems that have to run on Monday morning: the decisions were made by kernels, caches, and harnesses, not loss curves.

model-analysisagentic-aiinferenceevaluation

The Open-Source Agentic AI: Kimi K2.6

April 22, 2026 · 7 min read

An honest read of Moonshot's Kimi K2.6 from someone building production agentic systems: why the open weights matter more than the leaderboard, and where it still falls short.

model-analysisagentic-aiopen-sourceevaluation

Building Smarter AI Benchmarks

June 19, 2025 · 4 min read

What two controversial papers, The Illusion of Thinking and its rebuttal, taught us about measuring machine reasoning: many AI failures are benchmark design failures.

evaluationbenchmarksreasoning

5 RAG Failure Modes in Production

March 15, 2024 · 1 min read

Common ways RAG systems fail in week 2, and how to avoid them with evaluation and observability.

ragproductionevaluationlangfuse