Published onAugust 31, 2026llmagentsopenaiToo Many Tools: Testing Agent Tool DiscoveryEvaluate always-loaded tools, deferred discovery, and programmatic orchestration before expanding an agent tool catalog.
Published onJune 30, 2026agentsopenaievaluationAre Coding Agents Actually Saving Time? Measure Review and ReworkMeasure coding-agent productivity using accepted work, review effort, and later rework instead of generated code or perceived speed.
Published onApril 30, 2026llmopenaiai-engineeringLLM Routing and Fallback: Preserve the Contract, Measure the TradeoffBuild model routing and fallback policies that preserve application contracts while measuring quality, latency, and total cost.
Published onMarch 31, 2026llmagentsai-engineeringLong-Running Agent Harnesses: Plans, Artifacts, and Evaluator LoopsLearn how plans, durable artifacts, and evaluator loops help coding agents complete long-running software tasks without losing direction.
Published onJanuary 12, 2026llmtestingevaluationLLM Evaluation: Testing AI Systems That Actually WorkLearn how to evaluate LLM applications effectively. Covers evaluation frameworks, metrics, test set creation, and continuous monitoring strategies.