Published onSeptember 12, 2026llmagentsopenaiAnatomy of a Long-Running Agent: Session, Harness, and SandboxDesign long-running agents by separating durable session history, evolving harness logic, and disposable execution environments.
Published onAugust 31, 2026llmagentsopenaiToo Many Tools: Testing Agent Tool DiscoveryEvaluate always-loaded tools, deferred discovery, and programmatic orchestration before expanding an agent tool catalog.
Published onJuly 31, 2026llmagentsmcpMCP Tasks: Durable Asynchronous Tool CallsUnderstand the experimental MCP Tasks extension for long-running tools, including polling, cancellation, human input, and authorization.
Published onJune 30, 2026agentsopenaievaluationAre Coding Agents Actually Saving Time? Measure Review and ReworkMeasure coding-agent productivity using accepted work, review effort, and later rework instead of generated code or perceived speed.
Published onMay 31, 2026llmagentssecurityContaining AI Agents: Filesystem, Secrets, and Egress BoundariesDesign enforceable filesystem, credential, and network boundaries that limit what an AI agent can reach when model-level safeguards fail.