Anatomy of a Long-Running Agent: Session, Harness, and Sandbox
Design long-running agents by separating durable session history, evolving harness logic, and disposable execution environments.
Read the technical blogs, and use the chatbot.
Blog. Projects. Chatbot.
Inside the lab
Write-ups, experiments, and whatever comes out of them.
Long-form breakdowns on agents, retrieval, multimodal systems, evaluation, and the practical tradeoffs behind them.
ExploreSmall products and experiments that turn the notes into working interfaces, pipelines, and demos.
ExploreA usable chat surface for local and cloud models so the site is not just read-only commentary.
ExploreLatest writing
Fresh blog posts on agents, retrieval, multimodal systems, and the machinery behind shipping useful AI products.
Browse all postsDesign long-running agents by separating durable session history, evolving harness logic, and disposable execution environments.
Evaluate always-loaded tools, deferred discovery, and programmatic orchestration before expanding an agent tool catalog.
Understand the experimental MCP Tasks extension for long-running tools, including polling, cancellation, human input, and authorization.
Measure coding-agent productivity using accepted work, review effort, and later rework instead of generated code or perceived speed.