Build an Agent Workbench You Can Inspect
Run a fictional plan, draft and review. Find missing evidence, repair the brief and see why a changed task needs a fresh check.
Experiment 03 / Watch the work, not a typing indicator
An agent is easier to understand when you can see what it is trying to do and where it should stop. This workbench rehearses an invitation for a fictional Night Library event. You control the proposed guest count, the available evidence and the action boundary. Each step reveals an authored response; no live model runs and no messages are sent.
Let the first run reveal the problem
Begin with the default brief and run Plan, Draft and Check in order. The draft looks usable, but the review discovers that the room capacity has not been supplied. The workflow stops at missing evidence. That is a designed teaching case: the simulator follows fixed rules, so it demonstrates an interaction pattern rather than proving that any particular model will reliably notice the same problem.
Supply the approved room plan and run again. Thirty proposed guests still exceed the twenty-four-seat layout. Reduce the group to twenty-four or fewer and repeat the steps. The final label becomes ready for human review. It does not claim that a venue is booked, a person has approved the invitation or any message has left this page.
A changed brief needs a fresh review
Change the guest count after a completed run. The old steps disappear and the workbench tells you that the previous review was cleared. This is a small but consequential behaviour. A result checked for one set of inputs should not remain marked as reviewed after those inputs change. The interface makes the dependency explicit instead of carrying a green badge forward.
You can also remove the draft-only boundary. The check then identifies the requested send action as outside the exercise. The checkbox changes what the fictional requester asks for; it never grants this browser software a sending capability. Restore the boundary and run again. In a real agent system, tool permissions would need to enforce that separation as well as describe it on screen.
Design the work record before adding more tools
To build your own workbench, choose a task with an obvious finished artifact and a small number of facts. Ask the assistant to separate the plan, the draft and the evidence check. Give each stage a clear input and a visible output. Include one case that should stop, so you can inspect how an incomplete result is communicated.
Keep the model and the workflow separate in your design. A model may write the draft, while ordinary program logic tracks which inputs were reviewed and which tools are available. A human can judge meaning that a simple rule cannot. You do not need every part to be intelligent. You need to know which part is responsible when a result is incomplete or a requested action is out of scope.
A request you can adapt
Build a local rehearsal for an invitation workflow with Plan, Draft and Check stages. Use fictional records and authored outputs. Check that the supplied room capacity covers the guest count. Reset review whenever the brief changes. Show missing evidence and keep all outputs draft-only. Include a downloadable run record.
Export what actually happened
The run record contains the current brief and only the steps you revealed. Download after one step, then after three, and compare them. A plan alone should not become a completed workflow in the export. The labels also retain the fact that this is a simulation, so a file separated from the page does not imply that a real agent performed the task.
For a next version, use a different fictional workflow: preparing a reading list, organizing a small exhibit or reviewing a simple project brief. Add one new failure case at a time and make the recovery understandable. A useful agent experience gives people a way to notice a problem, supply what is missing and continue with an accurate record of the changed work.
Take this into your next task
Make each stage visible. Treat changed inputs as new work, and keep a reviewed draft distinct from an action in the world.
Questions people ask
Is this a live AI demo?
No. Astra in Codex helped create these original browser programs. Their controls run locally with authored rules; they do not call a model.
Can I keep what I make?
Yes. Download the current output, or explicitly save a world to your local creations shelf with its settings and camera view. The notebook lets you reopen worlds and export or import a library backup. Unsaved experiments last only while the page is open.
Sources and editorial notes
Reviewed 2026-09-10. Product capabilities depend on the app, plan, region and workspace settings. Workflows and prompts are authored teaching material, not recorded model results or measured time-saving claims.
Keep going
- Make an Evidence Atlas with AI
- AI Mission Control: From Idea to Verified Result
- From a Brief to a Working Demo with Astra
Related Power of AI pages
Keep reading with Start here, everyday uses, the tool guide, the writing workshop, and sources and standards.