An agent that does something small, every day, without breaking is more impressive than a demo that does something spectacular once. Small and dependable is the whole game.
Pick a narrow chore
Good first candidates share a shape: repetitive, low-stakes, and easy to check. Sorting incoming requests into categories. Summarising a feed you skim anyway. Checking a list for missing fields and flagging the gaps. All of these are reversible, and all of them produce output you can glance at and grade.
Bad first candidates touch money, send things to customers, or write to a system of record.
Four things it needs before you leave it alone
- A narrow brief. Write what the agent does, what it must never do, and what to do when it isn't sure. Ambiguity gets resolved at runtime, badly, if you don't resolve it now.
- Memory of what it already handled. Without it, agents redo work and re-notify people. A list of processed items is usually enough.
- A spend and step limit. Loops are the default failure mode. Cap the number of steps and the monthly cost.
- A report you actually read. One daily line: what it did, what it skipped, what it wasn't sure about. An agent nobody reviews is an agent nobody trusts.
Run it beside you first
For the first week, let it propose and let yourself approve. You are looking for the cases it gets confidently wrong, because those are the ones your brief hasn't covered yet. When a week passes with no surprises, loosen the leash one notch.
Here is the chore: [describe it]. Write me an operating brief for an agent doing this, including: what counts as done, three cases it should refuse and escalate, what it should log for each run, and what should stop it entirely.
Try this today
Name one chore that is repetitive, reversible and easy to check — and write the three cases where you would want it to stop and ask.

