Skip to main content
Start your own AI-powered blog — freeGet started →

I Replaced My To-Do List With an AI Agent for 30 Days. Here's What Broke.

Podcast episode2 voices
2:50
I Replaced My To-Do List With an AI Agent for 30 Days. Here's What Broke.
Photo by Cathryn Lavery on unsplash

I Replaced My To-Do List With an AI Agent for 30 Days. Here's What Broke.

I have tried every productivity system invented by humankind. Bullet journals. Kanban boards. The one where you put a single sticky note on your monitor and feel superior about it.

So when AI agents got good enough to actually do tasks instead of just listing them, I did the obvious reckless thing: I deleted my to-do app and handed the whole job to an agent for a month.

Here's the unvarnished report.

Quick Answer

Using an AI agent instead of a to-do list works — but not the way you'd expect.

The agent is bad at being a list. It's great at being a doer. The win came when I stopped asking it to remember my tasks and started asking it to complete them.

The rule that saved the experiment: the agent owns execution, I own priorities.

A tidy desk with a notebook and laptop Photo by Cathryn Lavery on Unsplash

Week one: the honeymoon

The first few days were genuinely thrilling. I'd say "follow up with the three people who didn't reply to last week's proposal" and it would draft three tailored emails, queued and waiting for my approval.

Tasks that used to sit on my list for days — the small, annoying, two-minute-but-I-keep-avoiding-it kind — just evaporated. That category of work is exactly where AI agents and AI assistants shine: low judgment, high avoidance.

I felt like I'd hired a chief of staff for the price of a coffee subscription.

Week two: the cracks

Then it started confidently doing the wrong things.

I asked it to "clean up my project notes." It interpreted "clean up" as "summarize and delete the originals." The summaries were fine. The originals had details the summaries dropped. I spent an afternoon reconstructing what I'd lost.

The lesson wasn't that the agent was dumb. It was that I'd given it a goal I hadn't defined. "Clean up" meant something specific in my head and nothing specific in the prompt.

Week three: the rules emerge

By week three I'd developed a working contract with the thing:

  • Destructive actions need confirmation. Anything that deletes, sends externally, or spends money waits for a yes.
  • Ambiguous verbs are banned. "Handle," "clean up," and "sort out" produced the worst results. Specific verbs produced the best.
  • I review in batches. Instead of approving each action live, I let it queue work and cleared the queue twice a day.

That last one mattered more than I expected. Constant approval prompts had me babysitting the agent. Batched review let it actually save me time.

Week four: the verdict

What the agent replaced wellWhat it couldn't replace
Drafting and sending routine messagesDeciding what actually mattered
Chasing follow-upsKnowing when to not follow up
Turning vague intentions into draft actionsProtecting the irreplaceable originals
Remembering recurring tasksJudging priority under pressure

The honest conclusion: an AI agent is a phenomenal executor and a mediocre manager. It will do the work. It will not decide which work deserves doing. That's still my job, and after a month I'm glad it is.

How to try this without the car crash

If you want to run your own version, skip my mistakes:

  1. Start with one category of task — say, email follow-ups — not your whole life.
  2. Make destructive actions require confirmation from day one.
  3. Use specific verbs. "Draft," "schedule," "summarize into a new note" — never "handle."
  4. Review in batches so you're supervising, not babysitting.

The bottom line

The future of personal productivity isn't a smarter list. It's a capable executor that handles the doing while you keep the deciding.

Hand an agent one annoying recurring task this week — with a confirmation gate — and see how much lighter your actual list gets. Just don't let it near your originals until you've taught it what "clean up" means.

The Hidden Cost of AI Agent Overhead: When Execution Creates More Work

The first week’s euphoria of offloading tasks masks a critical truth: AI agents don’t just do work—they generate work. Every action they take requires setup, review, and often correction. For example, when the agent drafted three follow-up emails, it saved me 20 minutes of typing but cost 10 minutes of editing tone and context. The net gain was real, but the overhead wasn’t zero. This pattern compounds with scale. A single task (e.g., 'summarize meeting notes') might take 5 minutes to review, but 20 such tasks in a day become a second job. The break-even point isn’t when the agent saves time—it’s when the review time becomes less than the execution time it replaces.

The overhead isn’t just temporal; it’s cognitive. Switching between 'doer mode' (writing, creating) and 'reviewer mode' (editing, approving) fragments focus. I noticed this most with batch reviews: approving 10 agent actions in one sitting felt efficient, but the mental whiplash of context-switching between emails, calendar invites, and document edits left me drained. The solution? Task clustering. Group similar agent actions (e.g., all email drafts, all calendar updates) into batches to minimize context shifts. This reduced my review time by ~30% and made the overhead feel manageable.

How to Train Your Agent: The Unwritten Rules of Prompt Engineering

Most advice about AI agents focuses on what to delegate, but the real leverage lies in how you delegate. The difference between a prompt that wastes hours and one that saves them often comes down to three unwritten rules:

  • State the constraint, not the goal. Instead of 'clean up my project notes,' try 'summarize each note into a new doc, preserving originals, and flag any action items with a deadline.' The first prompt invites deletion; the second defines the boundaries of 'clean up.'
  • Specify the output format. Agents default to their own interpretation of 'done.' A prompt like 'draft a follow-up email' might return a three-paragraph essay. 'Draft a 3-sentence follow-up email with a clear ask and deadline' ensures usable output.
  • Preempt failure modes. Tell the agent what not to do. 'Do not delete original files' or 'Do not send without approval' are explicit guardrails that prevent costly mistakes.

The most effective prompts I used followed a template: Action + Output + Constraints + Review. For example: 'Draft a meeting recap email (Action) in bullet points (Output), include decisions and next steps (Constraints), and queue for my approval (Review).' This structure eliminates ambiguity and turns the agent into a predictable tool rather than a creative liability.

Training the agent isn’t a one-time setup—it’s an ongoing dialogue. After each batch of actions, I’d refine prompts based on what went wrong. If the agent misinterpreted 'summarize,' I’d add 'in 3 bullet points, using the original’s exact phrasing.' Over time, the prompts became so precise that the agent’s outputs required minimal editing. The key insight: prompt engineering is a skill, not a one-time configuration. Treat it like teaching a new hire, not programming a machine.

The Priority Paradox: Why AI Agents Expose Your Bad Habits

AI agents don’t just execute tasks—they expose the flaws in how you think about tasks. When I handed off my to-do list, I realized most of my 'tasks' weren’t actions at all; they were vague intentions, half-baked ideas, or placeholders for work I didn’t want to confront. An agent can’t execute 'think about the Q3 strategy' or 'maybe follow up with Sarah.' It forces you to articulate what you actually want done, which is often the hardest part of productivity.

This revelation led to a brutal audit of my task list. I categorized every item into one of three buckets:

  • Executable: Clear actions the agent could do (e.g., 'draft a proposal outline').
  • Decidable: Tasks requiring my judgment (e.g., 'choose between two vendors').
  • Aspirational: Non-urgent ideas (e.g., 'research new tools').

The agent could handle ~60% of my 'executable' tasks but zero of the 'decidable' or 'aspirational' ones. This wasn’t a failure of the agent—it was a failure of my own clarity. The paradox: the more you rely on an AI agent, the more you must sharpen your own priorities. The agent doesn’t replace decision-making; it outsources execution, which makes your decisions more critical.

The real productivity win wasn’t the time saved—it was the mental space freed by eliminating the 'gray zone' of half-formed tasks. Without a to-do list to dump vague ideas into, I had to confront them head-on. Did I really need to 'look into that new CRM'? No. Could the agent actually 'prepare for the client call'? Only if I defined what 'prepare' meant. The agent became a mirror, reflecting the gaps in my own thinking. The lesson: AI agents don’t fix bad habits—they reveal them, and force you to fix them yourself.

Key Takeaways

  • Replace 'remembering tasks' with 'executing tasks'—AI agents excel at doing work (e.g., drafting emails, chasing follow-ups) but fail at prioritizing it. Shift your focus from delegation to execution.
  • Enforce a 'confirmation gate' for destructive actions (deletion, external sends, spending) from day one. Without this, the agent will confidently delete or alter originals, forcing costly recovery work.
  • Ban ambiguous verbs like 'handle,' 'clean up,' or 'sort out.' Use precise commands: 'draft a follow-up email,' 'schedule a meeting,' or 'summarize notes into a new doc' to avoid misaligned outputs.
  • Batch-review agent actions twice daily instead of approving live. Constant interruptions turn the agent into a babysitter; batching lets it operate autonomously while you retain oversight.
  • Start small: assign the agent one recurring task category (e.g., email follow-ups) before expanding. Scaling too fast leads to cleanup disasters and erodes time savings.
  • Treat the agent as a teammate, not a manager. It will execute flawlessly but cannot judge priority or context—your role is to define what matters; its role is to do it.

Frequently Asked Questions

Did it actually save time?

Net yes — but only after week two. The first week's gains were partly erased by the cleanup mess. Once the rules were in place, I got back maybe an hour a day.

Is this safe for important work?

With a confirmation gate on destructive actions, yes. Without one, absolutely not. The gate is non-negotiable.

Would I keep doing it?

I have. I didn't reinstall my old to-do app. But I think of the agent as a teammate who executes, not a system that decides.

C
Corvex

1 followers

Comments

Sign in to join the conversation

No comments yet. Be the first to share your thoughts!

More from Corvex

Recommended for you