Running Multiple AI Agents: Cognitive Load, AI Fatigue and Judgment
Running multiple AI agents raises cognitive load one approval at a time. How to manage AI fatigue, avoid burnout, and keep your judgment.
· By Guilherme Salgueiro
Nobody who runs several AI agents at once loses control in a single dramatic moment, because the agents never actually overrule you. What happens instead is quieter: you sign control away one approval at a time, and you rarely notice the point at which reviewing turned into accepting.
The pattern is easy to recognise once you have lived through a full day of it. The first result of the morning gets your full attention, so you read the diff, question an assumption and ask for a change. By the fourth agent and the eleventh approval you have started to skim, and by the twentieth you are clicking through, which means the work still carries your name while the judgment behind it left somewhere around approval number nine. Because that drift from deciding to accepting feels like productivity while it is happening, it is also the hardest kind of failure to catch.
The answer is not to run fewer agents out of caution, since used well they create real leverage. The answer is to recognise that supervising agents is a separate job with its own limits, and to design that job as deliberately as you design the work you hand off, because judgment scales with how you design the work, not with how many agents you run.
- Supervising AI agents is its own form of work, and early survey data suggests it tires people faster than simply using AI, with self-reported productivity peaking at around three tools at once.
- Once you are past your capacity, extra approvals stop adding safety and turn into rubber stamps, which is how ownership is actually lost.
- Your bottleneck is review rather than generation, so the number of agents you run should be limited by how many results you can judge before they go stale.
- Ownership comes from designing three things on purpose: the decisions only you make, your limit on work in progress, and the evidence every handoff must bring back to you.
- Practical habits help on both sides of the screen: one worktree per agent, one important change in focus at a time, notifications only when an agent needs you, and real breaks, ideally outdoors, to restore the attention that judging consumes.
- The research is still early, so treat every number in this piece as a starting point to measure for yourself rather than a rule.
It's the old context-switching problem, now running at model speed
If that slide from deciding to accepting sounds familiar, you have probably also met what comes with it: the slightly buzzing, slightly foggy feeling at the end of a day in which you shipped a great deal and cannot quite remember deciding any of it. Francesco Bonacci, founder of Cua AI, called this "vibe coding paralysis" early in 2026, and his description is worth reading slowly, because it explains why the most capable and enthusiastic people tend to fall into it first: "The paradox: the more capability you have, the more you feel compelled to use it. The more you use it, the more fragmented your attention becomes. The more fragmented your attention, the less you actually ship" (Fortune, March 2026).
Fragmented attention is not a new problem, and that is the most useful thing to notice about it. Anyone who has tried to write a strategy document while answering email and keeping one eye on Slack already knows the feeling, and the research on it is old and consistent. In 2009 the management researcher Sophie Leroy showed that when people switch away from an unfinished task, part of their mind stays behind, a lingering load she called attention residue, and their performance on the next task drops as a result. A year earlier, Gloria Mark of UC Irvine and her colleagues found that people who are constantly interrupted compensate by working faster, finishing the same work to the same quality but paying for it with more stress, frustration and effort after only twenty minutes.
What agents change is not the nature of the problem but its speed. Email arrived at the pace of other humans, so even a bad inbox gave you minutes between interruptions and let you decide when to look. Agents return at the pace of the model, which means five sessions that each finish in a few minutes can put a new result in front of you almost every minute, and every one of them leaves an unfinished task behind in your head. Each switch also ends in something heavier than a reply, because you are not acknowledging a message but judging work you did not write, so the load grows with the speed of the model, multiplied by the number of sessions you keep open, multiplied by the weight of each decision.
In March 2026, Boston Consulting Group put numbers on what that multiplication does to people. Its researchers surveyed 1,488 full-time US workers about how they used AI and how it left them feeling, and they called what they found "AI brain fry" (The Decoder, March 2026). Fourteen percent of AI users recognised it in themselves, rising to 26% in marketing teams, and that group reported 33% more decision fatigue and 39% more major errors, along with a stronger urge to leave, as intent to quit rose from 25% to 34%. Crucially, the strain followed oversight rather than usage: people with heavy oversight demands reported 14% more mental effort, 12% more mental fatigue and 19% more information overload, and self-reported productivity rose from one AI tool to two, peaked at around three and fell once people juggled four or more, which looks like the same kind of ceiling the multitasking research has described for decades.
Here is the twist that turns the study from a warning into something more hopeful. The same survey found that people who used AI to take repetitive work off their plate reported 15% lower burnout, so the technology that was frying one group's judgment was giving another group its energy back. The difference was not how much AI they used but whether it removed tasks or generated interruptions, and that suggests a simple exercise before you change anything else. List the agents and automations you run, and mark each one as either taking a task off your plate completely or sending you a fresh stream of results to check. The first kind is quietly paying you back, while the second kind is an interruption machine running at model speed, which makes that second list the place where your limits, your schedule and your attention belong.
The numbers deserve two caveats, since everything in the study is self-reported and BCG advises companies on AI transformation. The peak at three tools is therefore a signal rather than a law, and it will move with the stakes of the work, so someone running well-tested, low-risk agents may handle more while someone reviewing legal drafts may struggle with two. What holds up is the shape of the curve, in which capability keeps rising while your capacity to judge it flattens and then falls, or, as one senior engineering manager in the study put it, "I was working harder to manage the tools than to actually solve the problem."
The instinctive response to all of this is to add more checkpoints so that nothing slips past a tired reviewer, and that is exactly where the next trap is waiting.
More approvals can make you less safe, not more
Adding checkpoints feels like the responsible answer to a tired reviewer. If attention is slipping, surely the system should ask you more often? Hospitals tried a version of that idea long before anyone ran a coding agent, and how it played out is the clearest warning we have.
In the summer of 2010, a 60-year-old man lay in an intensive care unit with a head injury from a falling tree branch. For an hour his monitors raised alarm after alarm as his heart rate climbed and his oxygen fell, but staff only reached him once a critical alarm signalled that he had stopped breathing, and by then the damage could not be undone. The monitors had done their job. The people listening to them had simply stopped hearing.
When the Joint Commission looked at the wider pattern in its 2013 alert on alarm safety, it found the same story across US hospitals. A single patient could set off several hundred alarms a day, and up to 99% of them needed no action at all, so clinicians learned to tune them out. The commission named the most common factor behind deadly alarm failures "alarm fatigue".
Every agent approval is a small alarm. You read the first permission prompt of the day carefully, and you click the fortieth before your eyes have finished crossing it, not because you are careless but because attention runs out.
A 2026 paper, Oversight Has a Capacity, models exactly this for AI agents. It imagines two guards watching the same coding agent, one that asks its human to approve five actions a day and one that asks for five hundred. Most people pick the second as the safer one, yet by the three-hundredth routine approval that reviewer is clicking on autopilot, and the one dangerous action buried in the pile sails through. The paper finds that safety rises as you escalate more, peaks, and then falls, so escalating everything is worse than escalating selectively. It also points out the darker trick that follows: to sneak something past a reviewer, you don't need to hide it, only to bury it under enough harmless requests that it reaches someone who has stopped weighing anything.
This is a simulation rather than a study of real people, and the author says so. But it matches what hospitals learned the hard way, and what you already feel by mid-afternoon of a busy agent day.
An approval you don't really read is not oversight but a signature. Ownership rarely disappears because an agent overrules you; it disappears because you rubber-stamp it.
The hopeful part of the story is what hospitals did next, because they neither added more alarms nor ripped the monitors out. From 2014 the Joint Commission made alarm safety a national patient safety goal, and once hospital leaders had made it a priority, the next thing it asked of every hospital was surprisingly simple: decide which alarms actually matter. During 2014 each hospital had to name the signals that put patients at real risk, and by 2016 it also had to agree when an alarm could safely be switched off and who had the authority to change it. The goal was never silence. It was to make sure that when the alarm that mattered went off, someone would still be listening.
Go back to that intensive care unit for a moment. The tragedy was not that the monitor stayed quiet, because it shouted for an hour, but that it was one voice among hundreds and nobody could tell it apart from the noise.
Your agents are heading down the same road. Somewhere in today's stream of approvals there is one that genuinely matters, whether it is the migration that touches customer data, the email that goes to a real client or the change you cannot roll back, and if it reaches you as the forty-first request of the afternoon it will get the same tired click as the forty before it. Reading more carefully will not save you, because attention is exactly the thing that has run out. What saves you is designing the day so that the decision that matters arrives as the third thing you see rather than the forty-first, while everything that does not need you runs on its own or waits for a quieter moment. You cannot make your attention any bigger, but you can decide what deserves it, and the rest of this article is about how to make that decision on purpose.
Your bottleneck is review, not generation
To decide what deserves your attention, it helps to see where the pressure actually comes from, and one of the best explanations of it was filmed in 1952, about seventy years before anyone ran a coding agent. In "Job Switching", the episode that opened the second season of I Love Lucy, Lucy and Ethel take jobs in a chocolate factory and end up at the wrapping station. The rule is simple: wrap every chocolate that passes on the conveyor belt, and if a single one reaches the packing room unwrapped, you are fired. The belt starts gently and then speeds up, and within a minute the two of them are stuffing chocolates into their mouths, their hats and their blouses, anywhere at all, as long as nothing gets past them unwrapped. Then the supervisor walks in, sees an empty belt, decides they are doing splendidly and shouts to the kitchen: "Speed it up a little!"
It is one of the funniest scenes in television, and it is also an uncomfortably accurate picture of a day with five agents running. Generation is the belt, and its speed is now set by the models rather than by you, while your judgment is the wrapping, which still happens one piece at a time. Pere Villega describes the moment the belt overtakes you in From One Agent to Many (June 2026): one agent returns a patch and you review it, but five agents return five patches in the same hour, built on different assumptions, two of them touching the same interface and all five waiting for a decision. As he puts it, "The models may be faster. You are not."
What happens next is exactly what happened in the factory, because nobody openly gives up; they start hiding chocolates. A Microsoft study of 17 experienced developers found that people opt for "efficient, not perfect, oversight", skimming the output and letting passing tests stand in for proof that the code is correct. From the outside it looks like success, since the belt is empty and everything has been approved, so the natural conclusion is the supervisor's: start another agent and speed it up a little.
The hidden chocolates do not stay fresh, either. Finished work that waits for you starts to decay, because the codebase moves underneath it, you lose the context behind each task, and agents get restarted to fix conflicts that only exist because a result sat too long in the queue. That is why Villega's reframing is the most useful question in this whole topic:
"The useful question is not 'How many agents can I start?' but 'How many results can I responsibly decide upon before they decay?'" — Pere Villega
The teams that cope best have stopped trying to wrap faster. A small interview study of practitioners presented at ESEM 2026 found them moving routine checks into linting, tests and continuous integration, which is the equivalent of building a machine to wrap the plain chocolates, so that human attention is saved for the pieces that need it: architecture, intent and whether the change will still make sense in a year.
The real lesson of the scene is not that Lucy and Ethel needed faster hands. It is that nobody was in control of the belt. Villega's rule of thumb captures it: if your agents are waiting for tasks, that may be fine, but if finished work is waiting for your judgment, starting another agent is the wrong optimisation. Taking back control of the belt is what the next section is about.
Ownership means designing the belt, not running faster beside it
While American sitcoms were laughing at conveyor belts, Toyota was quietly building a very different one, and the idea behind it is the best answer we have to the chocolate factory. The Toyota Production System rests on two pillars, and the way Toyota describes the first one could have been written for anyone babysitting agents: its purpose is to eliminate "the need for people to be simply watching over machines". Machines stop by themselves the moment something is wrong, any worker can pull a cord to halt the line, and the factory makes only what is needed, when it is needed, in the amount needed. People are not there to watch everything. They are there to respond when something genuinely needs them.
You can design your own work with agents the same way, and it comes down to three decisions you make before the belt starts rather than after you are already drowning.
1. Decide which decisions are yours
On a Toyota line, everyone knows who can pull the cord and why. With agents, the equivalent is sorting work into tiers before you start, so that your attention goes to the decisions that need it and nothing else. Practitioner guides such as AgentPatterns use three, and the version below is a starting point to adapt rather than a standard.
| Tier | What it covers | Your role | Example | |---|---|---|---| | Fully delegated | Reversible, cheap to verify, covered by tests or checks | Read the summary, not the work | Formatting fixes, dependency bumps with green CI, first-pass research | | Delegated with checkpoints | Real changes, but bounded and recoverable | Approve the plan before, check the evidence after | A feature behind a flag, a refactor of one module, a draft for publication | | Human only | Irreversible, ambiguous, or about intent and trade-offs | You decide, and agents prepare options | Architecture, pricing, anything legal, what to build next, what to stop |
The table matters less than the act of filling it in, and the timing matters most of all. Write your human-only list at the start of the day, while you are fresh, because any decision you have not claimed in advance is one you will end up approving by default at four in the afternoon.
2. Make only what you can review
Toyota's second pillar, making only what is needed when it is needed, translates into a limit on how much unreviewed work you allow to exist at once. Villega suggests an admission test before starting another agent: the task should have a clear boundary, a known starting point, a visible finish line, an owner who decides whether the result is kept, and, most often forgotten, a realistic slot in your day to review it soon after it arrives. If you cannot say when you will look at the result, you are not delegating the task but deferring a decision and leaving it to go stale.
A sensible starting point is fewer finished results rather than many open ones, such as one or two implementation tasks and one research task running at the same time. Raise the limit only after you have watched where the work waits, and if it waits on you, the limit is already too high.
3. Let the work stop itself
The last idea is the most powerful, because it is what lets you stop watching. Toyota's machines detect their own problems and stop, and your agents can be set up to do the same by bringing evidence with every result instead of asking you to discover the problems yourself. A study of how experienced developers supervise coding agents (KAIST and Carnegie Mellon, September 2026) points to three habits that make this work: putting most of the effort into planning, when changing direction is still cheap; handing routine review to other agents, sometimes from a different model provider; and turning any guidance you keep repeating into a test, a hook or a line in an instruction file, so the check runs without you from then on. As one developer in the study put it, written guidance can be quietly deprioritised once the context fills up, whereas "a hook is code; it fires every time whether the model thinks it's relevant or not."
A simple way to make every result carry its own proof is to ask each agent to finish with a short handoff card:
Most of the card is evidence you can scan in seconds, but the last two lines are your stop cord. "Assumptions" shows you where the agent made a call on your behalf, and "Needs you" is the one decision that genuinely belongs to you. When those lines are empty on a task that should have had them, that is your signal to slow down and read properly.
Put the three decisions together and the day changes shape. Instead of standing beside a belt that keeps speeding up, you are running a line that produces only what you can absorb, stops itself when something is wrong, and calls you only when something genuinely needs you.
Keep your thinking in the loop, not just your signature
A well-designed line protects your attention, but there is a quieter way to lose ownership that no line can catch, because it happens inside your own head while you approve the work. Driving shows it best. In a study published in 2020, researchers at McGill University tested the spatial memory of 50 regular drivers and retested a small group of them three years later, and the more someone had leaned on GPS in between, the more their spatial memory had declined. The GPS still got every one of them where they were going, which is exactly what makes a decline like that so easy to miss.
AI offers the same trade for judgment. A study of 589 students and early-career knowledge workers (Frontiers in Psychology, July 2026) found that people who leaned on AI to do their core thinking, rather than using it as scaffolding, reported weaker independent judgment and lower motivation, even though both styles of use felt equally helpful in the moment. A review of 35 studies on automation bias adds that asking the AI to explain itself rarely helps; what protects you is doing some of the thinking yourself.
So the answer is not to switch the GPS off but to keep driving part of the route, and three habits that take a minute each do most of the work:
- Write your intent before you read the output, so you compare the agent's answer with yours instead of adopting its framing.
- Make the call first on your human-only decisions, then ask the agent to challenge it.
- Explain it back. If you cannot say in plain words why a result is right, you have accepted it rather than judged it.
A field guide for running several agents at once
Everything so far has been about principles, so here is how they look on an ordinary Tuesday. None of these tips needs a new tool, and each one removes a specific way the day goes wrong.
- Give every agent its own copy of the code. Two agents editing the same checkout will overwrite each other's work, which is why Anthropic's own Claude Code guide recommends running parallel sessions in separate git worktrees, so that concurrent edits never collide and each result can be reviewed on its own.
- Keep one important change in focus and parallelise the rest. Simon Willison, one of the most careful public voices on coding agents, describes his own rule in Embracing the parallel coding agent lifestyle: "I can only focus on reviewing and landing one significant change at a time", so the agents running beside it get research questions, explanations of how existing code works and small maintenance chores, the kind of work that does not compete for the same slice of judgment.
- Write the brief before you press go. Willison's second observation is that code built from your own specification is far less effort to review than code that "lands on your desk out of nowhere", because you already know what it was meant to do and only have to check that it did.
- Name each session after its goal. "Billing: fix refund rounding" tells you in one glance what a session is for, whereas five identical terminal tabs turn every return to your desk into a memory test.
- Notify on "needs you", not on "done". An agent that has finished can wait for your next review block, but one that is blocked on a decision is the only kind that deserves to interrupt you, so set your notifications accordingly and let the rest queue quietly.
- Merge one result at a time. Review and land each change before you open the next, so that later work starts from a base that already contains it and conflicts surface while they are still small.
Recover as if it were part of the job, because it is
All of those tips protect your attention while you work, and the other half of the job is giving it back afterwards. The simplest test of whether you managed it does not involve a dashboard: try reading a book at seven in the evening after eight hours of running ten parallel sessions. If the words slide past and you reread the same page three times, that is not a lack of discipline. It is the bill for the day, and it tells you more about your real limit than any productivity metric ever will.
The good news is that this kind of tiredness responds to small, unglamorous habits, and the research behind them is solid even if the effects are modest:
- Take real breaks, and make the hard days' breaks longer. A meta-analysis of 22 studies with 2,335 participants (PLOS ONE, 2022) found that micro-breaks of ten minutes or less modestly but reliably raised energy and reduced fatigue, but that getting your performance back after highly demanding work may take more than ten minutes, which is exactly the kind of work that judging agent output is.
- Move, or deliberately switch off. In a study of students during four-hour lectures, a break with light exercise left people with more energy and a relaxation break left them less fatigued, and both effects were still there twenty minutes after they returned to work.
- Go outside for half an hour. A 2025 meta-analysis of 80 studies found that time in nature restores overall cognitive capacity better than time in built environments, with the largest benefits at around thirty minutes. The effects are small and vary from study to study, but a walk in the park at lunch beats another coffee at the desk.
- Close one loop before you open the next. Remember Leroy's attention residue from earlier: unfinished tasks keep running in the background of your mind, so finishing or explicitly parking a result before switching is a recovery habit as much as a productivity one.
- Batch your reviews. The Berkeley researchers who followed a tech company for eight months suggested grouping AI work into set blocks with deliberate pauses before demanding decisions, which gives your mind stretches of time when nothing is asking it to judge.
- Shut the day down on paper. Before you stop, write down what is still running, what is waiting for you and what you decided, so that tomorrow starts from a note rather than from whatever your tired memory kept.
It is also worth being honest about the evidence, because nobody has yet measured your particular limit. The fatigue data is self-reported, the approval-capacity result is a simulation and the developer studies each involve between five and nineteen people, and a widely shared claim that errors rise 39% beyond three agents traces back to a weakly sourced review rather than to any of them. So treat every number in this piece as a hypothesis, and test it on yourself for two weeks by noting how many results you accepted, how many you genuinely reviewed and how often finished work waited for you. The gap between the first two numbers is your rubber-stamp rate, and watching it shrink is more satisfying than any agent count.
The work can carry your name. Make sure it carries you.
If this whole article had to fit on an index card, it would read something like this: claim your decisions before the belt starts, keep the line short enough to review, let the work stop itself, drive part of the route yourself, and take your breaks as seriously as your deadlines.
Lucy and Ethel finished that famous scene with their cheeks full of chocolate and a supervisor delighted by an empty belt, and plenty of agent days end in much the same way, with everything approved, very little genuinely decided and a strange tiredness that is hard to name. The ICU monitor in 2010 shouted for an hour and was not heard. The drivers in Montreal always arrived and slowly forgot the way. None of those stories is about a machine failing. Every one of them is about people who were present and no longer quite there.
Now picture the other version of the day. Fewer approvals, each one actually read. One cord pulled at exactly the right moment. A handful of results you understood well enough to defend. And at seven in the evening, a book that holds your attention, a dinner where you are listening rather than replaying diffs, and a decision about your own life made with the same mind you spent the day protecting.
That is what this is really about, and it was never throughput. The agents will run again tomorrow, as tireless as they were today. You get one mind, and every approval you sign spends a little of it. Spend it on the decisions that are truly yours, and the work will carry more than your name. It will carry you.
Frequently asked questions
What is cognitive load when supervising AI agents?
It is the mental effort of tracking several agents' context, reviewing their output, and approving their actions. It grows with every parallel agent, and when it gets too high, judgment degrades and AI fatigue sets in.
How many AI agents can one person supervise at the same time?
There is no proven number. In a BCG survey of 1,488 US workers, self-reported productivity rose up to about three AI tools used at once and fell after that, but the data is self-reported. A better limit is review capacity: run only as many agents as you can judge soon after their results arrive. If finished work is waiting for your judgment, adding another agent is the wrong optimisation.
What is AI brain fry?
AI brain fry is the term a 2026 BCG study used for mental fatigue caused by excessive use or oversight of AI tools beyond a person's cognitive capacity. Workers with heavy oversight demands reported 12% more mental fatigue and 19% more information overload. Those reporting brain fry also reported more decision fatigue and more major errors.
Why can more human approvals make AI agents less safe?
Every approval request spends the same finite pool of attention. Past a reviewer's capacity, approvals become automatic clicks, so a risky action buried in a stream of routine ones gets rubber-stamped. A 2026 modelling study found that safety peaks at an escalation rate below 'escalate everything'. It is a simulation, not a human study, but it matches decades of alarm fatigue in hospitals and alert fatigue in security operations.
How do you keep ownership of work done by AI agents?
Design three things on purpose: the decisions only you make, a work-in-progress limit set by how many results you can judge before they go stale, and the evidence each handoff must bring you. Put your judgment into planning and acceptance criteria, let tests and independent reviewers handle routine checks, and keep architecture, intent and trade-offs for yourself.
How do you restore focus after a day of supervising AI agents?
Treat recovery as part of the work. A 2022 meta-analysis of 22 studies found that short breaks modestly but consistently reduce fatigue and raise energy, though demanding work may need breaks longer than ten minutes. A 2025 meta-analysis of 80 studies found that nature's restorative edge over built settings is largest at about thirty minutes, though the effects are small. Batch reviews into set blocks, finish or park one task before switching, and end the day with a written note of what is still open.