Physicians will become managers of AI
Rounds already taught you how. Mostly.
Rounds already taught you how. Mostly.
Back to my AI setup that I previous said lived, like Harry Potter, in the cupboard under the stairs.
It has evolved in a couple of weeks.
Now it has a hierarchy. Agent Number One:
(1) is the agent I speak to the most -> and speech is the primary mode of input with text or generated artifacts in return;
(2) and manages a fleet of agents1 that often have specific roles.
However, since then, I've recruited agents to red-team Number One's plan (including Agent 1b); counselor agents; and an agent to specifically provide human feedback to the entire stack.
All of this has made me a manager of agents.
Managing AIs?
So what does this mean?
The word I keep landing on is one medicine pretends not to like: manager. But watch what an attending actually does on rounds, at the front of that travelling column of housestaff. She delegates (you take beds 4 through 9). She sets the priority order (sick before stable, discharges before noon). She verifies selectively (show me the gas, I want to see that tracing myself). She decides, patient by patient and trainee by trainee, how much rope each person gets. And she gives feedback that is supposed to change tomorrow's performance. That is management. We just called it supervision, never taught it formally, and let everyone absorb it by osmosis.
Here is the part that makes me laugh. In Canada we literally had this in the curriculum. The Royal College's CanMEDS framework included a role called the Manager for nearly two decades, and then in 2015 we renamed it Leader, partly because "manager" sounded too administrative, too middle-management, insufficiently aspirational. We retired the word about ten years before every physician was going to need it back.
The rest of the economy is not waiting for us
This is not a niche prediction from a gastroenterologist with a server in a cupboard. Microsoft's 2025 Work Trend Index surveyed 31,000 workers and coined a term for what every employee is about to become: the agent boss. Not a boss of people. A boss of AI agents: assigning them tasks, reviewing their output, deciding what gets delegated at all. Leaders in that survey expected their teams to be training and managing agents within five years. Jensen Huang put it even more bluntly at CES: the IT department of every company is going to become the HR department of AI agents. Onboarding them. Training them on the local vocabulary. Keeping them in line.
Swap two words and Huang has described a teaching hospital. Onboarding, orientation to local practice, graded responsibility, performance review. We have been the HR department of housestaff for a hundred years.
And medicine's fleet is already forming
The agents are not hypothetical anymore. This year Nature published DeepRare, a multi-agent system for rare disease diagnosis that coordinates more than 40 tools, generates a ranked differential, and shows its reasoning chain - expert reviewers agreed with 95.4% of those chains. Science published Biomni, a general-purpose biomedical research agent that composes its own workflows across 25 domains and runs them. A Mayo Clinic scoping review of agentic AI in healthcare is equally instructive for the opposite reason: across five databases they found only seven eligible studies, exactly one of which involved actual patients. The capability is arriving far ahead of the evidence, which frankly is the usual order of operations in our field.
Meanwhile NEJM Catalyst is already arguing about how to PAY for this, and their framing is the one that matters: AI that will "deliver clinician-grade care under the direction of a clinician." Read that phrase again. Under the direction of a clinician. That is a job description, and the job is management.
So what does managing actually mean?
I have been thinking about what I actually DO with this fleet, hour to hour. It is not coding. I have not written a meaningful line of code in this whole project. Drucker spent a career arguing that management is a practice you learn, not a personality you have, and that its core job is making other people's work productive. That description fits my evenings better than anything from computer science does.
When I strip it down, I do three things.
I brief. I tell an agent what I want, what the constraints are, and what finished looks like. The quality of what comes back depends almost entirely on how well I posed the task. This should have been familiar territory. It wasn't, at first. I have spent twenty years receiving vague consults ("rule out badness") and resenting them, but nobody ever sat me down and taught me to write a precise one. I learned briefing the way I learned everything managerial in medicine: by being on the receiving end of it done poorly.
I check. Not everything. That would defeat the purpose, and honestly there isn't time. An attending doesn't repeat every physical exam on rounds; she picks the ones that matter and takes the rest on trust. I now make that same choice dozens of times a day. Which agent outputs do I read closely, which do I skim, which do I hand to Agent 1b to attack before I ever see them. Some days I check less than I should. Other days I waste an hour re-verifying something that was fine. Getting that calibration right is most of the job, and nobody taught me this either.
I decide how much rope. Medicine has a beautiful word for this: entrustment. We built a whole framework for deciding when a trainee can do a task with less supervision, and the wisdom in it is that trust is never a global verdict. It is task by task. Full autonomy on the paracentesis, direct observation on the family meeting. I run the same judgement on my agents now. Number One can send a calendar hold on its own; it cannot send an email in my name. The rest of the economy is slowly rediscovering this idea. We have had it for years.
Briefing, checking, calibrating trust. That is what management means, at least the version of it I am living. And every piece of it came from the wards, not from the manual.
So teach it. On purpose.
Here is my actual point. Everything above, I learned by osmosis, over two decades, mostly by watching people who were good at it and occasionally being burned by people who were not. Osmosis is not a curriculum. It worked, sort of, when the only people we managed were other humans who could push back, fill in our gaps, and tell us when our instructions made no sense. Agents will not do that. They will cheerfully execute a bad brief at scale. The cost of unexamined management skill is about to go up.
So I think we should teach this overtly, and I mean in residency, not in some optional leadership elective in the last year of fellowship. A few concrete versions of what that could look like. Teach briefing the way we eventually taught handover: we ignored that skill for decades, then the safety data embarrassed us into structuring it, and now every trainee learns a format. Delegation deserves the same treatment. Teach checking as its own competency: how to sample another's work, what to verify personally, how to give feedback that changes the next draft rather than just grading the last one. And make entrustment of AI explicit rather than instinctive: if a resident is going to let an agent draft her discharge summaries, I want her able to say out loud what she has entrusted, on what evidence, and what she still verifies every single time. That is an assessable skill! We assess far softer things already.
And frankly the trainees might be the easy part. It is the mid-career attendings, me included, who are managing fleets right now with no training at all. Faculty development has to catch up to this, and the honest starting point is admitting that supervision was always management, and that we have been credentialing it on vibes.
The catch: managing what you never did
I will take one objection deeply, because it is the real one. Every good manager in medicine was once the intern. The attending can smell a bad overnight plan because she wrote a hundred of them badly first - I have written about this. But if the agents do the intern work, where does the next generation earn the judgement to supervise the agents? A group from Duke-NUS and Harvard writing in Nature Medicine calls the failure mode never-skilling: trainees who lean on AI in the formative years and never build the foundational reasoning that safe oversight requires. The automation literature has been warning us about this since before I went to medical school - Bainbridge's ironies of automation, newly urgent in the agentic era. The more the system does, the harder the leftover human work gets, because the leftover work is judgement. Management of agents cannot be the FIRST job. It has to be earned the way attending-hood is earned, on a gradient, with the manager still able to do enough of the work to judge it. IMO this is the central design problem for residency in the next decade, and almost nobody is working on it.
Back to the cupboard
So this is what the last couple of weeks have been teaching me. The hierarchy, the red-teamers, the feedback agent: none of it felt like computer science. It felt like running a clinical teaching unit. Medicine did not prepare me to write code. It absolutely prepared me to manage.
The cupboard under the stairs, it turns out, is a teaching hospital.
This essay first appeared on Dr. Grover's Substack.