Stop treating your agents like employees
The friendlier the label, the softer the check. New research finds that calling an agent a colleague makes its manager watch it less, and catch less.
The new hire has a name, a friendly avatar, and a line on the org chart. It never sleeps, it drafts faster than anyone on the floor, and this quarter it got a warmer welcome than most humans do. In a new study of 1,261 managers, 23 percent already work at organizations that list AI agents on the org chart like this. The same study carries an uncomfortable finding about what the welcome does: the moment identical work is labeled as coming from an "AI employee" rather than an "AI tool," the manager reviewing it starts checking it less, and catching less.
Harvard Business Review, M. Kropp, J. Bedard, E. Wiles, M. Hsu & L. Krayer, "Research: Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026.
This is a research brief on that one study: what it measured, where its effect actually lives, and what we think the honest conclusion is. We run this studio on an agent workforce, so the question is not academic for us. The industry's default answer to "how do we roll out agents?" is an HR ritual: a name, an onboarding deck, a seat at the standup. The evidence now on the table says the ritual has a price, and the price is paid in oversight.
Section OneThe experiment: one word, two labels
The research team (Matthew Kropp and Julie Bedard of BCG, Emma Wiles and Megan Hsu of Boston University) surveyed and ran experiments on 1,261 managers, HR and finance directors and executives, with a separate YouGov robustness survey of 1,500 more. The design was simple in the way good designs are: show managers identical drafts of work, and vary only the label on where the work came from. Sometimes the draft was presented as produced by an "AI tool." Sometimes by an "AI employee." Sometimes by a human employee. Same words on the page. Same errors planted in the text.
Wiles, E., Hsu, M., Bedard, J. & Kropp, M., "Putting AI on the Org Chart: Evidence on Delegation and Oversight," working paper, July 17, 2026.
One result is reassuringly human: managers exercised the most direct oversight when told the work came from a human employee. People check people. The question the study was built to answer is what happens to that instinct when the worker is a machine wearing a lanyard.
Across the full sample, the average effects on error catching were small. The authors say so plainly, and so will we. The finding that matters lives in a subgroup: the managers whose organizations had already institutionalized "AI employees." Those organizations had already put them on the org chart, already made the metaphor official. That was 23 percent of the primary sample (14 percent in the YouGov sample). For them, the label changed behavior.
Section TwoWhat the label did
Within that institutionalized subgroup, presenting identical drafts as the work of an "AI employee" rather than an "AI tool" reduced managers' monitoring intensity by 16 percent. Their review performance was 16 percent worse. They caught 18 percent fewer of the planted errors. And they leaned harder on someone else doing the checking: reliance on additional review rose by 22 percentage points, an increase of roughly 44 percent.
Wiles et al., "Putting AI on the Org Chart," working paper, July 17, 2026. Subgroup: managers at organizations that already have "AI employees."
Read the design again to feel the weight of that. Nothing about the work changed. Nothing about the errors changed. The only treatment was a noun. A word was enough to soften the human check by a sixth and let nearly a fifth more errors through, in exactly the organizations that had committed hardest to the metaphor.
The friendlier the label, the softer the check.
Section ThreeWhere the responsibility went
The second finding explains the first. When the identical work carried the "AI employee" label, managers in the institutionalized subgroup assigned about 9 percentage points less accountability to themselves, and about 8 percentage points more to "the AI system." The responsibility did not shrink. It moved. It migrated from the one party who can be accountable to the one party that cannot.
Wiles et al., "Putting AI on the Org Chart," working paper, July 17, 2026.
This is the quiet mechanics of anthropomorphism. Call something a colleague and you import the whole social contract that comes with colleagues: colleagues own their work, colleagues are answerable for their mistakes, colleagues would be insulted by line-by-line inspection. Every one of those imports is false for a language model. But the manager's behavior updates as if they were true. The framing hands the machine a kind of social credit it has done nothing to earn, and hands the manager an exit from the checking that the machine, of all workers, still needs most.
Responsibility did not disappear. It migrated to the one party that cannot carry it.
Section FourWhat this study is not
Now the honest counterweight, stated in the text and not in a footnote. This is one study from one team, published in two venues (the HBR article and the underlying working paper) with no independent replication yet. The average effects on error catching across the whole sample were small; the sharp numbers above are subgroup effects. A subgroup result from a single team is a warning shot, not settled law. We cite it as exactly that.
And the employee framing is not uniformly poison. Among managers whose organizations had not institutionalized AI employees, the "AI employee" label actually increased their stated comfort with the work. So the metaphor buys something real: adoption. People delegate more willingly to something that feels like a colleague. The study's structure suggests the uncomfortable trade hiding inside that comfort: the framing that eases delegation on day one is the same framing that, once it hardens into the org chart, correlates with a softer check. Comfort now, oversight later. That is a loan, and loans get collected.
Section FiveTreat them like systems, because they are
If the HR metaphor degrades oversight, what is the alternative? Not hostility: specification. An agent does not need what employees need. It needs what systems need: explicit, machine-parseable procedures instead of culture; a scoped mandate (what it may touch, what "done" means, what proof is required) instead of trust; logs a human can audit instead of a performance review; an escalation route instead of an open-door policy. Every one of those is a legibility artifact: something written down clearly enough that a machine can execute it and a manager can verify it.
This is not our invention. The engineering literature on running agents at scale reached the same conclusion from the opposite direction. OpenAI's account of a five-month agent-first build attributes its reliability to scaffolding (maps instead of manuals, executable feedback, repository knowledge as the system of record), never to treating the agent as a person. Anthropic's follow-up work on harness design goes further: the fix for an agent that grades its own work too kindly was structural separation of doer and checker: a generator building, a skeptical evaluator verifying. Structure the agent cannot route around, in place of trust it has not earned.
OpenAI, R. Lopopolo, "Harness engineering: leveraging Codex in an agent-first world," February 11, 2026; Anthropic, P. Rajasekaran, "Harness design for long-running application development," March 24, 2026.
Notice what both schools quietly assume: the agent is a system to be engineered, not a hire to be welcomed. The rollout that works looks like systems engineering (bounded scope, definition of done, verification gates), not like onboarding week. The BCG and Boston University data now supplies the missing half of the argument: the HR ritual is not merely unnecessary. In the organizations that embraced it hardest, it measurably weakened the one safeguard that matters, which is a human actually looking.
An agent doesn't need a welcome party. It needs a mandate it can parse.
We hold ourselves to this. The agents that research, draft, and ship inside this studio have no names, no avatars, and no seats on any chart. They have scoped briefs, hard guardrails written in plain files, and a founder who reviews the work as work, with the full monitoring intensity the "AI tool" label apparently preserves. The study gave us a number for something we had only felt: the moment you start thinking of the agent as a colleague, you stop reading its output like an editor.
In ClosingLegibility, inside and out
Everything we publish returns to one claim: machines can only act on what they can clearly read. We usually make that argument about the outside of a company: the model recommends the brand the web has made legible, and an invisible brand is simply not in the answer. This study makes the same argument about the inside. An agent workforce runs on legibility twice over: the agent needs procedures it can parse, and the humans need a frame that keeps their scrutiny switched on. Anthropomorphism fails both at once: it blurs the mandate and softens the check. The org chart is for people. Give the machines something better: a specification.
If you want to know how legible your own brand already is to the machines that now answer your customers' questions, you can measure it at /signal-index/ or write to us at /contact/.
Sources
- Harvard Business Review, M. Kropp, J. Bedard, E. Wiles, M. Hsu & L. Krayer, "Research: Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026. hbr.org/2026/05/research-why-you-shouldnt-treat-ai-agents-like-employees
- BCG (republication of the HBR article), "Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026. bcg.com/news/6may2026-why-you-shouldnt-treat-ai-agents-employees
- Wiles, E., Hsu, M., Bedard, J. & Kropp, M., "Putting AI on the Org Chart: Evidence on Delegation and Oversight," working paper, July 17, 2026. emmawiles.com/storage/ai_employee.pdf
- OpenAI, R. Lopopolo, "Harness engineering: leveraging Codex in an agent-first world," February 11, 2026. openai.com/index/harness-engineering
- Anthropic, P. Rajasekaran, "Harness design for long-running application development," March 24, 2026. anthropic.com/engineering/harness-design-long-running-apps
The Signal Index
How clearly can the AI era see you?
A free, transparent score of how AI and search find, understand and recommend you. Instant, from your domain.
Get your Signal Index →