The Signal FilesManaging the MachinesResearch Brief

Stop treating your agents like employees

The friendlier the label, the softer the check. New research finds that calling an agent a colleague makes its manager watch it less, and catch less.

~1,900 WordsFive Cited SourcesStop Trying To Be Invisible

The new hire has a name, a friendly avatar, and a line on the org chart. It never sleeps, it drafts faster than anyone on the floor, and this quarter it got a warmer welcome than most humans do. In a new study of 1,261 managers, 23 percent already work at organizations that list AI agents on the org chart like this. The same study carries an uncomfortable finding about what the welcome does: the moment identical work is labeled as coming from an "AI employee" rather than an "AI tool," the manager reviewing it starts checking it less, and catching less.

Harvard Business Review, M. Kropp, J. Bedard, E. Wiles, M. Hsu & L. Krayer, "Research: Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026.

This is a research brief on that one study: what it measured, where its effect actually lives, and what we think the honest conclusion is. We run this studio on an agent workforce, so the question is not academic for us. The industry's default answer to "how do we roll out agents?" is an HR ritual: a name, an onboarding deck, a seat at the standup. The evidence now on the table says the ritual has a price, and the price is paid in oversight.

Section OneThe experiment: one word, two labels

The research team (Matthew Kropp and Julie Bedard of BCG, Emma Wiles and Megan Hsu of Boston University) surveyed and ran experiments on 1,261 managers, HR and finance directors and executives, with a separate YouGov robustness survey of 1,500 more. The design was simple in the way good designs are: show managers identical drafts of work, and vary only the label on where the work came from. Sometimes the draft was presented as produced by an "AI tool." Sometimes by an "AI employee." Sometimes by a human employee. Same words on the page. Same errors planted in the text.

Wiles, E., Hsu, M., Bedard, J. & Kropp, M., "Putting AI on the Org Chart: Evidence on Delegation and Oversight," working paper, July 17, 2026.

One result is reassuringly human: managers exercised the most direct oversight when told the work came from a human employee. People check people. The question the study was built to answer is what happens to that instinct when the worker is a machine wearing a lanyard.

Across the full sample, the average effects on error catching were small. The authors say so plainly, and so will we. The finding that matters lives in a subgroup: the managers whose organizations had already institutionalized "AI employees." Those organizations had already put them on the org chart, already made the metaphor official. That was 23 percent of the primary sample (14 percent in the YouGov sample). For them, the label changed behavior.

Section TwoWhat the label did

Within that institutionalized subgroup, presenting identical drafts as the work of an "AI employee" rather than an "AI tool" reduced managers' monitoring intensity by 16 percent. Their review performance was 16 percent worse. They caught 18 percent fewer of the planted errors. And they leaned harder on someone else doing the checking: reliance on additional review rose by 22 percentage points, an increase of roughly 44 percent.

Wiles et al., "Putting AI on the Org Chart," working paper, July 17, 2026. Subgroup: managers at organizations that already have "AI employees."

Read the design again to feel the weight of that. Nothing about the work changed. Nothing about the errors changed. The only treatment was a noun. A word was enough to soften the human check by a sixth and let nearly a fifth more errors through, in exactly the organizations that had committed hardest to the metaphor.

The friendlier the label, the softer the check.

Why it mattersThe subgroup is not a footnote; it is a preview. The effect switches on precisely where the "AI employee" fiction has been institutionalized. Institutionalizing it is what the market is currently rushing to do, org-chart software, AI "headcount" and all. The 23 percent in this sample are simply early. The study's warning is aimed at everyone about to join them.

Section ThreeWhere the responsibility went

The second finding explains the first. When the identical work carried the "AI employee" label, managers in the institutionalized subgroup assigned about 9 percentage points less accountability to themselves, and about 8 percentage points more to "the AI system." The responsibility did not shrink. It moved. It migrated from the one party who can be accountable to the one party that cannot.

Wiles et al., "Putting AI on the Org Chart," working paper, July 17, 2026.

This is the quiet mechanics of anthropomorphism. Call something a colleague and you import the whole social contract that comes with colleagues: colleagues own their work, colleagues are answerable for their mistakes, colleagues would be insulted by line-by-line inspection. Every one of those imports is false for a language model. But the manager's behavior updates as if they were true. The framing hands the machine a kind of social credit it has done nothing to earn, and hands the manager an exit from the checking that the machine, of all workers, still needs most.

Responsibility did not disappear. It migrated to the one party that cannot carry it.

Section FourWhat this study is not

Now the honest counterweight, stated in the text and not in a footnote. This is one study from one team, published in two venues (the HBR article and the underlying working paper) with no independent replication yet. The average effects on error catching across the whole sample were small; the sharp numbers above are subgroup effects. A subgroup result from a single team is a warning shot, not settled law. We cite it as exactly that.

And the employee framing is not uniformly poison. Among managers whose organizations had not institutionalized AI employees, the "AI employee" label actually increased their stated comfort with the work. So the metaphor buys something real: adoption. People delegate more willingly to something that feels like a colleague. The study's structure suggests the uncomfortable trade hiding inside that comfort: the framing that eases delegation on day one is the same framing that, once it hardens into the org chart, correlates with a softer check. Comfort now, oversight later. That is a loan, and loans get collected.

The honest readTreat the average effect as the fact and the subgroup as the hypothesis worth acting on. If the effect replicates, the cost of the "AI employee" fiction lands precisely where adoption is deepest. That is where the stakes are highest. Waiting for replication before you tighten your review discipline is a bet we would not take with our own error rate.

Section FiveTreat them like systems, because they are

If the HR metaphor degrades oversight, what is the alternative? Not hostility: specification. An agent does not need what employees need. It needs what systems need: explicit, machine-parseable procedures instead of culture; a scoped mandate (what it may touch, what "done" means, what proof is required) instead of trust; logs a human can audit instead of a performance review; an escalation route instead of an open-door policy. Every one of those is a legibility artifact: something written down clearly enough that a machine can execute it and a manager can verify it.

This is not our invention. The engineering literature on running agents at scale reached the same conclusion from the opposite direction. OpenAI's account of a five-month agent-first build attributes its reliability to scaffolding (maps instead of manuals, executable feedback, repository knowledge as the system of record), never to treating the agent as a person. Anthropic's follow-up work on harness design goes further: the fix for an agent that grades its own work too kindly was structural separation of doer and checker: a generator building, a skeptical evaluator verifying. Structure the agent cannot route around, in place of trust it has not earned.

OpenAI, R. Lopopolo, "Harness engineering: leveraging Codex in an agent-first world," February 11, 2026; Anthropic, P. Rajasekaran, "Harness design for long-running application development," March 24, 2026.

Notice what both schools quietly assume: the agent is a system to be engineered, not a hire to be welcomed. The rollout that works looks like systems engineering (bounded scope, definition of done, verification gates), not like onboarding week. The BCG and Boston University data now supplies the missing half of the argument: the HR ritual is not merely unnecessary. In the organizations that embraced it hardest, it measurably weakened the one safeguard that matters, which is a human actually looking.

An agent doesn't need a welcome party. It needs a mandate it can parse.

We hold ourselves to this. The agents that research, draft, and ship inside this studio have no names, no avatars, and no seats on any chart. They have scoped briefs, hard guardrails written in plain files, and a founder who reviews the work as work, with the full monitoring intensity the "AI tool" label apparently preserves. The study gave us a number for something we had only felt: the moment you start thinking of the agent as a colleague, you stop reading its output like an editor.

In ClosingLegibility, inside and out

Everything we publish returns to one claim: machines can only act on what they can clearly read. We usually make that argument about the outside of a company: the model recommends the brand the web has made legible, and an invisible brand is simply not in the answer. This study makes the same argument about the inside. An agent workforce runs on legibility twice over: the agent needs procedures it can parse, and the humans need a frame that keeps their scrutiny switched on. Anthropomorphism fails both at once: it blurs the mandate and softens the check. The org chart is for people. Give the machines something better: a specification.

If you want to know how legible your own brand already is to the machines that now answer your customers' questions, you can measure it at /signal-index/ or write to us at /contact/.

Figure 01 · One Word, Two Labels
What "employee" did to the human check
−16%
Monitoring intensity. Identical drafts, identical errors. Only the label changed from "AI tool" to "AI employee." Subgroup: organizations that already institutionalize AI employees.
+22pp
Reliance on additional review. Roughly a 44% increase, in the same institutionalized-AI-employee subgroup. The manager checks less and hopes someone downstream checks more.
The check softens and gets outsourced. Source: Wiles, Hsu, Bedard & Kropp, "Putting AI on the Org Chart," working paper, 2026. A single, not yet independently replicated study.
Stop trying to be invisible.

Sources

  1. Harvard Business Review, M. Kropp, J. Bedard, E. Wiles, M. Hsu & L. Krayer, "Research: Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026. hbr.org/2026/05/research-why-you-shouldnt-treat-ai-agents-like-employees
  2. BCG (republication of the HBR article), "Why You Shouldn't Treat AI Agents Like Employees," May 6, 2026. bcg.com/news/6may2026-why-you-shouldnt-treat-ai-agents-employees
  3. Wiles, E., Hsu, M., Bedard, J. & Kropp, M., "Putting AI on the Org Chart: Evidence on Delegation and Oversight," working paper, July 17, 2026. emmawiles.com/storage/ai_employee.pdf
  4. OpenAI, R. Lopopolo, "Harness engineering: leveraging Codex in an agent-first world," February 11, 2026. openai.com/index/harness-engineering
  5. Anthropic, P. Rajasekaran, "Harness design for long-running application development," March 24, 2026. anthropic.com/engineering/harness-design-long-running-apps

The Signal Index

How clearly can the AI era see you?

A free, transparent score of how AI and search find, understand and recommend you. Instant, from your domain.

Get your Signal Index →

The Signal Files

Field notes on visibility, in your inbox.

The research behind how brands get seen now. The Signal Files, the moment they publish. No noise.

Double opt-in. Unsubscribe anytime.