The Augmented Team Manifesto: Seven Principles for Human-AI Teams
The Augmented Team Manifesto: Seven Principles for Human-AI Teams
A declaration about the unit of work that replaces the individual contributor — and the design decisions that make it work.
Adding AI agents to a team does not make the team better. Measured across 106 experiments and 370 effect sizes, human-AI combinations performed significantly worse than the best of either humans or AI working alone (Hedges' g = −0.23, 95% CI −0.39 to −0.07). That is the finding of a preregistered meta-analysis published in Nature Human Behaviour on 28 October 2024. Not "AI underperformed." The combination underperformed.
Be precise about what that does and does not say. Against humans working alone, the combinations won clearly (g = 0.64). Adding AI to a person helps. What fails is the assumption that adding AI to a team produces the best of both — on average it produces something worse than whichever side was already stronger.
Hold that next to the other number everyone quotes. Despite $30-40 billion in enterprise investment, 95% of generative AI pilots produced no measurable P&L impact, according to MIT's Project NANDA in July 2025. That study is preliminary and not peer-reviewed, and it has been criticized for a short measurement window. The direction of the finding has not been seriously contested.
Two findings, one conclusion: the constraint is not model capability. The constraint is team design.
For a century we built organizations on a single assumption — that individual performance is the unit of organizational capability. Hire great individuals. Measure individual contributions. Reward individual achievement. Build career paths for individual advancement. That assumption was always incomplete, but it was close enough to reality and convenient enough for management that it persisted.
It no longer holds. When a capable collaborator can be added to a team in minutes, the interesting question stops being who did we hire and becomes how is this team composed. The relevant unit is the augmented team: humans and AI agents working as one unit, with combined capability that neither side reaches alone.
That combined capability is not automatic. The meta-analysis is the evidence. Augmented teams that outperform are designed. The rest are just staffed.
These are the seven principles we design against. Argue with them. But if you are putting agents next to people this year, you need a position on all seven, because the failure modes are already documented.
Principle 1: Collective Intelligence Over Individual Brilliance
The old model: Hire the smartest individuals. Give them resources. Get out of their way. Organizational capability is the sum of individual capabilities.
The augmented team model: Design for collective intelligence. Organizational capability emerges from how intelligence — human and machine — combines. Composition beats accumulation.
Why this matters
The Nature Human Behaviour meta-analysis found something more useful than a headline. The losses were concentrated in decision-making tasks. The gains showed up in content creation tasks. And the direction depended entirely on who was better alone: when humans outperformed AI, the combination gained; when AI outperformed humans, the combination lost.
Read that again, because it is the whole design problem. Teams lose when the stronger party defers to the weaker one. A human overriding a more accurate model destroys value. An agent's output waved through by a reviewer who cannot evaluate it destroys value. Both are composition failures, not capability failures.
The smartest people in your building cannot match a well-designed augmented team. They will comfortably beat a badly designed one.
What to do about it
Assess how candidates work with agents, not just how they perform alone. The relevant skill is no longer solo output — it is knowing when to override a machine and when to defer to it. That is a judgment skill, and it is testable in an interview.
Build teams with intentional capability combinations. Ask which side of a given task is genuinely stronger, and design the workflow so the stronger side leads and the weaker side checks. Never the reverse.
Measure team outcomes, not individual throughput. When you reward individual output inside an augmented team, you are paying people to take credit for work the team did — and paying them to hide the agent that helped.
Principle 2: Fluid Roles Over Fixed Positions
The old model: Define job descriptions. Assign individuals to roles. Evaluate against role expectations. Structure is static.
The augmented team model: Roles are allocated dynamically. Humans and agents take on responsibilities based on current capability and current need. Composition adapts continuously.
Why this matters
Fixed roles were a rational response to expensive onboarding. You hired someone for a role, they learned it over months, they performed it for years. The investment in role-specific knowledge justified role stability.
Agents break that arithmetic. Onboarding to a new responsibility takes minutes. And leaders are already planning around it: in Microsoft's 2025 Work Trend Index — 31,000 knowledge workers across 31 countries — 82% of leaders said they are confident they will use digital labor to expand workforce capacity in the next 12 to 18 months, and the same share called 2025 a pivotal year to rethink key aspects of strategy and operations.
Capacity that can be reallocated in minutes should not be locked into a job description written last year. In Supanova, atoms progress through eleven capability tiers, L1 through L11, banded into contributors, leaders, and executives. A tier is not decoration. It sets what a team member has earned the right to do without you watching.
What to do about it
Replace rigid job descriptions with capability profiles: what a team member can do, not what they must do. Match capability to need each cycle instead of assuming last quarter's allocation still fits.
Give coordination to a named owner. A Chief of Staff atom coordinating operations, a Head of Agent Ops atom proposing new capacity — with a cost estimate and a skill-gap justification — for a person to approve or decline. Fluid roles need firmer coordination, not less.
Invest in agent progression the way you invest in human progression. Track it. Review it. Promote against demonstrated results. If your hiring plan is a confession of the capability you lack — and it usually is — your tier ladder is the record of the capability you actually built.
Principle 3: Context Abundance Over Information Scarcity
The old model: Protect information. Share on a need-to-know basis. Assume too much information creates confusion or risk.
The augmented team model: Maximize context availability. Performance scales with context quality. Design for abundance, not control.
Why is your best context still locked in someone's head?
Information scarcity was a reasonable accommodation to human cognitive limits. People could only hold so much; overloading them reduced performance. It was also always expensive. Coveo's Workplace Relevance Report found the average employee spends 3.6 hours a day searching for the information needed to do their job.
Agents have a different constraint profile. They do not get overloaded by volume; they get crippled by absence. An agent that does not know your pricing floor, your brand voice, why the last vendor was fired, or which customer is fragile this quarter will produce competent work that is wrong for your company.
So the old information-limiting instinct is now an active tax. Every piece of context withheld from your agents is capability withheld from your team. This is why we built a company-specific memory layer rather than shipping another model wrapper: strategy, decisions, and standards have to be structured, not just stored.
What to do about it
Document the things that never made it into a doc: strategic intent, cultural norms, why past decisions went the way they did, which relationships are load-bearing. Structure it for retrieval, not for reading. Our Question Universe exists for exactly this — it captures the context an agent needs by asking for it systematically instead of hoping someone writes it down.
Default to sharing. Restrict what is genuinely confidential, and for that material write a structured summary that carries the decision-relevant context without exposing the sensitive detail.
Staff context curation as a real function. Fund it, measure it, and give it an owner. Organizations with the richest context infrastructure will have the most capable augmented teams, and it will not be close.
Principle 4: Graduated Autonomy Over Binary Authority
The old model: You either have authority to decide or you do not. Authorization is binary. Escalation is the path when authority is missing.
The augmented team model: Autonomy is a spectrum. Different team members hold different autonomy levels for different decision types, and those levels move with demonstrated capability and risk.
Why does every ambiguous decision end up in your inbox?
Binary authority is expensive and the bill is already quantified. McKinsey surveyed more than 1,200 managers and found that 61% said at least half the time they spend making decisions is ineffective. For a typical Fortune 500 company that is roughly 530,000 days of managers' time a year — about $250 million in wages. Respondents who described decision-making as fast were also nearly twice as likely to call those decisions high quality. Speed and quality were not a trade-off.
Unlimited autonomy is the opposite failure. Decisions made without oversight cause harm that a gate would have caught. Some oversight is genuinely worth its cost.
Graduated autonomy dissolves the tension. Routine decisions execute immediately. High-stakes decisions meet a gate. The same framework governs a new hire and an L3 atom, which is the point — one system, applied consistently, with approval gates and an evidence trail behind the decisions that matter. In our own governance settings, approval thresholds are set as a minimum tier per decision type: tasks, subtasks, and whole projects each have their own bar.
What to do about it
Write down your decision types and the autonomy level attached to each. Most organizations have never made this explicit, which is precisely why every ambiguous decision defaults to escalation.
Put oversight inside the workflow rather than above it. A CFO atom surfacing spend against budget, a Head of Agent Ops atom watching performance and trust, a Chief of Staff atom routing an approval to the person who owns it — governance that runs at the speed of the work is the only governance that survives contact with the work. Note what those atoms do and do not do: they surface the number and route the decision. They do not quietly move the budget. Oversight that acts on its own is not oversight.
Expand autonomy on evidence. Demonstrated success raises the ceiling; failure lowers it. Apply the same rule to people and to agents, and say so out loud.
Principle 5: Continuous Calibration Over Static Evaluation
The old model: Evaluate performance annually or quarterly. Set goals at the start. Assess at the end. Adjust next cycle.
The augmented team model: Calibration is continuous. Signals flow constantly. Adjustment happens inside the cycle, not after it.
Why this matters
The annual review is a workaround for expensive data collection, and it never worked well even on its own terms. Gallup found that only 14% of employees strongly agree the performance reviews they receive inspire them to improve. Separately, Gallup found employees are 3.6 times more likely to strongly agree they are motivated to do outstanding work when their manager gives daily rather than annual feedback.
Now put an agent in that cycle. An agent can run a task hundreds of times between two quarterly reviews. Static evaluation does not just underperform for agents — it is arithmetically absurd. By the time a quarterly review notices a drift in output quality, the drift has shipped a few hundred times.
The information required for continuous calibration is already being generated. Every task carries a rating, an approval, a correction, or a rejection. The only real barrier is whether anyone routes those signals back into the work — which is where the quality question is actually settled, one corrected output at a time.
What to do about it
Instrument the feedback you already collect. Ratings, edits, and rejected drafts are performance data. If they die in a UI and never reach the next run, you are paying for signal and throwing it away.
Use governance atoms to synthesize rather than to report. A Head of Agent Ops atom watching performance across a hundred runs sustains an attention span no human manager can, and surfaces the exception instead of the dashboard. That is also how promotion should work: ours gates on completed tasks, quality score, and accumulated experience, not on a calendar.
Replace the annual review with an ongoing conversation. For humans this is a decade-old recommendation nobody implements. For agents you have no choice — the cycle is too fast to defer.
Principle 6: Purpose Alignment Over Task Assignment
The old model: Break work into tasks. Assign tasks. Confirm completion. Assume completed tasks add up to outcomes.
The augmented team model: Align the team on purpose. Let the team determine the activities. Evaluate against the purpose, not the checklist.
Why this matters
Task decomposition is what coordination costs look like when they are written down — and those costs have eaten the job. Asana's Anatomy of Work research found that 60% of a person's time at work goes to "work about work" rather than skilled work: chasing documents, updating status, relitigating priorities.
An augmented team should attack that 60%, not reproduce it. If your agents need every step specified, you have built a faster way to generate work about work. The coordination did not disappear — it just moved into your prompt.
Purpose also survives change in a way task lists do not. When circumstances shift, a task list is obsolete and a purpose is merely harder. This is the same failure the strategy deck has always had: the strategy was never the problem, the capacity to execute against it was.
What to do about it
Frame work as outcomes with owners: what are we trying to achieve, why does it matter, what does done look like, who decides. Then let the team compose the steps.
Keep a coordination layer that maintains alignment without specifying every move. A Project Manager atom turns a purpose into a plan and holds the thread across a long initiative — which is a different job from relaying status between people who already know it.
Evaluate the purpose. Did we achieve what we set out to achieve? A team that completed every task and moved nothing has failed, and your review process should be able to say so.
Principle 7: Ethical Integration Over Ethical Delegation
The old model: Ethics is a specialized function. Compliance officers and ethics committees own it. Everyone else focuses on performance.
The augmented team model: Ethical reasoning is embedded in the workflow. Both humans and agents participate. Accountability stays human.
Why this matters
Delegated ethics leaves a gap between policy and practice, and that gap is widening. Stanford's 2025 AI Index recorded 233 AI-related incidents in 2024, a 56.4% rise in a single year, and found that among companies "a gap persists between recognizing [responsible AI] risks and taking meaningful action."
Scale is what changes the stakes. An agent decision affects customers at a speed no human-only organization can match, which means a bad judgment does not stay a single bad judgment — it propagates across every interaction until someone notices.
But the same properties cut the other way. Agents apply a standard consistently, which humans under deadline pressure do not. They can flag the edge case rather than quietly resolving it. Consistency is the one thing a hand-off to a compliance committee never delivered.
What to do about it
Build the judgment in, not the rulebook on top. A rule list tells an agent what is forbidden; it does not tell it what matters here. The context work from Principle 3 is what makes ethical reasoning specific to your company instead of generic.
Route the hard call to a person. An agent that surfaces "this one is close, and here is why" is doing the job correctly. Design for that path and make it fast, or people will learn to bypass it.
Keep accountability human and named. Agents can support ethical reasoning. They cannot hold responsibility. Every high-stakes decision needs a person whose name is on it — and an evidence trail showing what they saw when they decided.
The Questions This Raises
"We ran an AI pilot and got nothing. Why would this be different?"
Because a pilot tests capability and these principles test design. The MIT finding — 95% of pilots with no measurable P&L impact — reported that externally sourced solutions succeeded about twice as often as internally built ones, and pinned the failure on the learning gap between tool and organization, not on model quality. That is a composition problem. A tool that nobody redesigned any work around produces exactly what it produced for everyone else: nothing measurable.
"Isn't graduated autonomy just a way to say the AI decides?"
No. It is a way to say which decisions the AI decides, at what threshold, with what evidence, and who owns the outcome. Binary authority is the model that actually hides the decision — everything either sails through unexamined or waits in a queue nobody is accountable for. Write the levels down and you can audit them.
"Does the augmented team mean fewer people?"
It means a different distribution of what people do. The work that shrinks is tactical execution. The work that grows is setting direction, defining standards, and evaluating output — and every one of the seven principles above puts more weight on human judgment, not less. Anyone who tells you an augmented team runs itself has not designed one.
Where This Starts
These principles are not incremental improvements to the individual performance model. They replace it. Our position on where this ends is already on the record — the autonomous organization is inevitable. The augmented team is the unit it gets built out of, and that part is happening now, in the gap between organizations that ran an AI pilot and organizations that redesigned a team.
If you lead an organization: pick one team and one purpose. Write down its decision types and autonomy levels. Give its agents the context you have been protecting. Measure the team, not the individuals. One team, done properly, teaches you more than a company-wide rollout.
If you are an individual contributor: the durable skill is knowing when to override a machine and when to defer to it. The meta-analysis says that judgment is where the value is won or lost. Build it deliberately.
If you build technology: design for teams, not users. Composition, context, graduated autonomy, and evidence trails are infrastructure — and almost nobody is shipping them.
The individual performance myth had a good century. Its replacement is not the AI agent. It is the team you build around one.
Manifesto Summary
Collective Intelligence Over Individual Brilliance — Design the composition so the stronger party leads and the weaker one checks; never the reverse.
Fluid Roles Over Fixed Positions — Allocate responsibility by current capability and current need, and track progression for agents as you would for people.
Context Abundance Over Information Scarcity — Withheld context is withheld capability; structure your company's knowledge for retrieval, not for reading.
Graduated Autonomy Over Binary Authority — Write down decision types and autonomy levels, put oversight inside the workflow, and expand autonomy on evidence.
Continuous Calibration Over Static Evaluation — Route the ratings, corrections, and rejections you already collect back into the next run.
Purpose Alignment Over Task Assignment — Align on outcomes and let the team compose the steps, or you will automate work about work.
Ethical Integration Over Ethical Delegation — Embed judgment in the workflow, route hard calls to a person, and keep accountability named.
Argue with any of the seven. But pick a position on all of them before you put agents next to people.
... Six
Sources
- When combinations of humans and AI are useful: A systematic review and meta-analysis — Nature Human Behaviour (2024)
- Humans and AI: Do they work better together or alone? — MIT Sloan
- The GenAI Divide: State of AI in Business 2025 — MIT Project NANDA (coverage)
- 2025 Work Trend Index Annual Report: The Frontier Firm is born — Microsoft
- Workplace Relevance Report — Coveo
- Report: Employees spend 3.6 hours each day searching for info — VentureBeat
- Three keys to faster, better decisions — McKinsey
- Give Performance Reviews That Actually Inspire Employees — Gallup
- How Effective Feedback Fuels Performance — Gallup
- How Work About Work Gets in the Way of Real Work — Asana Anatomy of Work
- AI Index Report 2025, Responsible AI — Stanford HAI