ResourcesAnalysis · AUGMENTED WORK

    What does it mean to augment work with AI?

    A practical definition of AI augmentation, grounded in research on task fit, human oversight, job design, and measurable outcomes.

    Prepared work modules arranged around a central human decision point

    In a study of 5,172 customer-support agents, access to a generative AI assistant increased the number of issues resolved per hour by 15 percent on average. Less experienced agents gained the most. The strongest agents saw much smaller improvements.

    That result is often presented as evidence that AI boosts productivity. The more useful lesson is narrower. The system helped in a particular workflow, with a particular knowledge base, for a particular group of workers. Its value depended on where it entered the work and what the people using it already knew.

    This is the practical meaning of augmentation. It describes how work is divided between people and machines, and whether that division produces a better result.

    A working definition

    AI augments work when it changes how a task or workflow is divided between people and machines in a way that improves a relevant outcome while preserving clear responsibility for judgment and consequences.

    The relevant unit is a task or workflow. A job usually contains many different activities. Some can be prepared by AI, some can be automated under defined conditions, and some depend on judgment that remains difficult to delegate.

    The outcome also needs to be explicit. Speed, quality, learning, customer satisfaction, consistency, and reduced coordination effort are different benefits. A system can improve one while damaging another.

    Responsibility cannot disappear into the interface. Someone needs to understand what the system did, where it can fail, and who carries the consequence of acting on its output.

    Improvement depends on task fit

    The uneven capability of AI systems is one reason broad productivity claims are unreliable.

    Fabrizio Dell’Acqua and colleagues studied 758 consultants completing realistic knowledge-work tasks. Participants using AI performed better on tasks that fell within the system’s capabilities. On a task outside that frontier, AI assistance made performance worse.

    The researchers called this a jagged technological frontier. The boundary between strong and weak performance does not follow an obvious line. Two tasks that look similar to a user can sit on different sides of it.

    This changes how teams should evaluate AI. Choosing a model is only one part of the design. Teams also need to decide which tasks enter the system, which context it receives, what output it produces, and how uncertainty becomes visible.

    The customer-support study offers another view of task fit. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond found that the gains were concentrated among less experienced and lower-skilled agents. The assistant appeared to help spread practices associated with stronger performers.

    A human reviewer does not guarantee a better system

    Many AI products describe themselves as human-in-the-loop. The phrase says little about how the loop works.

    Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone reviewed 106 experiments reporting 370 effect sizes. The combined systems performed better than humans alone on average. They also performed worse than whichever of the human or AI performed best alone.

    Collaboration has a design cost. People can accept weak recommendations, override strong ones, misunderstand confidence, or spend more time checking an output than the task would have taken directly.

    Review is useful when the reviewer has the information, authority, and attention needed to make a meaningful judgment. It becomes ceremonial when the interface encourages rapid approval or hides the reasons behind an output.

    This distinction matters for proactive systems. If an AI identifies work before someone asks, it can reduce the burden of noticing and reconstructing context. It can also create a stream of suggestions that someone must verify. Net value depends on precision, prioritization, and review cost.

    Work should be examined below the job title

    Discussions about AI and employment often treat a job as one unit. Daily work is more granular.

    The International Labour Organization and Poland’s National Research Institute developed a task-level index of occupational exposure to generative AI. Their 2025 analysis concluded that job transformation was more likely than full automation because most occupations still include tasks requiring human involvement.

    Consider a client renewal. The work may include monitoring the thread, locating a commitment from a meeting, checking a contract date, preparing a reply, deciding the commercial position, sending the message, and tracking the response. These activities belong to one business outcome, yet they have different requirements.

    An AI system might detect that the reply is overdue, retrieve the relevant history, and prepare a draft. A person may still decide whether the proposed terms are acceptable. A configured workflow might schedule an internal reminder without another approval. Each step has its own context, risk, and control.

    Five questions for evaluating augmentation

    1. What task or signal is being handled?

      Describe the actual work. “Help with email” is too broad. “Identify client threads awaiting a reply for more than three working days” can be evaluated.

    2. What does the AI prepare, recommend, or execute?

      Preparation, recommendation, and execution create different risks. A draft that waits for review differs from an external message sent by a configured workflow.

    3. Which judgment and consequence remain with a person?

      Name the person or role. Define what they need to know and what authority they have.

    4. How will a failure be detected and corrected?

      Failures include missed work, false alerts, incorrect context, weak drafts, and inappropriate actions.

    5. What evidence would show that the workflow is better?

      Choose an outcome before rollout and include the time required for review and correction.

    What remains unresolved

    Most studies observe individual tasks over a limited period. They tell us less about long-term skill development, dependency, team coordination, or how organizations should redesign roles. Results tied to a model or interface can also age quickly.

    There is even less direct evidence about proactive AI. Studies of chat assistants and tools used inside an existing task do not establish that systems which notice and prepare work before a prompt will improve follow-through. That claim needs its own measurement.

    AUGMTD develops proactive AI for workplace follow-through. Our interest in augmentation comes from building systems that connect work context, identify signals, and prepare useful next moves. That product thesis does not turn research on other systems into evidence about AUGMTD.

    The useful question for any team is concrete: does this allocation of work produce a better and more accountable result in the setting where it will actually run?

    PRIMARY SOURCES

    Research referenced

    1. Brynjolfsson, Li, and Raymond, “Generative AI at Work,” Quarterly Journal of Economics, 2025
    2. Dell’Acqua and colleagues, “Navigating the Jagged Technological Frontier,” Organization Science, 2026
    3. Vaccaro, Almaatouq, and Malone, “When combinations of humans and AI are useful,” Nature Human Behaviour, 2024
    4. International Labour Organization and NASK, “Generative AI and Jobs,” 2025
    Continue exploringHow proactive AI notices work before a prompt
    Product guide