AI Automation

AI and Skill Decay: How to Keep Your Team Sharp

Automation gives a team hours back, and it also removes the repetition that kept judgment in use. The 2026 evidence on skill decay is early and mostly short-term, but one finding is worth planning around: how AI is introduced shapes what people retain.

A well-built automation hands a team hours back every week. Invoices are coded, enquiries are triaged, and first drafts appear before anyone has opened the file. For a mid-size business that is real capacity. The question that rarely comes up at approval is what the routine used to do for the people doing it. Reconciling, drafting and triaging were not only output. They were repetition, and repetition is how people keep a feel for what normal looks like and when something is off.

This article looks at what the 2026 evidence says about AI and skill decay, which is thinner than most headlines suggest, and then at how a mid-size European business can keep human judgment in use. It is general information, not HR advice. The frame is resilience. An AI-resilient team is not one that avoids AI. It is one whose work is designed so that the skills needed to supervise AI stay in regular use.

Three problems that often get called skill decay

The phrase gets used for several different things, and mixing them leads to the wrong fix.

ProblemWhat it meansWhat the evidence looks likeTypical design response
Skill decayUnaided performance falls after a period of AI-assisted workShort experiments and one clinical observationPlanned practice without AI
Automation biasAccepting AI output without enough scrutinyDocumented in some contexts, according to the International AI Safety Report 2026Verification steps built into the workflow
Cognitive fatigueMental strain from excessive use or constant monitoring of AI toolsOne US survey based on self-reportLimits on tool count and oversight load

Practice addresses decay, workflow design addresses automation bias, and workload design addresses fatigue. A team can have one without the others.

What the evidence says, and where it is thin

Most of this research is recent and short-term, so the table gives each study's type, main finding and main caveat.

SourceType and sampleMain findingMain caveat
Budzyń and colleagues, Lancet Gastroenterology and Hepatology, August 2025Observational study at four endoscopy centres in Poland, 19 experienced endoscopists, 1,443 procedures done without AIShare of unassisted procedures that found at least one adenoma fell from 28.4% to 22.4% in the three months after AI was introduced, a drop of six percentage pointsNot randomised, so other changes during the study may have contributed. One AI system, experienced clinicians only
Bastani and colleagues, PNAS, 2025Field experiment with nearly 1,000 high school mathematics studentsThe group with a standard chat-style version of GPT-4 scored 17% lower once access was removed than students who had no access. A tutor version with safeguards, GPT Tutor, largely avoided the dropStudents learning, not employees working. One subject and one model generation
Liu and colleagues, arXiv preprint, April 2026Three randomised online experiments, 1,222 participants recruited and 1,060 analysed, fraction problems and reading comprehensionAfter a short period of AI help, people did worse once it was removed without warning, and skipped more problems, which the authors use as a measure of persistenceA preprint we could not confirm as peer reviewed. Very short exposure, laboratory-style tasks, and the skipping difference was not statistically significant in one experiment
Lee and colleagues, Microsoft Research and Carnegie Mellon University, CHI 2025Survey of 319 knowledge workers, who shared 936 first-hand examples of using generative AI at workHigher confidence in generative AI went with less critical thinking, and higher self-confidence went with moreSelf-reported perception, not measured skill, and an association, not causation. Microsoft sells AI assistants, so treat it as vendor-affiliated research
International AI Safety Report, February 2026Synthesis by over 100 independent experts. Its advisory panel was nominated by over 30 countries and international organisations, and does not endorse the contentAI use can affect cognitive skills such as critical thinking, and in some contexts people accept AI suggestions without checking them, shown in a randomised experiment with 2,784 participantsA synthesis, not new measurement. The report says the lack of long-term evidence makes persistent changes hard to identify
BCG, June 2026Survey of 70 senior executives worldwide plus interviews with about a dozen company leaders. BCG is a consultancy that sells related servicesHalf of respondents said they already observe de-skilling, and over 60% expected it to become a material threat within three to five yearsPerceptions of a small sample, not measured skill. We found no sampling description that supports generalising it

Primary sources: Budzyń et al., Lancet Gastroenterology and Hepatology, Bastani et al., PNAS, Liu et al., arXiv, Lee et al., Microsoft Research, International AI Safety Report 2026 and BCG, June 2026.

Read together, the studies support one narrow claim: after assisted work, performance without the assistance can be lower, and how much seems to depend on how the assistance is designed. In the school experiment, the standard chat version of GPT-4 was followed by lower scores once access was removed, while the version prompted to support learning largely avoided the drop. That is the most useful lesson for a business, because design is the part it controls.

What the studies do not show is that AI causes lasting skill loss in working teams. The samples are students, online participants and clinicians, and exposures last from minutes to a few months. The one workplace-style study is a self-reported survey. We did not find a long-term study of skill change in European business teams.

A note on the colonoscopy figure

You may see the clinical result quoted as detection being six percent lower. The Lancet paper reports a fall from 28.4% to 22.4% in the share of unassisted procedures that found at least one adenoma. That is six percentage points, or about 20% in relative terms, and the two readings are not interchangeable. The International AI Safety Report quotes the result as about 6 percentage points.

The study was observational. Its authors acknowledge that factors other than AI may have influenced the result, commentators pointed to organisational changes and higher procedure volumes during the period, and the authors call for randomised trials. The study also sits beside a body of trial evidence that the same kind of tool can raise detection while it is in use. A meta-analysis of 44 randomised trials, cited in the journal's accompanying commentary, suggested an absolute gain of about eight points. The question is not whether the tool is useful, but what happens to the people when it is not there.

What the surveys add

The Microsoft Research survey adds a pattern worth noting. Workers with more confidence in their own skills reported more critical thinking, and those with more confidence in the AI reported less. It is self-reported and associational, so it cannot show cause. One cautious reading is that skills are part of what keeps checking in place, which makes them part of what makes AI safe to use, not a competitor to it.

The BCG executive survey is a different kind of evidence. It shows that senior leaders are worried and most have not yet acted: BCG reports that about one in ten companies has an organisation-wide response and about a third have not discussed the issue explicitly. It does not show that skills have fallen, and with 70 respondents it is a signal, not a market measurement.

For Europe: what the data can and cannot say

Eurostat's survey of enterprises with ten or more employees found that 20.0% of EU enterprises used at least one AI technology in 2025, up from 13.5% in 2024, a rise of 6.5 percentage points. By size class, the shares were 17.0% for small enterprises, 30.4% for medium enterprises and 55.0% for large enterprises. The survey covers a range of AI technologies, not only chat assistants, so it measures exposure, not skill change.

We did not find EU-wide, firm-level data on skill decay specifically. The one European study in the evidence above is clinical. The colonoscopy research was run in Poland and tells us about endoscopists, not accountants or customer service teams. Medium-sized firms are now at roughly three in ten for AI use, so exposure is no longer confined to large enterprises, while the evidence on skills is still being built. That is a reason to design deliberately now, while it is cheap, not a reason for alarm.

For context: the US picture

The most-cited workplace study of AI strain comes from the US. Researchers from Boston Consulting Group, writing in Harvard Business Review in March 2026, surveyed 1,488 US workers at large companies. They define AI brain fry as mental fatigue from excessive use or oversight of AI tools beyond one's cognitive capacity. Participants described mental fog, slower decision-making and headaches.

Two points matter. First, the study concerns fatigue, not lost skill. Second, it rests on a survey, and the authors tie the strain to excessive use or constant monitoring of AI tools, with more errors, decision overload and intention to leave among the costs reported. US workplace practices differ from those in most European settings, so we treat this as an early signal for workload design, not a forecast for your team.

Why resilience is the right frame

If skills are what makes AI output safe to use, protecting them is part of the business case for automation, not a brake on it. Someone has to notice when a classification is wrong, when a draft misreads the client, or when an approved exception should have been refused. That ability comes from practice. As AI agents take on multi-step work, a person makes fewer small decisions along the way, which makes oversight skills valuable and easy to neglect. We cover how agents work in What Is Agentic AI?.

The distinction between routine execution and judgment is the same one we draw in Can AI Replace an Employee?. Automate the execution, and keep judgment in regular use. Roles change as a result, which we discuss in How AI Restructures Teams, Not Just Headcount. Beginners are a special case, because they have less judgment to protect and more to build, and we cover that in Junior Roles and AI: How to Keep Your Talent Pipeline Alive. This article is about the rest of the team, who can drift into approving what the system suggests because it is fluent and fast.

A framework for keeping judgment in use

This is a Kubera AI planning heuristic, not a validated method. It draws on the evidence above, which is early, and the specific choices have not been tested in controlled trials, so we would treat any implementation as a pilot to measure.

The Kubera Skill Resilience Map has four steps:

StepQuestion to askWhat to doExample output
1. Find the judgment skillsWhich skills decide quality when the AI is wrong?For each role, name the decisions where a wrong call is costly, such as exceptions, approvals and client adviceA one-page list of skills to protect for each role
2. Set the human-first pointsWhere should a person form a view before the AI weighs in?Label tasks as human first, AI first with review, or AI off, and keep the AI-off list shortA task map with three labels
3. Build practice into the workflowHow will these skills be used on purpose?Rotate unaided tasks, run error-spotting exercises on training items with known errors, run occasional failure drills, and require a short written reason before an approvalA monthly practice calendar and a decision log
4. Measure and adjustHow will we know the skills are holding?Track observable signals such as known errors caught in practice items, time to independent sign-off, and the quality of written reasons, and review them each quarterA quarterly review of three signals

A safeguard on error-spotting exercises: place known errors only in a controlled training queue or a safe test process that the team knows about. Do not insert errors covertly into real credit decisions, medical conclusions or client documents. A wrong item that reaches a customer or patient is real harm, and testing staff without their knowledge raises employment and data protection questions that need legal advice.

Step 2 reflects the one design lesson that appears across the experiments: when people formed a view before or alongside the AI, the damage looked smaller than when they leaned on it as a crutch. We would not read that as established for workplaces, which is why Step 4 exists.

Where automation helps build resilience

A well-built workflow can create the structures the map needs, instead of adding them as extra chores. It can route routine volume to AI and send exceptions to a person who must write a short reason before approving, so the decision log builds itself as a by-product of the work. It can run a separate practice queue of anonymised past items with known errors, clearly marked as training, so reviewers stay sharp and the team gets a measurable catch rate without touching live decisions. It can rotate who handles a task unaided each week, and cap how many AI tools or agents one person oversees at once, which speaks to the risk of oversight overload.

Starting from clear processes makes all of this easier, which is the point of AI Automation Doesn't Fix Chaos — It Scales It: How to Know If Your Business Is Actually Ready. The practice calendar and the three signals fit naturally into a phased plan like the one in How to Build an AI Roadmap for a Small Business, and they can sit beside the financial measures in How to Measure ROI from AI Automation, so that skill signals are reviewed with the same discipline as savings.

What the workflow can look like

This is a design pattern we would propose in a project, not a description of a delivered client system.

StageWhat happensWhat it produces
1. Standard tasksAI processes routine items that meet defined rulesCompleted items with a confidence flag
2. Complex casesItems outside the rules go to an employee, who reviews them before anything is sent or approvedA named reviewer for each exception
3. Decision logThe employee records the decision and a short reasonA searchable record of decisions and reasons
4. Quality and skill reportA monthly report on exception volume, overrides, reason quality and practice resultsInput for the quarterly review in Step 4

What other organisations report doing

BCG's June 2026 article describes practices at several organisations. These are company practices as BCG reports them, not independently measured results. At Shell, junior staff reportedly frame the problem and produce a baseline analysis before using AI to refine it, and BCG reports early results of better question quality and faster progress to independent work, though it publishes no figures. At CNIL, the French data protection authority, managers reportedly assess employees' ability to challenge AI outputs, not just to use them. An Indian bank reportedly runs an AI-free session on the first Friday of each month, and Salesforce reportedly pairs confident AI users with less confident colleagues.

None of this shows these practices work. It shows organisations treating the question as a design problem.

Where this plays out in practice

Illustrative scenario, not a specific Kubera client: a mid-size European distributor has a small finance team that approves customer credit limits. After automation, AI prepares each assessment and an analyst approves it. Within months, approvals are faster, written reasons are one line long, and a system outage leaves two analysts unsure how to assess an unusual customer by hand.

The firm keeps the automation and applies the map. It lists the decisions where a wrong call is costly, such as large limits for new customers. It sets those as human first, so the analyst records a view before opening the AI draft. It runs a weekly practice pack of anonymised past assessments containing one known error, kept apart from live approvals and announced to the team as training. It also rotates one unaided assessment per analyst each month, and requires a reason of at least two sentences for any limit above a set threshold. It reviews three signals each quarter: known errors caught in the practice pack, time to independent sign-off for newer analysts, and the quality of written reasons. It treats this as a pilot reviewed after two quarters. This is design intent, not a measured result.

FAQ

Does AI make people worse at their jobs? The evidence is early and mixed. Some experiments and one clinical observation show lower performance without AI after assisted work, but exposures were short and tasks narrow. A practical first check is to list which decisions in your team now start from an AI draft, since those are where unaided practice has dropped most.

How should the colonoscopy figure be read? Unassisted detection fell from 28.4% to 22.4%, which is six percentage points, or about 20% in relative terms, not a six percent relative drop. The study was observational, so it cannot show that AI caused the fall.

Is this a reason to avoid AI? No. The same tools raised detection in many trials while in use, and the school experiment found that design changed the outcome. A sensible first step is one mapping session per team to decide which decisions stay human first.

What is automation bias, and how is it different from skill decay? Automation bias is the tendency to accept AI output without enough scrutiny. Skill decay is lower unaided performance after a period of assisted work. A common warning sign is a review queue where almost every item is approved unchanged and written reasons keep getting shorter.

What is AI brain fry, and is it the same thing? It is a term from a 2026 survey of 1,488 US workers for mental fatigue from excessive use or oversight of AI tools. It concerns strain, not lost skill, and it comes from a survey, so we treat it as a prompt to check how many tools and agents one person oversees.

Do these findings apply to European mid-size firms? As an early signal, yes. As a measurement of your team, no. We found no EU-wide firm-level data on skill change, and the one European study here is clinical.

Which skills are most at risk? BCG's executives rated judgment and problem framing as most exposed, but that is a perception from a small sample. A safer assumption is that any skill that stops being practised is at risk, which is why Step 1 asks you to name your own.

What can we do without slowing the automation down? Build the safeguards into the workflow. Require a short written reason before approvals, run practice items with known errors in a separate training queue, rotate unaided tasks, and keep the AI-off list short. The time cost depends on volume and design, so measure it during the pilot instead of assuming it is small.

How do we know whether it is working? Track known errors caught in practice items, time to independent sign-off, and the quality of written reasons. Review them each quarter and treat the first two quarters as a pilot.

Can we test staff by hiding errors in real work? No. Use a training queue or test process that the team knows about, built from anonymised past cases. Hidden errors in live credit, medical or client work can cause real harm, and covert testing damages trust.

How does this connect to an automation project? Directly. The same workflow that automates routine volume can route exceptions to a person, log reasons and run practice items. We raise skill resilience during design, not after launch, because automation changes what each role practises day to day. Rolling it out alongside training is covered in How to Onboard Your Team to AI Automation.

If you would like to work out which routine work in your business should move to AI, which judgment should stay in regular use, and how to build that into the workflow from the start, that is design work worth doing before the automation goes live.

Discuss your automation project →

Back to blog