Two professionals reviewing work together, representing the mentorship and knowledge transfer that builds expert judgment over time
    ·8 min read·Enterprise AI

    AI Is Replacing the Work That Made Your Experts

    Last updated July 18, 2026

    AI is making human judgment the scarcest skill in any organization. It's also automating the entry-level work through which expert judgment is built. BCG estimates 50-55% of US jobs will be reshaped in two to three years. The capability pipeline no one is measuring closes first.

    There's a question many professionals are carrying but not saying out loud. It's not whether AI will replace their job. It's whether they're still getting better at it.

    That's a different kind of anxiety. Quieter, harder to name. And it's the one that actually matters more - for individuals trying to build a career, and for the organizations that depend on those individuals making calls under pressure.

    AI is doing more of the work. Output is faster. Teams look productive. And underneath all of that, something is changing that won't show up on any efficiency dashboard until it does.

    The Paradox Nobody Is Naming

    In May 2026, the World Economic Forum published a jobs and skills roundup with a finding that deserves far more attention than it received.

    As AI handles prediction, pattern-spotting, and routine analysis, "what remains scarce and valuable is human judgement - the ability to weigh trade-offs, navigate ambiguity and decide what 'good' looks like in context." That's WEF's framing. The concept they introduce is "judgement work" - the distinctly human layer that AI prediction cannot replace.

    Here's what the same piece says next, and this is the part that should make every executive uncomfortable: "there is a paradox at the centre of this transition: the same technologies that increase the value of judgement may also erode the pathways through which people develop it."

    AI is making expert judgment more valuable. And it's simultaneously automating the work through which expert judgment is built. That's the trap. And most organizations are walking into it while their efficiency metrics look excellent.

    What Expert Judgment Actually Is - and How It Gets Built

    Before the mechanism, it's worth being specific about what we mean. "Judgment" is the kind of word that sounds important until someone asks you to define it.

    In a customer success context, it's the practitioner who can look at a customer's engagement pattern and know the relationship is drifting before any health score shows it - because they've seen what that drift looks like across fifty accounts, and they've been wrong enough times to know when to trust the signal. In operations and delivery, it's the manager who can feel when a project is going to slip before the dependency map shows a red flag - because they've watched that same early pattern compound into something worse in three engagements before this one. In any leadership context, it's the person who can tell the difference between a team that's stretched but functional and one that's three weeks from breaking - because they've been in both, and the difference is legible to them in ways that don't fit neatly into a status report.

    None of that capability is trained. It's accumulated.

    It's the product of the slow, repetitive, effortful work professionals do early in their careers - the first drafts, the detailed research, the manual analysis, the conversations where they read a situation wrong and someone explains why. The tedious, unglamorous, often frustrating work that teaches you how to think in a domain before you're trusted to make calls in it.

    That is precisely the work AI is now doing.

    The Mechanism of the Problem

    BCG research from March 2026 estimates that between 50 and 55% of US jobs will be reshaped by AI in the next two to three years. Not replaced - reshaped. The task mix shifts: routine and rule-based work moves to AI; humans are asked to do more interpretive, relational, and judgment-intensive work.

    That sounds like good news. It is, for the people who already have the judgment. The problem is who is developing it next.

    I've watched this pattern across CS and operations teams at Zendesk, Adobe, and across enterprise organizations. The practitioners who make your toughest judgment calls today didn't arrive with those instincts. They built them doing work that is now being automated. The CSM who can now evaluate a customer portfolio in thirty minutes built that instinct during eighteen months of manual health-score reviews, logging every interaction, sitting in on calls where they mostly listened and occasionally asked the wrong question. The delivery manager who can smell a timeline problem three weeks out earned that read through a year of watching projects they thought were fine turn out not to be.

    Automate that pipeline, and the person arriving today gets faster output. They do not get the same development.

    The most precise warning I've read on this comes from economist Carl Benedikt Frey, cited in the WEF piece: "if firms respond to AI by hiring fewer junior lawyers and analysts, training less, and assuming the machine will handle the first draft, they erode the very expertise needed to check the machine's output. The organization will look leaner until the hidden error surfaces in public."

    That's the actual risk. Not that AI makes a mistake. That when AI makes a mistake, nobody left has the background to catch it.

    What Organizations That Get This Right Actually Do

    The difference between organizations building a sustainable capability pipeline and those quietly depleting it comes down to a few specific decisions.

    Organizations preserving judgment development:
    - Keep some tasks deliberately human even when AI could do them faster - specifically the tasks that build pattern recognition
    - Build explicit mentorship structures around judgment transfer: what do I notice, how do I evaluate it, where am I uncertain
    - Ask practitioners to produce their own read of a situation before seeing the AI-generated summary - not always, but enough to keep the muscle active
    - Treat "here's where I'd question this output" as a required part of how AI-assisted work gets reviewed

    Organizations optimizing purely for output speed:
    - Automate all entry-level work immediately as AI capabilities improve
    - Measure efficiency gains without measuring the capability pipeline behind them
    - Treat AI as a replacement for the slow work rather than augmentation of the fast work
    - Reduce junior hiring and mentorship investment as AI coverage expands

    The second path looks better on every short-term metric. It produces the capability gap Frey describes - invisible until it isn't.

    The Dashboard Shows Green Until It Doesn't

    I've had CS leaders tell me: *"We're moving faster than ever. I just don't know if we're right."* That sentence has stayed with me longer than most things I've heard in the last two years.

    The efficiency gains are real. In my current operating environment, across customer operations and delivery teams, I've watched account coverage ratios expand substantially when AI is embedded well. That's not in dispute. The capability to manage more with the same team is genuine, and the value is genuine.

    That's the part nobody shows in the case study.

    What doesn't appear on the dashboard: whether the people operating those expanded ratios have the judgment to recognize when something is wrong that the tool missed. Whether the analyst generating AI-assisted reports has enough background to question whether the framing is right. Whether the CSM whose portfolio has doubled has the instinct to feel the renewal that looks fine but isn't. Speed compounds. So does the absence of judgment behind it.

    The early warning signs are quieter than most leaders expect. Not a dramatic AI failure - a team that produces output faster than they can evaluate it. A junior practitioner who gives excellent AI-assisted answers but struggles to explain the reasoning without the prompt. A pattern of decisions that look reasonable individually and strange in aggregate, because no one had the vantage point to see across them. These signals are easy to miss in an environment where the headline metrics are moving in the right direction.

    What the Leadership Layer Needs to Hold

    I'm not sure there's a clean solution to this, and I'm skeptical of frameworks that make it sound like one. But there are decisions that compound the problem versus decisions that slow it.

    For CS, operations, and delivery leaders who are accountable for outcomes that depend on human judgment under pressure, the practical implication is direct: your AI tools are only as reliable as the people checking their output. That checking capability is a function of the judgment those people have developed. Judgment develops through experience with the slow, effortful work. When you automate the slow, effortful work without a deliberate plan for developing judgment in parallel, you're writing a check your future team may not be able to cash.

    This is not an argument against deploying AI. It's an argument for being specific about which parts of the work you protect and why. Not every inefficiency should be optimized away. Some of it is how people become capable. Losing that distinction is the version of this problem that compounds past the point where it's easy to recover from.

    The question worth asking - not just at the executive level, but at the team level, in the operational reviews where these decisions actually get made - isn't "how much can we automate?" It's "what are we automating that we can't rebuild once it's gone?"

    Frequently asked

    If AI makes judgment more valuable, shouldn't organizations naturally want to develop it more?+

    They should, and most will say they do. The gap is between stated priority and structural reality. When entry-level tasks are automated, junior staff no longer get the volume of low-stakes, real-condition decisions that build pattern recognition. Training programs can describe what good judgment looks like - they cannot substitute for the experience of being wrong in real situations and understanding why. Valuing judgment and building it are different investments, and most organizations are only making one of them.

    How do you actually tell whether your team has strong judgment or strong AI-assisted output?+

    Ask them to work through a problem before they see the tool's output. The practitioner with genuine judgment walks you through their reasoning, names where they're uncertain, and identifies the assumptions they'd question. The one dependent on the output can describe it accurately but struggles with the reasoning behind it. The distinction is visible in conversation. It's just rarely the conversation leaders are having before they need to make a call under pressure.

    Didn't previous automation waves raise the same concern - and didn't expertise adapt?+

    The concern was real in each wave and expertise did adapt - but each wave automated at a different level. Calculators changed how arithmetic was done, not how analysis was framed. AI is operating at the reasoning and framing layer, which is different in kind. Frey's specific warning is about this: when AI generates the first draft of analysis, the practitioner doesn't build the mental models that come from constructing it themselves. The historical analogies hold partly, not fully. The degree of exposure depends on how deeply the automated task was connected to how judgment developed.

    Is this primarily a risk for junior professionals, or does it affect experienced people too?+

    Both, through different mechanisms. For junior professionals, the pathway into expertise narrows as entry-level work disappears. For experienced practitioners, the risk is gradual atrophy - a slowly eroding capacity to work without the tool, which usually surfaces as discomfort with problems that don't have an AI-generated starting point. Most won't notice it until they're operating in a context where the tool isn't available, or fails. That's typically too late to diagnose cleanly.

    What's one concrete thing a leader can do right now that actually matters?+

    Before AI summarizes the customer call, the account review, or the project status - have the practitioner tell you what they're seeing. Their read first, tool output second. Do this consistently enough, and you create a practice where judgment is exercised rather than substituted. You also create the data to notice when someone's independent read is getting systematically weaker. It's a small structural decision. Over two years, it makes a large capability difference. ---

    About the author

    Varun Goel
    Varun Goel

    NovaTransform

    Varun Goel has spent his career at the point where enterprise strategy meets the reality of execution - at Adobe, Zendesk, and enterprise operations. He works with business leaders on customer success, digital growth, and operational scale, and writes about the gap between what the playbook says and what actually happens in the room.

    Customer SuccessGTM StrategyAI InnovationDigital TransformationLeadership & ScalingStakeholder Engagement
    View full profile

    More from the blog