The jagged frontier: why your most confident AI users fail silently

The jagged frontier is the uneven, invisible line between tasks AI does well and tasks it quietly does badly. Your most confident users are often the ones who cannot see the line, so they trust the tool exactly where they should check it. Counting licences or course completions hides this. The fix is to measure how people actually work with AI, place each person on a fluency map, and price the gap in rupees. That is what turns AI training from a hopeful cost into a number finance can approve.
What is the jagged frontier?
In a now well-known field experiment, researchers from BCG and Harvard Business School (Dell'Acqua et al., later published in Organization Science) gave hundreds of consultants access to a frontier AI model. On tasks that sat inside the model's capability, the AI users were faster and produced markedly higher-quality work. On a task deliberately designed to sit just outside it, AI users were more likely to be wrong, because the model handed them a confident, polished answer that was subtly off.
The researchers called this the jagged frontier. AI is excellent on one side of an invisible line and unreliable on the other, and the line is uneven: it moves from task to task, and it is not marked. Nobody gets a warning that they have crossed it.
Why your best people are the most exposed
The intuition is that heavy users are your safest users. With the jagged frontier, the opposite can be true.
- Confidence outruns calibration. Someone who uses AI all day trusts it more, so they verify it less, which is exactly the wrong instinct near the edge of what it can do.
- The failures are silent. A wrong answer that looks right does not announce itself. It gets pasted into the deck, sent to the client, or shipped in the code, and the cost surfaces later, somewhere else, detached from its cause.
- Speed disguises it. AI makes the wrong answer arrive faster and cleaner than a human would produce it, and polish reads as competence.
This is why a headcount of people using ChatGPT tells you almost nothing about whether your team uses it well. The person generating the most output can be the one quietly shipping the most subtly wrong work.
Why the usual metrics miss it
Most learning dashboards measure inputs, and every input metric is blind to the frontier.
- Licences bought measure access, not skill. A seat is not a habit.
- Course completions measure attendance, not whether anyone changed how they work the following Monday.
- Self-report surveys measure confidence, and confidence is the very thing the jagged frontier turns against you.
The training itself is not the missing piece. Around 81% of Indian GCCs already run internal generative-AI training programmes (EY GCC Pulse Survey), and EY upskilled 44,000 of its own people before taking an AI academy to the market. Volume is not the gap. Knowing who has actually crossed the frontier and who is stuck below it is the gap.
How to see the frontier in your own team
You cannot fix an invisible line by running more generic training at it. You make it visible by measuring behaviour instead of opinions. Three steps.
- Score real fluency, not a quiz. A short conversational diagnostic reveals how a person actually works with AI: what they use it for, where they stop, whether they check the output, and whether they can tell a good answer from a plausible wrong one.
- Place everyone on a five-level map, from absent to multiplier. Now the jagged frontier stops being an abstract idea and becomes a picture of your specific team, with the exposed people visible.
- Price the gap in rupees. Map each level to the hours AI can realistically save and a sober productivity lift, and moving people up the map becomes an investment decision rather than a hope.
Discernment, not prompting, is the skill that separates a multiplier from a confident guesser.
What this means for your AI budget
India is expected to need roughly a million more AI-skilled professionals in 2026 (Nasscom), and AI-enablement roles are among the fastest-growing categories on Indian job boards. The spend is coming whether or not anyone can prove it works. The only question your CFO will ask, a quarter later, is whether it did.
A team that can name who sits where on the frontier, and what closing that gap is worth, walks into that review with an answer instead of an anecdote. That is exactly what our AI Litmus diagnostic builds in two weeks: a team maturity heatmap, the top three gaps holding you back, and a board-ready rupee ROI figure. We are taking 10 founding pilots this month at a founding rate. If proving AI ROI is landing on your desk this quarter, that is the conversation to have now.
See this on your own teams.
A private walkthrough, calibrated to your roles. About two weeks.
Frequently asked
What is the jagged frontier in AI?
It is the uneven, invisible line between tasks a frontier AI model does well and tasks it does badly. BCG and Harvard researchers coined the term after finding that AI made consultants better on tasks inside the line and worse on a task just outside it, without the users being able to tell the difference.
Why do confident AI users make more mistakes?
Because confidence leads them to verify less, and the jagged frontier's failures are the ones that look right. Near the edge of what AI can do, the model returns a polished, plausible answer that is subtly wrong. A heavy user who trusts the tool is less likely to catch it than a cautious one.
How do you measure whether a team actually uses AI well?
By scoring behaviour rather than opinions. A structured conversational diagnostic looks at how someone works with AI, whether they verify output, and how they handle tasks near the edge of AI's capability. Scoring is consistent across people, so you can place a whole team on one maturity map and see where the risk sits.
Does more AI training fix the jagged frontier?
Not on its own. Generic training adds volume, not discernment. You first have to measure where each person actually sits, then target training at the specific gap, then re-measure to prove the change. Otherwise you are spending against a problem you cannot see.

Shobhit Khandelwal is the founder of VMS Culture Labs, on a mission to measure what most leaders only guess at: how fluently their teams truly work with AI, and the hidden cost of how people behave at work. He is out to replace workplace guesswork with evidence, and build the kind of workplaces the next generation deserves.
Connect on LinkedInOccasional, high-signal notes on measuring AI fluency and culture cost. No spam.