Skip to content
All guides
AI Litmus7 min read10 August 2026

Workslop, shadow AI, vibe coding: every AI buzzword is a people problem

By Shobhit Khandelwal·Founder, VMS Culture Labs
Workslop, shadow AI, vibe coding: every AI buzzword is a people problem
The short answer

Every term the AI conversation keeps producing describes a human behaviour, not a model capability. But measure AI fluency is where most advice stops being useful, because fluency is not one number. It is at least five distinct capabilities, and two people with an identical average can need opposite interventions. The more valuable read is not the score at all: it is why more AI use is stuck, which is one of about seven barriers, only one of which training fixes.

None of these words describe a model

Workslop, coined in Harvard Business Review, for AI output that looks finished and quietly pushes the real work onto whoever receives it. Vibe coding, Andrej Karpathy's term for building by accepting what the model suggests without really reading it. The jagged frontier, from the field experiment with BCG consultants, where AI lifted performance on one task and degraded it on a task that looked identical. Shadow AI, where a study of nearly 48,000 workers across 47 countries found more than half hide their AI use from managers.

There is no benchmark for workslop. No model card reports a shadow AI score. Every one of those words describes a decision a person made. Your competitor has the same models you do, from the same handful of vendors, on the same terms, so the model is not the variable. The behaviour is.

How AI buzzwords sort into two questionsWorkslop, vibe coding and the jagged frontier are capability signals, answered by AI Litmus. Shadow AI, AI fatigue and AI-first backlash are consequence signals, answered by Behavioral Intelligence.WorkslopVibe codingJagged frontierShadow AIAI fatigueAI-first backlashCAPABILITYCan this role get trustworthywork out of the tool?Answered by AI LitmusCONSEQUENCEWhat is the behaviour aroundAI costing you, in money?Answered by Behavioral Intelligence
Every term sorts into a capability signal or a consequence signal. Each has a different owner and a different first move.

Measure AI fluency is where most advice stops being useful

Everyone now agrees you should measure fluency. Almost nobody says what fluency is, so it collapses back into one number on a dashboard. That number is the problem, because fluency is not one capability. It is at least five, and they fail independently.

  • Prompting and direction: can they describe a task well enough that the tool can do it.
  • Critical thinking and judgment: do they verify, and can they recognise a confident wrong answer.
  • AI and tool literacy: which tools, how deeply, and do they know where capability ends.
  • Workflow integration: is AI in the daily work as a repeatable step, or one clever thing they did once.
  • Growth and team influence: do they teach it forward, or is their skill trapped in one person.

The first three are close to the four capabilities in Anthropic's AI Fluency framework, and the ladder underneath them is aligned with the UNESCO AI Competency Framework and the US Department of Labor's AI literacy work. We did not invent this vocabulary. We made it measurable per role.

Now look at what an average does to it. Two people, same overall figure, nothing else in common:

Two people with the same average fluency score and opposite problemsOne person scores high on prompting and low on judgment, so they ship fast and ship wrong. The other scores high on judgment and lower on prompting and workflow, so their work is trustworthy but does not compound. Both average 3.6, and each needs a completely different intervention.Ships fast, ships wrongStrong prompting, weak judgmentPrompting5Judgment2Tool literacy4Workflow4Influence3AVERAGE 3.6Trustworthy, not compoundingStrong judgment, weak workflowPrompting3Judgment5Tool literacy3Workflow3Influence4AVERAGE 3.6
Identical averages, opposite problems. One is a workslop factory. The other is trustworthy but nothing they do compounds. A single score reports them as the same person.

The person on the left is strong at prompting and weak at judgment. They produce a lot, fast, and some of it is wrong in ways that only surface downstream. That is not a training-needs finding, it is a risk finding, and it is exactly the profile that a usage dashboard rewards. The person on the right verifies everything and is safe to trust, but their good work lives in their own head and never becomes a repeatable step anyone else can use.

The average is not a summary of those two people. It is the deletion of the only thing you needed to know about them.

The same applies to the ladder. Absent, experimental, functional, integrated, multiplier are not grades, they are five different states with five different next moves. The one worth counting is multiplier: the people who build things other people then use. Most companies cannot name theirs, which means they cannot promote them, protect them, or copy what they did.

The score is not the useful part. The reason is.

Here is the finding that changed how we built the product. When you ask people properly, the thing blocking more AI use is almost never a mystery, and it is usually not skill. It is one of about seven things, and each one has a completely different fix.

Seven reasons AI use stalls, and the fix for eachThe barriers are access, skill, trust, policy, time, relevance and fear. Training is the correct response only to the skill barrier. Trust and fear are behavioural and carry a financial cost.AccessProvision the licence. No training needed.SkillTrain. This is the only row training fixes.TrustGovernance and verified examples in their own work.PolicySay out loud what is actually allowed.TimeRedesign the workflow. They are not idle.RelevanceRole-specific use cases, not general demos.FearA leadership problem with a rupee cost.GREEN = TRAINING FIXES IT · ORANGE = BEHAVIOURAL, PRICED BY BEHAVIORAL INTELLIGENCE
The barrier decides the intervention. Blanket AI training spends the same budget on all seven rows and correctly addresses one of them.

Read that as a budget document. If the real barrier in your operations team is access, training is theatre and a licence is the whole fix. If it is policy, people are not unskilled, they are waiting for permission that nobody has given in writing. If it is time, they are not idle, the workflow is wrong and no course changes that. If it is relevance, they sat through a demo about a job that is not theirs and concluded AI is not for them.

This is why we resist selling a score. A score tells a leader how they are doing. A barrier tells them what to do on Monday, and it is the difference between a training budget and a targeted one.

Three questions about tools, which only work asked together

Alongside the fluency read, there are three things worth knowing about every person's toolset, and their value comes from being asked in the same conversation:

  • What they actually use for real work, in their words, not what procurement says they have.
  • What they have access to and do not use. You are already paying for this, and it is the cheapest recovery available to you.
  • What they wish they had. This is the only honest buy signal in the building, because it comes from the person doing the work rather than the vendor's roadmap.

Set those against what the company actually provisions and the answers stop being opinions and become decisions. Paid and unadopted is a deployment problem, not a purchase. Wanted and absent is the buy case. Used but used badly is the train case. Most AI budgets get argued without any of the three on the table.

Two more reads come from the same conversation and matter more than the score. Whether the person is resistant, willing or eager, because a willing team with an access barrier is the fastest win available to you and a resistant team is a communication problem you have not had yet. And the governance read: whether they paste confidential material into public tools, and whether anything gets verified before it ships. We record that as a behaviour pattern and never the content itself.

One design decision underneath all of it: this is a conversation, not a test, and no real name ever reaches the model. A test measures test-taking, and people manage their answers when they think a score is going to their manager. The honesty is the product. Everything above is worthless if people perform for it.

Two of those barriers are not capability problems at all

Look back at the barrier figure and notice the two orange rows. Trust and fear are not fixed by any curriculum. They are what happens when a rollout lands badly: people who believe the tool is auditioning for their job, people who stopped disclosing which tools they use because the first company-wide message sounded like a warning, managers absorbing a verification burden nobody costed, teams told to be AI-first who heard be fewer.

That is not an AI measurement problem. It is a behavioural one with a price attached, and it is what Behavioral Intelligence exists to put a number on, in the unit finance already argues in. It is also why we will not sell the fluency half on its own. Measure capability alone and you get a technically fluent, quietly resentful company. Measure the behaviour alone and you understand your culture beautifully while losing on execution.

There will be another buzzword next quarter. Sort it into a capability signal or a consequence signal and you will already know which of these two things it belongs to, and who owns the fix.

The words will keep changing. The two questions underneath them will not.

See this on your own teams.

A private walkthrough, calibrated to your roles. About two weeks.

Frequently asked

What is workslop?

Workslop is a term coined in Harvard Business Review for AI-generated work that looks polished but lacks substance, so the effort of fixing it transfers to whoever receives it. It is invisible to usage dashboards, because someone producing workslop registers as highly active. In fluency terms it is a specific profile: strong prompting paired with weak critical judgment.

Why is a single AI fluency score not enough?

Because fluency is at least five capabilities that fail independently: prompting and direction, critical thinking and judgment, AI and tool literacy, workflow integration, and growth and team influence. Two people can share an identical average while one ships fast and wrong and the other is trustworthy but never scales what they know. They need opposite interventions, and the average deletes the difference. A useful read is per dimension and per role.

What actually stops teams from using AI more?

In practice it is one of about seven barriers: access, skill, trust, policy, time, relevance or fear. Only the skill barrier is fixed by training. Access is a provisioning fix, policy is a clarity fix, time is a workflow fix, relevance needs role-specific use cases, and trust and fear are behavioural problems that carry a financial cost. Identifying the barrier per team is what turns a blanket training budget into a targeted one.

Shobhit Khandelwal
Shobhit Khandelwal
Founder, VMS Culture Labs

Shobhit Khandelwal is the founder of VMS Culture Labs, on a mission to measure what most leaders only guess at: how fluently their teams truly work with AI, and the hidden cost of how people behave at work. He is out to replace workplace guesswork with evidence, and build the kind of workplaces the next generation deserves.

Connect on LinkedIn
Get the intelligence briefing

Occasional, high-signal notes on measuring AI fluency and culture cost. No spam.

Keep reading