Enterprise AI Readiness & Competency Diagnostic

Before you spend another dollar on AI.

Which of your workflows are actually worth changing, can your organization support the change, and can your people execute it? Three questions, in that order. Most assessments skip the first and get the third wrong.

The problem

Most enterprises are not failing at AI because the technology doesn't work.

They are failing because they bought tools before they knew what problem they were solving, had no baseline to measure against, and assumed that giving people access to a chatbot was the same as building capability.

88%

of organizations report regular AI use in at least one business function — but only about a third have begun scaling it.

McKinsey, State of AI Global Survey, n=1,993, 2025
39%

report any EBIT impact at all. Of those, most attribute less than 5% of total EBIT to AI. Roughly 6% qualify as high performers.

McKinsey, same survey
30–55%

of purchased AI seats reach sustained weekly active use in year one — pushing effective cost per active user to 1.8–3× list price.

Redress Compliance, ~25–35 enterprise estates, 2024–2025
>40%

of agentic AI projects are projected to be cancelled by end of 2027 — escalating costs, unclear value, inadequate risk controls.

Gartner, June 2025

The organizations capturing value are not the ones with the best tools. They are the ones that changed how work is done. High performers are roughly three times more likely to fundamentally redesign workflows, three times more likely to have senior leadership demonstrating real ownership, and at least three times more likely to scale AI agents across functions.

A note on the number you have probably seen

We do not lead with "95% of AI pilots fail."

You have likely seen that claim, from MIT's GenAI Divide report. We do not use it, because its methodology does not support it. The "zero return" finding rested on 52 interviews the report itself described as directionally accurate rather than measured, and it defined success narrowly as measurable KPI-backed deployment within six months.

The honest version is less dramatic and more useful: most organizations are getting something from AI, almost none can prove how much, and the ones capturing real value redesigned their work rather than buying more licences.

That gap — between spend and provable return — is what this diagnostic measures.

What it produces

Two scores that are never combined into one.

A strong infrastructure score does not compensate for a weak workforce, and the reverse is equally true. They answer different questions and lead to opposite action plans.

Organizational readiness

Can the business support AI in production? Data, integration, ownership, process, governance, and the ability to measure whether anything actually worked. Seven weighted domains.

Human capability

Can the people actually do the work? Four domains, scored by role group — and measured by observation, not by survey. Tool depth and judgment are scored independently.

Why we refuse to blend them. A single number lets an organization with excellent infrastructure and an untrained workforce land in the same tier as one with capable people and broken data. Those two need opposite interventions. Combining the scores destroys the only diagnostic information the assessment produced. The two are read together in a posture matrix that produces one of four recommendations — not an average.

The recommendation

Four postures. Four different plans.

Scale

Readiness strong · capability strong

Both foundations hold. Move to workflow redesign, integrated deployment, and measured autonomy in defined areas. The constraint is ambition, not capability.

Enable

Readiness strong · capability weak

Infrastructure is ahead of the people. More tooling will not help and will increase waste. Capability build first, with measured before and after. Expect current licence utilisation to be poor.

Unblock

Readiness weak · capability strong

The people are ahead of the plumbing. This is the shadow-AI generator: capable staff routing around a stack that fails them. Fix data access and integration urgently — this posture is unstable and the capability will leave.

Sequence

Readiness weak · capability weak

Stop discretionary AI spend. Baseline three processes, remediate data, build foundational literacy — in that order. Anything else is expensive theatre.

How it works

You do not score yourself.

Scores come from four independent evidence streams, weighted and reconciled by the assessor. Every score is anchored to a named artifact, interview, observed task, or telemetry figure. An unevidenced score is downgraded automatically.

Stream 1

Artifact review

We request ten specific documents — data catalogue, integration diagram, 90-day licence utilisation, AI policy, rollout post-mortems, process documentation, cost baselines, vendor data-processing terms, the named AI owner, and prior business cases with their actual results. Whether you can produce them is itself a scored signal.

Stream 2

Structured interviews

Role-specific scripts across executive, technical, operational, security, and finance stakeholders. Participant selection is not delegated entirely to the sponsor; we require a defined proportion of assessor-selected participants, including frontline staff in the candidate workflows.

Stream 3

Practical capability assessment

A sampled cohort completes a 30-minute observed work task from their own role. We watch what they do. Self-reported AI competency is systematically inflated, and no survey detects the most dangerous profile in an organization — the confident user who does not verify.

Stream 4

Telemetry

Where available we pull actual usage data: seat activity, query volumes, feature engagement, spend against budget. Numbers, not narrative.

Operating principles

Eight non-negotiables.

No happy ears

The engagement surfaces uncomfortable findings. A diagnostic that validates the sponsor's existing plan was not worth commissioning.

Evidence beats enthusiasm

Every score is anchored to a named artifact, interview, or observed task. "Leadership feels we're a 4 on data quality" is not a score.

Measure the before, or don't claim the after

If a process has no baseline — hours, error rate, cycle time, cost — no efficiency claim about it can ever be verified. Baselining is the first work, not the last.

Tool access is not capability

A licence is an input. We measure what people actually produce and whether they can tell when the output is wrong.

The bottleneck is rarely the model

It is almost always data access, permissions, undocumented process, or the absence of anyone accountable for the outcome.

Redesign beats automation

Pointing AI at an existing broken process makes the process faster and still broken. The value is in restructuring the work.

Shadow AI is a symptom, not a crime

When capable people route around the sanctioned stack, the stack is the problem. We measure the behaviour without punishing the people, because it is the most honest signal in the building.

The right answer is sometimes "don't"

Some workflows should not be automated. Naming those protects the credibility of everything else we recommend.

What you get

Ten leadership-ready deliverables.

DeliverableWhat it does
Executive readout memoBoard-level summary, posture, and the single recommended move
Organizational readiness scorecardSeven domains, weighted and scored, with a per-domain evidence log
Human capability scorecardFour domains scored by role group, including practical assessment results
Gate reportThe hard blockers, ranked, with what it takes to clear each one
Shadow AI exposure reportWhat is actually being used, by whom, with what data
Value opportunity registerTen candidate workflows, sized and ranked by provable value
Top-three business casesThe three best candidates, costed against a measured baseline
Spend efficiency reviewWhat current AI spend is returning, and what to cut
90-day action planNamed moves, named owners, dates, and gates
Buy / build / wait / stop decision briefThe leadership call, with the reasoning behind it

The commercial argument

In one paragraph.

A company of 150 people does not have a CIO's office, a strategy function, or an internal team whose job is to check whether the AI plan holds. It has a leadership team with day jobs, a vendor with a compelling deck, and a decision worth several hundred thousand dollars in licences, integration work, and the headcount that gets hired or not hired around it.

The diagnostic costs a fraction of that decision. It exists to tell you, before you commit, whether the work is worth changing, whether the business can support the change, and whether your people can execute it — and to tell you plainly if the answer is not yet.

Baseline it. Prove it's worth doing. Then build the capability to do it.

What we need from you

These are engagement conditions, not preferences.

The diagnostic does not work without access. If access is restricted, we decline the engagement — a diagnostic conducted through a filtered view produces a filtered answer, and both parties waste the money.

Sponsor time

An executive sponsor available for two sessions and reachable throughout.

The artifacts

The ten documents on the request list — or a written acknowledgement of what does not exist. "We don't have it" is a valid and highly informative answer.

Interview access

Across executive, technical, operational, security, and finance roles, including assessor-selected frontline participants.

Assessment slots

Thirty-minute practical sessions for a sampled cohort, scheduled and attended.

Telemetry

Read access to licence and usage data for the AI tools already in use.

Unedited findings

Agreement that findings are delivered to the sponsor unedited.

Next step

The disciplined step between AI enthusiasm and AI spend that returns something.

Tell us what you have already bought and what you expected it to do. We will tell you whether a diagnostic would help, and what it would take.

Email contact@shockpoint.io and tell us what decision you are trying to make.