Enterprise AI Readiness & Competency Diagnostic
Before you spend another dollar on AI.
Which of your workflows are actually worth changing, can your organization support the change, and can your people execute it? Three questions, in that order. Most assessments skip the first and get the third wrong.
The problem
Most enterprises are not failing at AI because the technology doesn't work.
They are failing because they bought tools before they knew what problem they were solving, had no baseline to measure against, and assumed that giving people access to a chatbot was the same as building capability.
of organizations report regular AI use in at least one business function — but only about a third have begun scaling it.
McKinsey, State of AI Global Survey, n=1,993, 2025report any EBIT impact at all. Of those, most attribute less than 5% of total EBIT to AI. Roughly 6% qualify as high performers.
McKinsey, same surveyof purchased AI seats reach sustained weekly active use in year one — pushing effective cost per active user to 1.8–3× list price.
Redress Compliance, ~25–35 enterprise estates, 2024–2025of agentic AI projects are projected to be cancelled by end of 2027 — escalating costs, unclear value, inadequate risk controls.
Gartner, June 2025The organizations capturing value are not the ones with the best tools. They are the ones that changed how work is done. High performers are roughly three times more likely to fundamentally redesign workflows, three times more likely to have senior leadership demonstrating real ownership, and at least three times more likely to scale AI agents across functions.
A note on the number you have probably seen
We do not lead with "95% of AI pilots fail."
You have likely seen that claim, from MIT's GenAI Divide report. We do not use it, because its methodology does not support it. The "zero return" finding rested on 52 interviews the report itself described as directionally accurate rather than measured, and it defined success narrowly as measurable KPI-backed deployment within six months.
The honest version is less dramatic and more useful: most organizations are getting something from AI, almost none can prove how much, and the ones capturing real value redesigned their work rather than buying more licences.
What it produces
Two scores that are never combined into one.
A strong infrastructure score does not compensate for a weak workforce, and the reverse is equally true. They answer different questions and lead to opposite action plans.
Organizational readiness
Can the business support AI in production? Data, integration, ownership, process, governance, and the ability to measure whether anything actually worked. Seven weighted domains.
Human capability
Can the people actually do the work? Four domains, scored by role group — and measured by observation, not by survey. Tool depth and judgment are scored independently.
Why we refuse to blend them. A single number lets an organization with excellent infrastructure and an untrained workforce land in the same tier as one with capable people and broken data. Those two need opposite interventions. Combining the scores destroys the only diagnostic information the assessment produced. The two are read together in a posture matrix that produces one of four recommendations — not an average.
The recommendation
Four postures. Four different plans.
Scale
Readiness strong · capability strongBoth foundations hold. Move to workflow redesign, integrated deployment, and measured autonomy in defined areas. The constraint is ambition, not capability.
Enable
Readiness strong · capability weakInfrastructure is ahead of the people. More tooling will not help and will increase waste. Capability build first, with measured before and after. Expect current licence utilisation to be poor.
Unblock
Readiness weak · capability strongThe people are ahead of the plumbing. This is the shadow-AI generator: capable staff routing around a stack that fails them. Fix data access and integration urgently — this posture is unstable and the capability will leave.
Sequence
Readiness weak · capability weakStop discretionary AI spend. Baseline three processes, remediate data, build foundational literacy — in that order. Anything else is expensive theatre.
How it works
You do not score yourself.
Scores come from four independent evidence streams, weighted and reconciled by the assessor. Every score is anchored to a named artifact, interview, observed task, or telemetry figure. An unevidenced score is downgraded automatically.
Stream 1
Artifact review
We request ten specific documents — data catalogue, integration diagram, 90-day licence utilisation, AI policy, rollout post-mortems, process documentation, cost baselines, vendor data-processing terms, the named AI owner, and prior business cases with their actual results. Whether you can produce them is itself a scored signal.
Stream 2
Structured interviews
Role-specific scripts across executive, technical, operational, security, and finance stakeholders. Participant selection is not delegated entirely to the sponsor; we require a defined proportion of assessor-selected participants, including frontline staff in the candidate workflows.
Stream 3
Practical capability assessment
A sampled cohort completes a 30-minute observed work task from their own role. We watch what they do. Self-reported AI competency is systematically inflated, and no survey detects the most dangerous profile in an organization — the confident user who does not verify.
Stream 4
Telemetry
Where available we pull actual usage data: seat activity, query volumes, feature engagement, spend against budget. Numbers, not narrative.
Operating principles
Eight non-negotiables.
No happy ears
The engagement surfaces uncomfortable findings. A diagnostic that validates the sponsor's existing plan was not worth commissioning.
Evidence beats enthusiasm
Every score is anchored to a named artifact, interview, or observed task. "Leadership feels we're a 4 on data quality" is not a score.
Measure the before, or don't claim the after
If a process has no baseline — hours, error rate, cycle time, cost — no efficiency claim about it can ever be verified. Baselining is the first work, not the last.
Tool access is not capability
A licence is an input. We measure what people actually produce and whether they can tell when the output is wrong.
The bottleneck is rarely the model
It is almost always data access, permissions, undocumented process, or the absence of anyone accountable for the outcome.
Redesign beats automation
Pointing AI at an existing broken process makes the process faster and still broken. The value is in restructuring the work.
Shadow AI is a symptom, not a crime
When capable people route around the sanctioned stack, the stack is the problem. We measure the behaviour without punishing the people, because it is the most honest signal in the building.
The right answer is sometimes "don't"
Some workflows should not be automated. Naming those protects the credibility of everything else we recommend.
What you get
Ten leadership-ready deliverables.
| Deliverable | What it does |
|---|---|
| Executive readout memo | Board-level summary, posture, and the single recommended move |
| Organizational readiness scorecard | Seven domains, weighted and scored, with a per-domain evidence log |
| Human capability scorecard | Four domains scored by role group, including practical assessment results |
| Gate report | The hard blockers, ranked, with what it takes to clear each one |
| Shadow AI exposure report | What is actually being used, by whom, with what data |
| Value opportunity register | Ten candidate workflows, sized and ranked by provable value |
| Top-three business cases | The three best candidates, costed against a measured baseline |
| Spend efficiency review | What current AI spend is returning, and what to cut |
| 90-day action plan | Named moves, named owners, dates, and gates |
| Buy / build / wait / stop decision brief | The leadership call, with the reasoning behind it |
The commercial argument
In one paragraph.
A company of 150 people does not have a CIO's office, a strategy function, or an internal team whose job is to check whether the AI plan holds. It has a leadership team with day jobs, a vendor with a compelling deck, and a decision worth several hundred thousand dollars in licences, integration work, and the headcount that gets hired or not hired around it.
The diagnostic costs a fraction of that decision. It exists to tell you, before you commit, whether the work is worth changing, whether the business can support the change, and whether your people can execute it — and to tell you plainly if the answer is not yet.
What we need from you
These are engagement conditions, not preferences.
The diagnostic does not work without access. If access is restricted, we decline the engagement — a diagnostic conducted through a filtered view produces a filtered answer, and both parties waste the money.
Sponsor time
An executive sponsor available for two sessions and reachable throughout.
The artifacts
The ten documents on the request list — or a written acknowledgement of what does not exist. "We don't have it" is a valid and highly informative answer.
Interview access
Across executive, technical, operational, security, and finance roles, including assessor-selected frontline participants.
Assessment slots
Thirty-minute practical sessions for a sampled cohort, scheduled and attended.
Telemetry
Read access to licence and usage data for the AI tools already in use.
Unedited findings
Agreement that findings are delivered to the sponsor unedited.
Next step
The disciplined step between AI enthusiasm and AI spend that returns something.
Tell us what you have already bought and what you expected it to do. We will tell you whether a diagnostic would help, and what it would take.
Email contact@shockpoint.io and tell us what decision you are trying to make.