29 July 2026
How to measure whether AI training actually stuck
If you cannot describe a changed Tuesday, you measured the catering, not the training.
Who this is for
This is for the L&D, HR, or transformation lead who will be asked, three months after a cohort, "did it work?" The default pack is a smile sheet, a completion rate, and a quote from someone who liked the facilitator. None of those tell you whether the team now checks an output before it leaves the building.
If your only planned measure is "would you recommend this course," stop and write a better one before you book the room.
The business problem
Training that cannot be evidenced will be cut, or worse, repeated as if the first pass landed. Sponsors who bought AI seats already have a finance story they do not like. McKinsey reported in 2025 that only 39% of organisations see any EBIT impact from AI, while 88% use it in at least one function and only about a third have scaled past pilots. The Conference Board reported in 2026 that only 33% of workers received employer AI training in the last six months. TalentLMS reported in 2025 that 49% of workers say AI is moving faster than their company's training.
Those figures describe a market that spends on tools and under-invests in capability. If you then measure capability with a happiness score, you will not know whether you closed the gap. World Economic Forum reporting in 2025 found that 63% of employers name skill gaps as the top barrier to transformation. A smile sheet does not move that number.
This is not a request for a fabricated Experrt ROI case. We do not publish client percentages here because we will not invent them. The method is what you can run on your own work.
What not to treat as evidence
Do not treat these as proof the training stuck:
- "I would recommend this course."
- Attendance alone, if nobody was asked to do the work afterwards.
- A quiz that restates the slides.
- Seat usage going up for two weeks because people were being watched.
- A single champion's anecdote.
Those can be supporting colour. They are not the result.
What to collect instead
Pick a small set of signals you can observe without a research department.
1. A before-and-after work sample. Same type of task: a draft, a summary, a check-list. Score it against a written standard (sources, numbers, policy, refusal). Do this on five or ten pieces, not a census. Embedding AI in Daily Workflows already uses the job as the unit of practice. Reuse that artefact.
2. A checking behaviour you can count. For example: proportion of AI-assisted drafts that include a named check, or a record that a person reviewed the output. AI output verification at work is the briefing for that standard. You are measuring whether the standard is in use, not whether people enjoyed hearing about it.
3. Manager observation. A one-line note in the month after the cohort: is the team refusing bad outputs, or still shipping fluent errors? That is why Managers, not champions, after you buy the tools sits next to this one. If the manager was not trained, your measure will be noise.
4. The training record itself. Attendance, the assessed task, the grade or observed outcome, the export. AI Governance and Oversight for Managers is where sponsors learn what they will be asked to show. A missing register means you cannot even say who was in the room.
McKinsey has also reported that leaders with highly AI-fluent teams are 3.9 times more likely to capture enterprise value. You will not prove 3.9 times anything from one cohort. You can prove that a named team now produces checked work. That is enough for a quarter. Do not wait for a perfect measurement model. Ten dated files and a manager note beat a slide with a round percentage and no source.
How to set the measure before you buy
Write the success line into the brief. "Within four weeks, this team can show five work samples that meet the checking standard, and the manager can describe the escalation." If a vendor cannot tell you how the course produces those samples, they are selling a day, not a change.
Do not promise EBIT. Do not promise licence utilisation. Those have too many owners. Promise a behaviour you can see on the work.
What to insist on
- Success defined as behaviour on real work types, written before the quote.
- An assessed artefact the organisation keeps.
- A four-to-six week look-back with the manager, in the diary now.
- No vanity dashboard as the only KPI.
- Honesty about sample size. Five good files beat a fake 40% improvement.
Experrt will run the cohort and give you the attendance and assessment pack. We will not write a case study that pretends your margin moved because of a two-day course. If you want the look-back designed, say so in the brief.
What to do this quarter
Choose one team, one work type, and one checking standard. Capture three samples before the cohort. Run the training. Capture three after. Sit with the manager at week four and write down what changed and what did not. File the pack next to the policy. If nothing changed, do not buy a second course until you know why the first one missed the work.
Contact Experrt with the team and the work type. We will tell you whether a cohort can produce that evidence, and what it cannot.
Book a conversation
Experrt runs live, in-house cohorts. If this briefing matches a gap you already have, talk to us about scope, dates and the record the programme should produce.
Contact Experrt