Security awareness training metrics: measure more than completion

Measure security awareness with a practical metric dictionary covering participation, decisions, retention and reporting, plus fair comparisons and clear limits.

CyberPlay editorial team · Published · Updated · 8 min read

Guide and exercises in English

Scene from Ghost Protocol.

Expand image

From the CyberPlay Ghost Protocol gallery. Illustrative game scene; any interface text shown is in English.

Useful security awareness training metrics answer a specific question. Completion tells you whether an assigned activity was finished. A score can describe performance within that activity. A later scenario can examine whether the learner can apply a principle after time has passed. Workplace reporting describes another part of the system. Combining those measures into a single reassuring percentage hides the differences that make them useful.

Start with a decision the programme aims to support, define the evidence you can reasonably collect and decide what would prompt an improvement. This guide provides a metric dictionary and an illustrative evaluation design. It contains no customer outcome figures. Its examples show how to reason about measurements without treating game progress, quiz results or simulated phishing outcomes as proof that incidents have been prevented.

What you’ll take away

  • Define every numerator, denominator, time window and exclusion.
  • Separate participation, in-activity performance, retention and workplace behaviour.
  • Compare scenarios of similar difficulty and describe differences.
  • Use results to improve learning and processes; avoid unsupported breach-reduction claims.

1. Begin with the decision and evaluation question

Suppose the intended behaviour is verifying an unusual payment change through an existing trusted contact. The evaluation question might be: can finance participants select that route in a changed scenario after the session? You can then design the prompt, record the first decision and compare the reasoning with a simple scoring guide. That is more informative than asking whether the session was popular.

NIST SP 800-50 Rev. 1 includes evaluation and metrics within a maintained learning programme. Use the results to decide what to change: an unclear instruction, an inaccessible activity or a missing contact directory. A metric that has no plausible consequence for programme decisions may be unnecessary to collect.

Phishing Detective 3D gameplay: reviewing a supplier invoice.

Expand image · Game screenshot · English interface

  1. Compare the changed details

    Check which payment details changed and whether the request matches the expected invoice and work.

  2. Compare with trusted records

    Compare with a trusted invoice, then verify changed bank details through the supplier contact already on record.

Phishing Detective 3D gameplay: reviewing a supplier invoice.

Section sources: Building a Cybersecurity and Privacy Learning Program, SP 800-50 Rev. 1

2. Use a metric dictionary before building a dashboard

The definitions below are examples for a local evaluation plan. State whether the unit is a person, assignment, session or message. A learner may start several game sessions for one assignment, so session counts and employee counts cannot be substituted for each other. Document the window and eligibility rules beside the number.

2. Use a metric dictionary before building a dashboard
MeasureExample definitionWhat it can support
CompletionEligible assigned people who completed by the cut-off / eligible assigned peopleWhether the planned activity reached the intended group.
First-decision accuracyParticipants selecting the defined appropriate action on their first attempt / participants with a valid first attemptPerformance on the presented decision.
Reasoning qualityResponses meeting a defined explanation rubric / responses assessedWhether participants can explain a verification or reporting step.
Delayed applicationParticipants making the appropriate choice on a changed later scenario / participants with a valid follow-upApplication under the conditions of that follow-up.
Simulation reportingRecipients with a qualifying report / recipients with confirmed delivery, under the stated rulesReporting during that simulation.
Time to first useful reportElapsed time from defined campaign start to the first report meeting the stated criteriaAn exercise-level view of how quickly actionable information arrived.
Improvement closureAgreed programme or process actions completed by due date / actions dueWhether identified barriers were addressed.

3. Interpret completion before interpreting learning

A low completion rate may reflect unclear invitations, shift schedules, inaccessible controls, leave or missing devices. Separate those explanations from someone starting an activity and deciding not to finish. Record an eligibility rule before the reporting period so exclusions do not quietly change to improve a result. Show the count alongside the percentage, especially for small teams.

For example, an assignment aimed at twenty people may include two on extended leave. Decide whether they remain in the current denominator or receive a later due date and explain the choice consistently. Do not compare the resulting percentage with another team whose absence rules differ. The number only becomes interpretable once the operational context is visible.

4. Distinguish first attempts from supported practice

A learner who receives hints and retries until succeeding has completed useful practice, but the final result is not the same as an independent first decision. Keep those states separate where the activity supports it. If it does not, describe the available result honestly rather than reconstructing an unsupported first-attempt metric.

Write a small rubric for explanations. For a payment-change task, a complete response might identify the unusual change, choose a contact from an established record and obtain the required approval before acting. A partial response might recognise suspicion but verify through the same untrusted conversation. Let facilitators compare a few fictional responses before scoring, so the rubric is applied consistently.

Participation: Did people take part? Knowledge: Can they explain the rule? Retention: Can they apply it later? Workplace action: Do relevant behaviours change in context?

Expand image

Different measures answer different questions. Conceptual measurement framework; no customer results are shown. Original CyberPlay explanatory diagram.

5. Account for scenario difficulty

The NIST Phish Scale provides a method for rating the human detection difficulty of phishing emails. It considers cues and how the premise aligns with the recipient’s context. A generic prize message and a well-timed supplier request are not interchangeable test items. Raw click rates alone can therefore give a misleading comparison between campaigns.

For other kinds of scenario, document the features likely to affect difficulty: available information, time pressure, familiarity, device, language and the plausibility of choices. You do not need to claim a validated difficulty score where one does not exist. You do need to explain material differences when comparing performance across versions or teams.

Phishing Detective 3D gameplay: inspecting an incoming request.

Expand image · Game screenshot · English interface

  1. Read the requested action

    Identify what the message asks you to do before judging its familiar name or appearance.

  2. Verify through known contacts

    Use an established contact route to verify an unexpected request, even when the sender seems familiar.

Phishing Detective 3D gameplay: inspecting an incoming request.

Section sources: NIST Phish Scale User Guide

6. Use a small, transparent evaluation design

Here is an illustrative design for the payment-verification objective. Give participants a fictional baseline request and record the first action and reasoning. Teach the approved process, then let them practise with feedback. Later, present a changed request with comparable information and record a new independent response. Preserve the version and interval with the results.

An improvement can support a statement about performance on those exercises. It does not by itself isolate the effect of training: familiarity, parallel communications, staff changes or process improvements may contribute. If the organisation needs a causal evaluation, plan an appropriate comparison and analysis with qualified support before collecting data.

6. Use a small, transparent evaluation design
StageRecordInterpretation
BaselineScenario version, first action and reasoningStarting performance on this task.
PracticeFeedback offered and completed attemptsOpportunity to learn and use the process.
Follow-upChanged scenario, interval and independent responseLater application in the exercise.
ReviewConfounders and improvement actionsWhat can reasonably be concluded and changed.

7. Treat reporting as a joint employee-and-response process

More reports may reflect better awareness, more malicious messages, an easier reporting button or a noisy campaign. Fewer reports may reflect filtering or reduced opportunity, not better judgement. Pair reporting counts with context and define what makes a report useful. Distinguish a report receipt from a completed investigation and avoid asking employees to prove maliciousness before raising a concern.

Consider response capacity too. If reports enter an unattended mailbox, training people to report faster will not solve that bottleneck. Measure whether reports reach the right team and whether employees receive helpful acknowledgement when appropriate. Use authorised workplace records with a clear purpose and access policy rather than building an unnecessary store of personal message content.

Ransomware Reaction gameplay: preparing a useful incident report.

Expand image · Game screenshot · English interface

  1. Include time and device

    Report the observed symptoms, when they appeared, and the affected device through the organisation's approved reporting route.

  2. Distinguish observation from diagnosis

    Separate direct observations from suspected causes so responders can investigate without treating an early guess as fact.

Ransomware Reaction gameplay: preparing a useful incident report.

8. Read research that challenges easy success claims

A research preprint by Rozema and Davis, based on a large study at a US financial technology firm, reported no significant main effects of its training interventions on click or reporting rates, while phishing difficulty predicted behaviour. Those results concern that study and its tested approaches. They neither evaluate CyberPlay nor establish that every possible training design has the same effect.

The useful lesson is to test assumptions. Engagement, completion and an appealing interface do not eliminate the need for evaluation. Read the intervention, population, outcome definition and study limits before borrowing a result. A source headline is not enough to justify a product promise, whether the headline favours or criticises training.

Section sources: Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction

9. Know what CyberPlay records can and cannot show

CyberPlay’s session model records game-session status and results, including a reported percentage and pass state where supplied, together with game-specific metrics and details. These are learning-activity records. Availability and interpretation depend on the game and the relevant account or organisation view; a field in the data model does not mean every game supplies the same evidence.

Treat a game’s percentage, progression or completed run as evidence about that activity. A separate changed scenario is needed to examine delayed application. Workplace phishing-reporting behaviour and business outcomes need separate authorised collection. Do not infer that an employee prevented an attack, or quantify a reduction in organisational risk, from a game score.

10. Publish a compact report with its limits

A useful monthly report states the objective, audience, activity versions, participation, demonstrated decisions and actions taken. Add the main limitation near the result. For small cohorts, use counts and qualitative observations instead of dramatic percentage changes. Keep individual results visible only to the people who need them for the stated purpose.

  • What decision did the programme aim to support?
  • Who was eligible, and who could actually access the activity?
  • Which attempt, scenario version and time window produced the result?
  • What does the evidence show, and what remains unmeasured?
  • Which process or learning change will follow?
  • Who owns that change and when will it be reviewed?

Put the decision into practice

Try a phishing scenario and define one observable decision you could evaluate separately from completion or score.

Explore phishing games

Sources and further reading

  1. Building a Cybersecurity and Privacy Learning Program, SP 800-50 Rev. 1 — NIST. Accessed 2026-09-13
  2. NIST Phish Scale User Guide — NIST. Accessed 2026-09-13
  3. Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction — Research preprint, arXiv. Accessed 2026-09-13

Keep exploring

All articles

Contact · About