Your Assessment Questions Are Evidence. Write Them That Way.

Most guidance on writing assessment questions comes out of education. In safety-critical industries, it aims at the wrong outcome.
THE SHORT VERSION
After an incident, your question set is what gets produced. Write for that moment.
Recall questions test whether a worker knows the rule. Scenario questions test whether they recognise the moment it applies. Incidents live in the gap.
In a multilingual workforce, a badly worded question tests English comprehension and records the result as competence.
One correct answer, eighteen months ago, is not a result.
A question is a record you may have to defend
In construction, energy, manufacturing, mining, and an other high risk industries, assessment questions have a second life.
After an incident, the question set is what gets produced. "Was this worker competent to carry out that task" becomes "show us what you asked them, show us what they answered, and explain why you believed that proved anything."
A question is not just a teaching instrument. It is a record you may one day have to defend.
That changes the standard. Everything below follows from it.
1. Start with what the assessment is actually for
The formative / summative / diagnostic taxonomy is useful in a classroom. On a site, the real categories are different:
Verification before access: the worker cannot start until they pass and have a demonstrated understanding. The highest-stakes category, and the one most organisations write their weakest questions for.
Cadence based refreshers: periodic reassessment confirming that knowledge held and identification of gaps.
Incident-triggered reassessment: something happened, and you need to establish what the crew actually understands and close any gaps.
Third-party credential confirmation: the card says they are qualified. Your assessment establishes whether they can apply it in your environment, to your standards.
Each needs a different difficulty, a different length, and a different tolerance for failure.
The most common mistake: one question bank, used for all four.
2. The anatomy of defensible assessment questions
The classic rules survive the reframe. Their justification changes.
Principle | The real test |
Clarity | If a third party cannot see what was being asked, the answer proves nothing. |
Task relevance | Align to what the worker will do, not to what the video covered. Not always the same thing. |
Unambiguity | Can someone outside your organisation see why each wrong answer is wrong? |
Difficulty matched to risk | Housekeeping and energy isolation should not be equally easy to pass. |
Fairness | See section 5. In a multilingual workforce this is the failure mode, not a footnote. |
3. Question types, honestly assessed
Type | Best for | Watch out for |
Multiple choice | Factual recall, fast marking, analysis at scale | Rewards recognition of a familiar phrase. Write distractors a real worker might choose. |
True / false | Quick checks, confirming a specific prohibition | 50% is a pass by chance. Never let it carry an access decision alone. |
Short answer | Genuine understanding, in the worker's own words | Expensive to mark consistently. Reserve for what matters most. |
Scenario-based | Judgement under real conditions | The hardest to write well. See below. |
4. Why scenario questions are worth the effort
Work does not present itself as a rule-retrieval problem. It presents as a situation.
The lift is behind schedule. The wind is picking up. The correct fitting is in a container nobody can find, and the supervisor is on the other side of the site.
Most serious incidents are not caused by workers who did not know the rule. They are caused by workers who did not recognise that this was the moment it applied or who recognised it and got talked out of it.
Recall questions test the first half of that. Scenario questions test the half that actually hurts people.
They also carry more evidentiary weight. A correct answer to "how often must a harness be inspected" can be dismissed as a memorised fact. A correct answer to "you arrive at the anchor point and find this, with the crew waiting on you what do you do next" is much harder to argue away as a lucky guess.
How to write one that works
Ground it in a real site condition your people would recognise, with details that fit your environment.
Include the pressure. Schedule, weather, a missing part, a crew standing around. A scenario without pressure is a definition wearing a costume.
Make every wrong answer plausible. Something a competent, tired, well-intentioned person might genuinely do. Shortcuts, not straw men.
Allow exactly one defensible action, and several tempting ones.
Keep it to three or four sentences. Past that, you are testing reading stamina.
Map each distractor to a specific misconception, so item analysis tells you what to retrain, not just who failed.
Scenarios are the hardest question type to write, which is why most assessment banks contain almost none of them. That has now changed.
5. Fairness means language and language means competence
"Avoid cultural and gender bias" is borrowed from academic testing. It is far too soft for our industry.
On a modern site, bias in an assessment usually means one specific thing: you tested English comprehension and recorded the result as competence.
Consider a worker who has isolated energy sources correctly for twenty years, and fails a question because it runs to forty words with a subordinate clause in the middle. You have not identified a competence gap. You have created a false record, and the record is the thing you will be asked to stand behind.
Write at the reading level of the workforce you actually have. Deliver the assessment in the worker's own language. Then the result means what you need it to mean.
6. Pitfalls with one correction
Leading questions that signal the expected answer.
Double-barrelled questions that ask two things and accept one answer.
Overcomplicated language where plain wording would do.
Negatives - and here the standard advice is wrong. It does not survive contact with safety, where prohibitions are the whole point and "which of the following is not permitted in a confined space" is a legitimate, necessary question. The real rule is narrower: never use a double negative, and never bury the negation mid-sentence. Signal it clearly and put it where it cannot be missed.
7. One correct answer is not a result
Competence decays. Standards change. Crews turn over.
Assessment is not an event at the end of a course. It is a cadence and the shape of that cadence is itself part of what you will be asked to justify.
How LUMA1 handles this
Scenario questions, generated for you. Our AI reviews your video content and generates questions from it automatically including full scenario-based questions built from the situations and conditions in the source material. Scenario writing has always been the bottleneck that kept organisations stuck on recall. It is not a bottleneck any more.
Generation is only half of it. Because these questions carry evidentiary weight, every generated question enters a review workflow before it reaches a worker. Questions carry version history. Each one traces back to the exact moment in the source content it came from.
When someone asks why a worker was cleared to start, the answer is a record. Not a recollection.
Editing is quick. Every question is available in the worker's own language.
If your current assessment bank is made entirely of recall questions, that is the place to start.
John Hudson Co-Founder & CEO, LUMA1




Comments