Skip to content

Most companies assess with multiple choice because it is the only thing that can be marked at scale. AI removes that constraint: open answers, recorded conversations and practical tasks can be scored for ten thousand learners a day, and every one of them gets feedback on their own work.

The problem

A multiple-choice test tells you who can recognise the right answer. It tells you almost nothing about whether someone can hold the conversation, structure the explanation or make the judgement call — which is what the training was for. Organisations know this and assess with tick-boxes anyway, because marking two thousand written answers is not a thing a team of four can do.

The consequence is quiet but expensive: training is evaluated on the wrong thing, and learners receive a number instead of guidance they could act on.

How it works

Assessment items are written to elicit behaviour — "the customer has just told you the competitor is 12% cheaper; write what you would say next" — and scored against a rubric your subject-matter experts sign off. The model is calibrated first against a sample your team has marked by hand; we report the agreement rate honestly, and tune until it is good enough to trust for the purpose at hand.

Each learner receives written feedback on their own answer: what worked, what was missing, what to do differently. Scores that sit near a pass/fail boundary, or that the model itself flags as uncertain, are routed to a human moderator before any consequence follows.

What changes at work

  • Assessment finally measures the thing the programme was about, so the evaluation is worth reading.
  • Every learner gets individual feedback — in our experience the single most-noticed improvement by participants.
  • Cohort patterns appear while the programme is still running: if forty per cent of a region cannot structure the discovery question, the facilitator knows before the follow-up day.
  • Recruitment and promotion screening can use work-sample questions instead of aptitude proxies, at volumes that were previously impossible.

How we prove it

Two numbers, both published to the client: agreement between AI scores and human scores on a held-out sample, and the correlation between assessment results and the business behaviour they are meant to predict. If the second is weak, the assessment is measuring the wrong thing and we rewrite it — that is a normal part of the work, not a failure.

What it cannot do

It cannot decide a person's career on its own. Language models score confidently even when wrong, they can be nudged by fluent nonsense, and they carry whatever bias sits in the rubric they were given. Everything consequential — certification, promotion, exit — goes through a human who can see the AI score, the answer and the moderation note, and who signs their name to the decision.

Frequently asked

What can AI assess that a normal quiz cannot?

Anything that is not multiple choice: a written explanation of how the candidate would handle a customer, a recorded sales pitch, a filled-in planning template, a role-play conversation. These are the assessments that predict job behaviour, and they are exactly the ones most companies drop because marking them at scale is impossible.

How accurate is the scoring?

Accurate enough to rank and to give feedback, not accurate enough to be the sole basis of a consequential decision. We calibrate the model against a sample your team has scored by hand, report the agreement rate, and route anything near a pass/fail boundary to a human moderator. Where the decision affects someone's job, a person signs it.

Can we use our own rubric?

Yes — and we prefer it. Your rubric, your weights, your pass marks. If you do not have one, we build it with your subject-matter experts before the first learner is assessed.

How fast are results?

Feedback is returned in minutes rather than weeks, which is the point. Learners read feedback when they still remember the answer; managers see cohort patterns while the programme is still running and can act on them.

Does this work for regulated or certification assessments?

It supports them; it does not replace them. AI marks the practice and the formative attempts, a human moderates the summative one, and the audit trail records who decided what. Any certification claim rests on the human decision.

Tell us about the team you want to develop.

We come back within one working day with a first-cut approach — content, facilitation, technology and analytics in the right mix.

Request a proposal