CEFR-Compliant AI Scoring: How Talketet Built a Reliable Language Testing Tool
August 14, 2026

CEO, Talketet

Table of Contents
- Most standardized English and language tests for recruitment have weak validity
- How to choose the most reliable language testing tool for hiring staff
- How did Talketet build a CEFR-compliant AI that scores like a human rater?
- Language testing and validation: what academic research actually proves
- How the Talketet language assessment has been validated against bias
- Reliable language assessment for high-volume recruitment
Most language assessments used in recruitment rely on weak academic validation.
The reason is structural. In most cases, the language test is one test type inside a bigger pre-employment package. It sits next to cognitive tests, personality questionnaires, and situational judgment items.
The package is sold as one product, and the validity evidence is presented as one product too. So the language part inherits credibility it never earned on its own.
The Talketet tests for recruitment are backed by deep academic validation and are designed for hiring contexts.
This way, we give recruiters a high-quality language test at a price ideal for volume hiring.
Most standardized English and language tests for recruitment have weak validity
Multiple choice is still the default for language tests in the recruitment market.
A candidate reads a sentence, picks the missing preposition, and a score comes out.
These grammar quizzes do not test whether a candidate can handle a customer complaint on the phone or write a clear email to a client.
The gap matters because a recruiter needs to know if the candidate can speak and write at the level the role requires.
Testing only what a candidate can read or listen to, and then assuming they can speak at the same level, is a guess.
Until recently, there was an excuse for this. Testing speaking and writing at scale meant human raters, training, calibration and cost.
AI changed that. Tools like Talketet can test speaking and writing the way a trained rater would.
How to choose the most reliable language testing tool for hiring staff
Four questions separate a deeply validated test from a confident one.
- Does the test score speaking and writing skills, or only grammar, listening and reading skills?
- Is the CEFR alignment written down somewhere a buyer can read it?
- Has scoring consistency been measured?
- Who did the validation and did they have a commercial interest in the result?
Almost any vendor will say yes to the first three. The useful move is to ask for the document behind each yes.
The Talketet language test has been validated by independent language researchers from leading European universities. Let us know and we will share the documentation with you.
How did Talketet build a CEFR-compliant AI that scores like a human rater?
The starting point is the construct, meaning what the test is built to measure. Our framework combines the CEFR descriptors with Processability Theory.
Processability Theory describes how learner language develops through a fixed sequence of stages, one after the other. That lets the assessment check whether the level a candidate shows fits the way a second language is really learned.
Item design follows from that. We keep multiple choice to a minimum, because it can be guessed. Wherever possible, we replace it with items where the candidate has to speak or write.
Those answers are then scored against an explicit rubric with fixed weights: grammar, content, vocabulary, comprehension, cohesion, plus fluency and pronunciation for speaking.
Scores are mapped to CEFR levels for each skill and for the test as a whole.
Consistency is where automated scoring could fail, so we measured it directly. We scored the same responses again and again, then checked how the results changed between runs. We found the parameters consistent in more than 95% of cases.
This research proved that the AI behind Talketet can give CEFR-compliant results like human assessors would do.
The full method, the weights, and the results are set out in the papers in our research section.
Language testing and validation: what academic research actually proves
Deep validation is rare in language testing. It usually happens only in high-stakes exams, where a single score can decide whether a university admits a student.
These tests publish their research and explain exactly what they measure and how they score it.
They are also priced for one person taking one exam, so they are too expensive for recruitment.
At Talketet, we wanted the same level of quality in our tests with a scalable price for recruitment.
We did this work with linguistics researchers from the University of Pavia and the University of Rome Tor Vergata.
The results were presented at CLiC-it 2025 (Italian as a second language) and at the AIA (Italian Association of English Studies) and GSCP (Study Group on Spoken Communication) conferences (English).
The validation process went through peer review; our articles are public, and you can find them.
Linguistics researchers with no commercial interest validated our framework and technology.
How the Talketet language assessment has been validated against bias
Bias shows up in a language assessment when candidates who share a trait that has nothing to do with proficiency get different scores as a group.
In hiring, that is both a legal problem and a measurement problem, and the EU AI Act treats employment assessment accordingly.
The first bias check in the published validation looked at speaker gender. We ran the same spoken responses with a male voice and a female voice, across all scoring parameters, to see whether the model was reacting to the speaker rather than to the performance. Gender had no relevant effect on the scores.
You can find more information in this paper.
This study makes Talketet one of the few AI language testing tools for recruitment validated against bias.
Reliable language assessment for high-volume recruitment
In the past, recruiters hiring at high volume had to choose between two options:
- an expensive, validated language test with human assessors
- a cost-effective but inaccurate multiple-choice language test
With Talketet, they now get the best of both worlds:
A deeply validated language test with volume-based pricing that scales for recruitment.
Language tests built around your needs
We measure how your candidates speak and write in your real work context. Customised language assessment at scale.
Book a Demo
