Short answer: To choose AI recruitment software for frontline hiring, start from your own evidence gap: find which job requirements your process decides on without real evidence, and pick the type of tool that moves them from claim to evidence. Then ask every provider the same ten questions, and check that a person stays responsible for every hiring decision.
Most AI recruitment software is bought after a demo. The demo shows what the tool can do. It rarely shows whether that is the problem you actually have.
For frontline and high-volume hiring, that mismatch is expensive. A team that struggles with late certificate checks buys an interview tool. A team whose real gap is a shift-pattern mismatch buys a CV screener. Both tools work. Neither fixes the problem.
The Signal Gap, a term Radius Hire introduced in The Radius Papers 01 (2026), is the distance between what a hiring decision rests on and what predicts the outcome. Choosing well starts with finding yours.
The 10 questions at a glance
- What does the system do with each answer: record, transcribe, summarise, score or rank?
- Who writes the questions and the criteria, and can we change them for each role?
- Is any candidate rejected or moved forward without a person reviewing their case?
- How are candidates told they are dealing with an AI system, and what are they told?
- Does the system infer emotions, or assess accent, appearance or background?
- What validity evidence do you have, for which roles, with what sample size, period and population?
- How do you test for differences in outcomes between groups, and can we see the results?
- Where is candidate data stored, for how long, and who can access it?
- Is a data processing agreement available, and who are your subprocessors?
- How will the system support our duties as a deployer under the AI Act?

On this page: Why this matters now · Where to start · What good looks like · The 10 questions · Practical checks for frontline roles · How to run a pilot · How Radius Hire answers · FAQ
Why this matters now
AI used to recruit or select people is high-risk under Annex III of the EU AI Act. After the Digital Omnibus on AI (Regulation (EU) 2026/1744), the main high-risk duties apply from 2 December 2027. Some rules already apply: the ban on inferring emotions at work from biometric data such as face or voice (since 2 February 2025), AI literacy duties (since 2 February 2025), and the duty to tell people when they are interacting with an AI system (since 2 August 2026). Our guide to the EU AI Act and recruitment sets out what employers must do, and when.
A later date changes the timing, not the direction. A tool you adopt now will be judged under the GDPR and equal treatment law today. Whether the AI Act's high-risk duties also reach a system already in use before December 2027 depends on how much it changes after that date. Ask counsel. Either way, the records you need are the same.
Selection researchers say they want to see newer methods, such as asynchronous video interviews and scoring based on language processing, compared with traditional ones. A body of validity evidence on them has not yet built up (Sackett and colleagues, 2023). Treat any provider's validity claims with care until independent evidence exists.
Where should you start when choosing AI recruitment software?
Start from your own process, not from the demo: find which requirements your hiring decides on without real evidence. To find them, run a one-hour audit on your highest-volume role. For each requirement, note the evidence your process holds at the moment of decision: verified, demonstrated, structured, claimed or not covered.
Then match what you find to the type of tool that could help. Many products combine several types.
| Type of tool | Fixes this gap | What still needs a person | What to watch |
|---|---|---|---|
| Candidate verification | Requirements with legal weight, such as certificates, are only claimed at the decision | A legal basis, reviewing each check, handling errors | Proportionality, and the stricter rules for identity documents (below) |
| Structured or asynchronous interviews | Important requirements, such as a 05:30 start or the shift pattern, are never asked, or asked differently each time | Writing the questions and the criteria for a good answer | Validity evidence is still limited; check the emotion ban (question 5) |
| Recruiter decision support | Evidence is collected but never judged against a standard, such as phone-screen notes that differ per recruiter | Rating the evidence and deciding | Recommendations can become decisions in practice if nobody checks them |
| Structured assessments | Knowledge that can be tested before hire, such as site safety rules, is never tested | Checking the test fits the role | Some methods suit inexperienced applicants poorly |
| Skills or requirement matching | Nobody can see which requirements each candidate supports | Setting good requirements | Only as good as the requirements and the evidence behind them |
| CV and application screening | The only bottleneck is reading hundreds of applications | Remembering that the output still rests on claims | Years of experience is a weak predictor |
Based on The Radius Papers 01: A New Standard for Frontline Hiring, Table 7.1. No type is proven better than the others for frontline hiring outcomes.
A note on identity checks. In the Netherlands, an employer may check an applicant's identity by looking at the original document, but may not copy, scan or photograph it, or ask for the citizen service number (BSN), until the person is hired (Autoriteit Persoonsgegevens). Certificates are the check you can realistically move before the offer.
This step changes the question you ask in the demo. Instead of "what can your tool do?", you ask "which of these requirements will your tool move from claim to evidence, and how?"
What good looks like
Good technology moves requirements up the evidence levels without taking the decision away from people. Check for all eight:
- Every question and check ties back to a written, job-related requirement.
- Every candidate for a role gets the same questions, even at volume (how to run structured interviews at high volume).
- Answers are recorded against criteria written in advance.
- Checks happen before the decision where lawful and proportionate.
- The reviewer can see the evidence behind every output.
- Candidates know when they are dealing with AI, what is recorded and who decides.
- A named person reviews the evidence for every candidate who is decided on, can question, correct and overrule any output, and is recorded as the person who decided.
- The tool does not decide on its own who is hired or rejected.
Our guide to AI in hiring sets out what technology should evaluate and what people must decide.
What questions should you ask an AI recruitment software provider?
Ask every provider the same ten questions, in writing, and compare the answers side by side.
1. What exactly does the system do with each answer: record, transcribe, summarise, score or rank?
Why it matters: A structured interview asks every candidate the same job-related questions and rates each answer against a standard written in advance. A tool that shows each answer next to its criteria keeps that design visible. A single opaque score, a ranking or automatic progression puts more weight on the tool, which makes human oversight and the GDPR rules on automated decisions more important to get right.
A good answer includes: a plain description of every step, and reviewer access to the answers and evidence behind any output.
Red flag: a score or ranking nobody can trace back to what the candidate said, or vague answers about who can see the evidence.
2. Who writes the questions and the criteria, and can we change them for each role?
Why it matters: The methods at the top of the research rankings are built around the specific job (Sackett and colleagues, 2023). Generic question sets may not fit your roles.
A good answer includes: questions and criteria you write or approve for each role, and can change.
Red flag: one generic question set for every role, which you cannot see or edit.
3. Is any candidate ever rejected or moved forward without a person reviewing their case?
Why it matters: Under Article 22 of the GDPR, a decision about a candidate that has significant effects may not be based solely on automated processing, unless a narrow exception applies, such as necessity for entering into a contract. Our guide to human oversight in hiring sets out what real review looks like.
A good answer includes: no, with a named person who reviews each case, can overrule the output and is recorded as the person who decided.
Red flag: "The AI makes the decision, so you save all that time." That puts a solely automated decision in the process.
4. How are candidates told that they are dealing with an AI system, and what are they told?
Why it matters: Since 2 August 2026, Article 50 of the AI Act requires providers to design AI systems that interact directly with people so those people are told, unless it is obvious. The GDPR also requires information about how data is used (Articles 13 and 14).
A good answer includes: the exact words candidates see or hear at the start, what is recorded, who decides, and the languages the information is available in.
Red flag: disclosure only in the terms and conditions, or only in a language many of your applicants do not read.
5. Does the system infer emotions, or assess accent, appearance or background?
Why it matters: Since 2 February 2025, Article 5(1)(f) of the AI Act has banned AI systems that infer the emotions of people at work. The ban covers inferring emotions from biometric data such as a face or voice, with exceptions for medical or safety reasons. The European Commission's guidelines, which are not binding, say it also covers job candidates (C(2025) 5052 final).
A good answer includes: a clear no, in writing, covering emotions, tone of voice, accent, appearance and background.
Red flag: any analysis of tone, emotion, accent or appearance, whatever it is called.
6. What evidence of validity do you have, for which roles, with what sample size, period and population?
Why it matters: An average is not a promise. Even for structured interviews, results vary widely between settings (Sackett and colleagues, 2023). A figure without a sample, period and population cannot be judged.
A good answer includes: a study on comparable roles, with the sample size, period, population and outcome measured (for example supervisor ratings or 90-day retention), ideally independent.
Red flag: an accuracy or validity figure with no study behind it.
7. How do you test for differences in outcomes between groups of candidates, and can we see the results?
Why it matters: No tool removes bias on its own. Fairness can only be shown by monitoring outcomes. Comparing groups can involve special categories of personal data, which needs a lawful basis.
A good answer includes: what is monitored, how often, on what lawful basis, and the results you can see.
Red flag: "Our AI is unbiased" or "eliminates bias". No verified evidence supports that for any screening method.
8. Where is candidate data stored, for how long, and who can access it?
Why it matters: The GDPR requires that personal data is kept no longer than necessary (Article 5(1)(e)). The Autoriteit Persoonsgegevens says it is customary to delete an unsuccessful applicant's data within four weeks of the procedure ending, or up to one year with their consent. You need a documented retention period for recordings, transcripts and verification results.
A good answer includes: the hosting location, a default retention period with automatic deletion, and which roles can open recordings.
Red flag: no clear answer on retention and deletion of recordings.
9. Is a data processing agreement available, and who are your subprocessors?
Why it matters: The provider processes candidate data on your behalf. You remain responsible for it.
A good answer includes: a data processing agreement you can read before you sign, and a current list of subprocessors.
Red flag: no subprocessor list, or a data processing agreement offered only after signing.
10. How will the system support our duties as a deployer when the AI Act high-risk rules apply?
Why it matters: Article 14 requires providers to build high-risk systems so the people overseeing them can understand their limits, stay alert to over-reliance and override the output. Article 26(2) requires employers to give that oversight to people with the competence, training and authority to use it. From 2 December 2027, employers using high-risk recruitment AI must also monitor the system, keep the logs under their control and inform candidates (Article 26).
A good answer includes: the features your reviewers use to see the evidence and override outputs, logs you can keep, and instructions for use.
Red flag: a promise that the tool takes care of your AI Act duties for you. Deployer duties stay with the employer.

Practical checks for frontline roles
The ten questions cover governance. These four cover whether the tool works for your candidates and your team.
- Languages: can candidates complete it in the languages they actually speak?
- Phone first: can they finish it on a phone, without downloading an app?
- Completion: what share of candidates finish it, and where do they drop out? Measure this in your own pilot.
- Systems and cost: does it connect to your applicant tracking system, and is it priced per vacancy, per candidate or per user?
If you are a staffing agency
Ask two more questions:
- Can we use this under our own name without becoming the provider? An agency that offers a screening system to clients under its own name, or substantially modifies it, may take on provider duties (Article 25). Check this with counsel.
- How are candidate data and evidence shared with clients, and under which GDPR roles?
Our guide to the EU AI Act and recruitment covers the agency position in more detail.
How to run a pilot
- Decide in advance which requirements the tool should move up the evidence levels, based on your audit.
- Measure a baseline first: time to decision, candidate drop-out, checks completed before the offer, early leaving, attendance and a supervisor rating.
- Change one thing at a time where you can, and read results over several hiring rounds. Ten hires is not enough to judge.
- Repeat the audit at the end of the pilot and compare the count of requirements at each evidence level.
How Radius Hire answers these questions
Radius Hire should be held to these ten questions like any other provider. Here is how we answer them today. Where the answer is "not yet", we say so.
- Q1: What it does with answers. Radius Hire asks every applicant the same structured questions, built on the requirements the employer sets, and shows the recruiter the evidence behind each requirement, labelled Proven, Claimed or Missing. It does not rank candidates or reject anyone on its own. A recruiter decides. The paper records evidence at five levels: verified, demonstrated, structured, claimed and not covered. Radius Hire's product shows three labels: Proven (broadly verified), Claimed and Missing (not covered). We also show a count, such as "meets 4 of 5 requirements with evidence". It summarises evidence a recruiter can open. It is still the tool's judgement, not a prediction, and it does not order candidates: they appear in the order they applied.
- Q2: Questions and criteria. The employer sets the requirements for each role. Every applicant answers the same structured questions built on them.
- Q3: Human review. Radius Hire does not reject candidates on its own and does not decide who is hired. Recruiters review, challenge and overrule what it presents.
- Q4: Candidate information. Radius Hire's interview is a structured voice interview. Candidates are told at the start that the interviewer is automated and that a person reviews the result. They can take it in seven languages (Dutch, English, Polish, Romanian, Ukrainian, Russian and French), and can ask to speak with a recruiter instead.
- Q5: Inference. Radius Hire's interviewer does not infer emotions or assess accent, appearance or background.
- Q6: Validity evidence. Radius Hire has not published validity results. We make no performance claim.
- Q7: Group outcome testing. Radius Hire does not yet provide group outcome reports. Employers who want to compare outcomes between groups need a lawful basis for it, and we recommend legal advice before starting.
- Q8 and Q9: Data. In Radius Hire, interviews can be kept as a recording, a transcript and a summary. Candidate data is hosted in the EU under a data processing agreement. The employer sets how long interviews are kept, and we recommend following the Autoriteit Persoonsgegevens norm in question 8.
- Q10: Deployer support. Under Annex III we treat Radius Hire as a high-risk AI system and design it for human oversight ahead of 2 December 2027. Today that means a recruiter sees the evidence behind every requirement, can overrule anything Radius Hire presents, and makes every decision.
FAQ
How do you choose AI recruitment software?
Start from an audit of your own process: find which job requirements your hiring decides on without real evidence, and pick the type of tool that moves them from claim to evidence. Then ask every provider the same ten questions and check that a person reviews and decides every case.
What are the red flags when buying AI hiring software?
Be careful with claims that the AI makes the decision, validity figures with no sample, period or population, promises that the tool is unbiased, and any analysis of emotion, tone, accent or appearance. Vague answers on who sees the evidence or how long recordings are kept are red flags too.
How much validity evidence is enough?
There is no fixed threshold. Ask for a study on comparable roles with sample size, period, population and the outcome measured, ideally independent, then check it against your own pilot results.
Should an AI tool score or rank candidates?
Designs differ, and neither scoring nor ranking has published evidence of better hiring outcomes. A single score, a ranking or automatic progression puts more weight on the tool's output, so human oversight matters more.
Which type of AI hiring tool should we choose?
It depends on your gap. Verification helps when requirements with legal weight are only claimed at the decision, structured interviews help when important requirements are never asked, and decision support helps when evidence is collected but not judged against a standard.
Sources
- Radius Hire (2026). The Radius Papers 01: A New Standard for Frontline Hiring. Sections 2, 6.2, 6.4, 7.1 to 7.7 and 8.
- Sackett, P. R., Zhang, C., Berry, C. M., and Lievens, F. (2023). Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors. Industrial and Organizational Psychology, 16(3), 283 to 300. Open access. DOI 10.1017/iop.2023.24
- Regulation (EU) 2024/1689 (AI Act), as amended by Regulation (EU) 2026/1744: Articles 5, 14, 25, 26 and 50, and Annex III. EUR-Lex
- European Commission (2025). Guidelines on prohibited artificial intelligence practices, C(2025) 5052 final, 29 July 2025 (first published 4 February 2025), para. 254. Not binding. European Commission
- Regulation (EU) 2016/679 (GDPR), Articles 5, 13, 14 and 22. EUR-Lex
- Autoriteit Persoonsgegevens. Personal data of applicants. Autoriteit Persoonsgegevens
