🔥

Contracting

Explore the new contractor management module

🔥

Contracting

Explore the new contractor management module

🔥

Contracting

Explore the new contractor management module

🔥

Contracting

Explore the new contractor management module

Blog

AI Candidate Scoring: What is it, how does it work, and how to use it responsibly? The ultimate guide (2026)

AI candidate scoring

Last updated:

AI Candidate Scoring: What is it, how does it work, and how to use it responsibly? The ultimate guide (2026)

Innovations

Iwo Paliszewski

Iwo Paliszewski

You search Google for ‘AI Candidate Scoring’ and, more often than not, you are presented with one of two answers. The first goes: artificial intelligence reads CVs, assigns a matching percentage to the candidate and highlights the top talent. The second warns that the algorithm takes over the recruiter’s decisions, creates soulless rankings and can automatically shut someone out of a job opportunity. Both answers simplify the topic so much that instead of clarifying it, they merely build further misunderstandings.

AI Scoring does not have to be a magical autopilot selecting the ‘best people’, nor a black box passing judgment. In a well-designed process, it is a tool supporting the assessment of a specific application or candidate against the requirements of a specific project. It can present a numerical score, for example from 0 to 100, but it can also provide a contextual assessment: a profile summary, strengths, weaker areas, information gaps and questions worth asking during an interview. It delivers the greatest value when it does not replace human judgment, but helps structure, speed up and better justify it.

This guide answers the question ‘what is AI Candidate Scoring’ as it should be answered in 2026: comprehensively, practically and without marketing shortcuts. We explain exactly what is assessed, where the score comes from, how scoring differs from AI Matching, CV parsers and disqualifying questions, how to set up solid criteria, what data can be analysed, where the risks lie, and what GDPR and the EU AI Act say about AI used in recruitment.


Quick Definition
AI Candidate Scoring is an AI-supported assessment of a specific application or candidate against the requirements of a given recruitment project. The result can be numerical, contextual, or a combination of both. A good scoring system does not just show a number; it also explains what influenced the rating, highlights the strengths and weaker areas of the profile, and suggests questions for further verification. It is not the same as AI Matching, which is designed to search the database and suggest candidates who fit a project.


Table of Contents

1. What is AI Candidate Scoring?

2. What exactly does AI Scoring assess: the human, the candidate or the application?

3. Why has AI Scoring become important right now?

4. Numerical and contextual assessment - two dimensions of good scoring

5. How does AI Scoring work step-by-step?

6. What data can AI Scoring use?

7. Why are project requirements more important than the AI model itself?

8. Must-haves, nice-to-haves and disqualifying criteria

9. AI Scoring vs. keyword search

10. What does a score from 0 to 100 mean?

11. Strengths, weaknesses and interview questions

12. Missing information is not the same as a lack of competence

13. AI Scoring vs. the traditional candidate scorecard

14. AI Scoring vs. AI Matching - a key distinction

15. AI Scoring vs. CV parser, AI summary and killer questions

16. AI Scoring in practice - three assessment examples

17. Benefits for the recruiter and the recruitment team

18. Benefits for agencies, in-house HR and Hiring Managers

19. Benefits and risks from the candidate’s perspective

20. Where does AI Scoring deliver the most value?

21. How to implement AI Scoring in your recruitment process?

22. Most common mistakes when using AI Scoring

23. Limitations of AI Scoring - what does the system not know?

24. Bias, discrimination and the illusion of objectivity

25. Human-in-the-loop - what does real human oversight mean?

26. AI Scoring and GDPR

27. AI Scoring and the AI Act

28. How to choose an AI Scoring tool?

29. How to measure the effectiveness of AI Scoring?

30. How does AI Scoring work in Recruitify?

31. The future of AI Scoring

32. FAQ - Frequently Asked Questions

33. Summary - AI Scoring does not choose the person, it supports the evaluation


1. What is AI Candidate Scoring?

AI Candidate Scoring is a method of supporting the assessment of an application or a candidate's profile using artificial intelligence. The system analyses the available information about the candidate, compares it to the requirements of a specific project and prepares a result designed to help the recruiter in their next steps. Depending on the solution, the result can take the form of a number, percentage, category, description or a set of several elements. In Recruitify, this includes a score from 0 to 100, a contextual description of the candidate, a comparison of strengths and weaker areas, and suggestions for questions to ask during the interview.

The key phrase here is 'against the requirements of a specific project'. AI Scoring should not evaluate whether someone is a good or bad candidate in general. There is no universal professional value of a human being that can be honestly wrapped up in a single number. The exact same candidate might get a high score in a project looking for someone to scale B2B sales in the Polish market, and a significantly lower score in a project requiring experience in international enterprise sales. This does not mean their skills changed between the two measurements. The context and objective of the evaluation changed.

That is why it is safest to think of AI Scoring not as an evaluation of a person, but as a structured analysis of how well the currently available information fits. This phrasing is longer, but much more precise. The system does not see the entire professional history, potential, character, motivation or behaviour of the candidate in their future job. It sees a specific set of data and compares it against a specific set of criteria. Its utility therefore depends entirely on the quality of both sides of this comparison.

It is also worth distinguishing scoring from ranking itself. Scoring is the process of evaluating and explaining a result. Ranking is one of the possible ways to use this result, for example, by sorting applications from the highest to the lowest score. You can have scoring without an automatic ranking, and you can also create a ranking based on simple rules that have nothing to do with AI. In practice, ATS systems often combine both elements, but from the perspective of a responsible process, they should not be treated as synonyms.


The Golden Rule
AI Scoring should answer the question: ‘What in the available data suggests this candidate fits this role, what raises doubts, and what still needs to be verified?’, rather than: ‘Does this person deserve the job?’.


2. What exactly does AI Scoring assess: the human, the candidate or the application?

In everyday language, people usually talk about 'candidate scoring', but within a recruitment system, the object of evaluation can be both an application and a candidate assigned to a project. This distinction has practical significance. An application is a specific submission to a specific job opening. It can include the CV sent in response to the advert, answers from the application form, the application source, and additional documents or information provided during submission. A candidate, on the other hand, is a person who already exists in the database, who can be added to a project by a recruiter, sourced, or reconsidered in a new recruitment round without submitting a new application.

AI Scoring can therefore assess an incoming application, but it can also be run for a person found previously in the database or added to a project by a recruiter. In both cases, the evaluation should be anchored in the requirements of that specific recruitment. However, the scope of available data may differ. For an application, the system often has a fresh CV and answers to application questions at its disposal. For a database candidate, it can use the profile, previous documents, LinkedIn data, and information gathered in the system. If some of the data is outdated, the score may also be less up-to-date.

This leads to an important conclusion: a score is not a permanent label attached to a person. It should not be logged in anyone's mind or in the organisation's culture that 'Anna is a 64-point candidate' or 'Piotr is a 91 per cent candidate'. The correct phrasing is: 'Based on currently available data, in relation to the requirements of Project X, the system assigned Anna's application a score of 64/100'. The exact same profile in another project might receive a completely different score. Even within the same recruitment, the score can change if the recruiter refines the requirements, the candidate updates their details, or a new CV is uploaded.

This approach also guards against one of the biggest dangers of automation: turning a supporting score into a label that defines a person. In recruitment, the anchoring effect is very powerful. When a recruiter sees 92 points, they may start looking for confirmation of that high score. When they see 48, they might read the profile more critically. That is why explanations are needed alongside the numerical score, and users should be trained to treat the score as the starting point of their analysis, not the conclusion.


3. Why has AI Scoring become important right now?

Candidate scoring is not an invention of the generative AI era. Recruiters have been using scorecards, Excel spreadsheets, killer questions, weighted criteria and simple point systems for years. In many companies, the person managing the recruitment was already assigning points to candidates for experience, language skills, availability or knowledge of a specific technology. The problem was that such a process was time-consuming, highly dependent on user discipline, and often ended up as a paper-based methodology that nobody consistently used after receiving the hundredth application.

The shift today is not so much the existence of a point system itself, but rather the ability to automatically analyse massive amounts of unstructured text. Classic rules are good at dealing with 'yes' or 'no' answers, a specific number of years, or an exact certificate name. Modern language models can additionally compare the meaning of requirements with project descriptions, responsibilities and achievements written in different words. Thanks to this, scoring does not have to be limited to matching exact phrases. It can prepare a draft summary, highlight potential strengths and weaknesses, and suggest questions to help clarify ambiguities.

At the same time, the recruitment environment has changed dramatically. Easy applying, one-click apply forms, automatic alerts and generative tools that help build CVs mean that submitting an application takes much less effort than it did a few years ago. In many processes, the number of applications grows faster than a team's capacity to read them thoroughly. However, a larger volume does not mean more people who actually fit the role. Recruiters must find valuable profiles among documents that are random, mass-sent, linguistically very similar or tailored to match the job advert's keywords.

In such an environment, time is not the only issue. A human does not read the two-hundredth CV with the same freshness as the first. Concentration levels drop, pressure rises, and initial evaluation is increasingly based on a few of the most visible elements. A candidate who applied in the morning might get more attention than an equally good person whose document landed in the system at the end of the day. AI Scoring can help apply the same analytical framework to all submissions, but this does not guarantee automatic objectivity. If the criteria are flawed, the system will replicate the error more consistently than a human.

Scoring has therefore become important right now because three trends have converged: a growing volume of applications, increasing language analysis capabilities, and pressure for a faster, more measurable process. This does not mean every recruitment needs AI. It does mean, however, that teams must find a way to maintain evaluation quality in a scenario where manual document screening increasingly becomes the bottleneck.


AI Scoring is not the answer to a lack of recruitment strategy. It is an attempt to bring structure to evaluation where scale, pace and the volume of data make consistent human work difficult.


4. Numerical and contextual assessment - two dimensions of good scoring

The most visible element of scoring is usually the number. A score of 84/100 allows you to spot an application quickly, compare it with others and set the review order. It is particularly convenient when there are dozens or hundreds of profiles in a project. However, the number alone says surprisingly little. It does not explain which criteria were met, which ones carried the most weight, what was missing from the documents, or whether a lower score is due to a genuine mismatch or simply a lack of information.

That is why a good AI Scoring system should have two complementary dimensions. The first is numerical and structuring. It helps set the order of work and quickly flag profiles that require attention based on the adopted criteria. The second is contextual and explanatory. It shows what the system found in the data, how it linked that information to the role, which elements it considered strengths, which ones were weaknesses or unconfirmed, and what still needs checking.

In Recruitify, users get a score from 0 to 100, a candidate description in the context of the project, strengths, weaknesses and suggested interview questions. These elements are not just decoration for the percentage. They are precisely what allows the recruiter to judge whether the score makes sense. If someone gets 88 points but the justification does not confirm a critical must-have, the recruiter should spot the issue. If a profile has 58 points mainly because it lacks information on team scale or salary expectations, the correct response might be a quick screening question, not a rejection.

Contextual evaluation also helps reduce the anchoring effect. A large number displayed on the screen easily becomes the first and most dominant piece of information about a candidate. A human starts looking for arguments to back up the score instead of analysing the profile independently. When they see the sources of the score, unconfirmed areas and questions next to the number, it is much easier to treat the scoring as a working hypothesis rather than a final verdict.

However, contextual results must use cautious language. ‘The candidate managed a team of ten’ can be a fact derived from the CV. ‘The candidate is an excellent leader’ is a conclusion that the document does not prove. ‘No information on budget management was found in the available materials’ means a lack of evidence, not a proven lack of competence. A good tool and a well-trained user distinguish between facts, interpretations and information gaps.

It is possible to have purely numerical scoring, but it will lack transparency. It is also possible to have a purely descriptive evaluation, but with a high volume of applications, it is harder to use for prioritisation. The greatest value comes from combining both layers.


The number organises. The explanation allows you to evaluate. The question allows you to verify.


5. How does AI Scoring work step-by-step?

While different systems may use different models, algorithms and interfaces, the practical AI Scoring process can be presented as a sequence of several linked stages. Crucially, the evaluation does not start with the CV. It starts with defining the role and deciding what information matters for this specific recruitment.

1. The recruiter creates a project and describes the requirements. The system needs a point of reference. A job title alone or the full text of a job ad is usually not enough. Requirements should specify genuine must-haves, nice-to-haves, acceptable alternatives and questions that cannot be resolved solely on the basis of documents.

2. Criteria are checked and approved by a human. AI can help structure the description, but it should not decide on its own that exactly five years of experience, working at a specific company or graduating from a particular university is essential. The recruiter and the Hiring Manager must confirm that the criteria are job-related, proportionate and verifiable.

3. The system gathers available data about the application or candidate. In Recruitify, this can include information from the CV, LinkedIn profile and answers provided in the application form. The scope of data may vary between individuals, so the score must always be read in the context of the material's completeness.

4. Data is structured and mapped to the criteria. The system can recognise not only identical words but also semantic connections. A requirement for managing a sales team can be linked to a description of running a business development department. However, such a conclusion should remain auditable by the recruiter.

5. The AI distinguishes between what is confirmed and what is unknown. A good solution should not turn silence into a negative answer. A lack of information about Kubernetes is not the same as a declared lack of knowledge of Kubernetes. On the other hand, information from a form stating that a candidate can only start work in six months is a concrete signal that contradicts a requirement to start within a month.

6. A numerical score and contextual evaluation are generated. The system can award a score from 0 to 100, prepare a profile description, point out strengths and weaknesses, and suggest questions for further discussion. The result should be a trail of analysis, not an automated decision to advance or reject.

7. The recruiter checks the justification and source material. If a score seems too high or too low, the cause must be determined. The issue could lie in the requirements, data extraction from the CV, model interpretation or simply the incompleteness of the profile. A human must have the ability to make a different decision than suggested by the scoring order.

8. The result leads to action. The recruiter can start with high-match profiles, plan a call, ask a follow-up question or go back to the Hiring Manager to refine the brief. Scoring creates no value if it ends with a colourful number that nobody acts on.

9. The organisation monitors quality and impact. The team should track where scoring works well, when it generates inaccurate conclusions, how often recruiters disagree with the assessment, and whether lower scores are leading to valuable people being automatically overlooked. Implementation does not end when you toggle the feature on. That is when the responsible management of its impact begins.

In practice, this process may take a few seconds on the system's side, but its quality is the result of prior human work. The better the brief, the more adequate the data and the more conscious the user, the greater the utility of the result.


[SCREENSHOT: example of AI Scoring flow in Recruitify - project requirements, score 0-100, summary and questions]


6. What data can AI Scoring use?

The quality of the scoring depends not only on the AI model, but primarily on the quality, relevance and scope of the input data. In Recruitify, the evaluation can take into account information from the CV, LinkedIn profile and answers provided in the application form. These sources are not equivalent and should not be mechanically lumped together. Each shows a different piece of the profile and has its own limitations.

The CV usually provides the most structured professional history: job titles, companies, dates, responsibilities, education, skills and achievements. However, it is a selective document. The candidate decides what to fit on one or two pages, may shorten older experience, omit a project deemed less relevant, or use role titles specific to their previous organisation. A CV can also be outdated or generic, prepared without knowledge of a specific project's criteria.

A LinkedIn profile can contain a newer employment history, broader project descriptions, recommendations or skills omitted from the CV. However, it should not automatically be treated as more reliable. It is also created by the user, can be incomplete, outdated or written in a marketing style. Its value lies primarily in adding context and spotting discrepancies, not in serving as external proof of truth.

The application form allows you to collect information that is often missing from both the CV and LinkedIn. It can cover availability, salary expectations, preferred working model, right to work in a given country, willingness to relocate, required licences or experience in a very specific area. A candidate usually does not write in their CV whether they can start work in August, whether they accept two days a week in the office, or whether they have worked with a specific version of a system before. However, they can answer such questions in the form.

It is precisely the combination of sources that allows you to build a fuller picture. Imagine a candidate for a Finance Manager role. Her CV describes team leadership and month-end closing, LinkedIn shows a more recent promotion, and in the form, she confirms knowledge of a specific ERP system and availability in the required timeframe. Scoring based solely on the CV would miss the last two elements. At the same time, the system should not extrapolate anything that no source confirms.

Data should be interpreted according to three basic states: criterion confirmed, criterion probably unmet, and lack of sufficient information. This distinction is more important than the technical number of sources. Two outdated profiles do not create better proof than one fresh and specific answer.

It is also worth establishing a hierarchy of relevance and rules for resolving contradictions. What should the system do if a CV indicates employment ended in May, but LinkedIn shows ongoing employment? Should a form response submitted yesterday take precedence over a CV prepared six months ago? A good process does not hide such discrepancies under a uniform score. It signals them to the recruiter and turns them into questions for verification.

Combining multiple sources does not mean you should analyse everything that is technically available. The data minimisation principle requires using only the information necessary for a specific purpose. Age, photos, family situation, health status, origin or other sensitive areas should not affect scoring just because they can be found in a document or profile. More data does not always mean a better evaluation. Sometimes it just means more noise and a higher risk of biased conclusions.


The best scoring does not analyse the largest amount of data. It analyses data that is adequate, up-to-date and needed to evaluate the requirements of a specific project.


7. Why are project requirements more important than the AI model itself?

The most advanced model will not fix a poorly defined recruitment. If project requirements are vague, contradictory, unrealistic or contain biases, AI will only apply them faster and more consistently. It is precisely this consistency, often presented as an advantage of automation, that can become a problem: a human reading a few CVs might eventually notice that a criterion makes no sense. The system will keep repeating it for every person until someone changes it.

A good requirement should be specific, related to actual work and verifiable in the available data or during a later stage of the process. ‘A minimum of three years of experience in managing ERP implementation projects’ is more useful than ‘extensive project experience’. ‘Independently acquiring B2B clients with a contract value over €100k’ gives more context than ‘strong sales skills’. ‘Knowledge of English at a level allowing for negotiation’ describes a business need better than just ‘C1’, if the organisation does not intend to check the level with a formal test.

Requirements should not simply copy the entire job advert. An advert also serves a communication and employer branding function. It contains a description of the company, benefits, scope of responsibility and engaging language. Scoring, on the other hand, needs clear evaluation logic: what is a mandatory condition, what increases fit, what experience is equivalent, what can be learned after hiring, and what is absolutely non-negotiable.

The weight of criteria is also critical. Five years of experience is not always twice as good as two and a half years. A lack of one technology might be easy to catch up on, while a lack of a legally required licence actually prevents starting work. Scoring should reflect the business significance of criteria, not the number of words dedicated to them in the ad. If the ‘nice-to-have’ list is longer than the mandatory requirements, the system should not automatically allow the sum of minor additions to overshadow the lack of a core competence.

Before launching scoring, it is worth conducting a quick calibration with the Hiring Manager. Instead of only asking ‘who are we looking for?’, it is better to establish: what is this person supposed to deliver in the first six months, what experience genuinely increases the chance of success, which gaps are acceptable, and which ones make hiring impossible. Only on this basis can you create criteria that hold value for both AI and people.

8. Must-have, nice-to-have and disqualifying criteria

One of the most important stages of preparing scoring is separating three categories of requirements. Must-haves are conditions absolutely necessary to perform the job or start it within the set timeframe. Nice-to-haves increase fit, but their absence should not close the door to the process. Disqualifying criteria, however, are binary conditions where failure to meet them can justify automatically or semi-automatically directing an application to a separate pipeline - provided they are legal, proportionate and unambiguous.

In practice, organisations overuse the must-have category. A Hiring Manager might deem experience at a specific company, exactly five years of work, knowledge of an internal tool or graduation from a preferred university as essential, even though none of these are conditions for successfully performing the role. If AI Scoring receives such a list uncritically, it will reinforce a narrow 'ideal candidate' profile and lower the scores of people with diverse but highly valuable career paths.

Disqualifying criteria should be applied with extreme caution. They make sense where the answer is binary and directly related to the ability to do the work: a legally required licence, the right to work in a given location, availability during required hours, or willingness to work in a model the organisation cannot change. Do not turn elements requiring interpretation, such as 'good cultural fit', 'sufficient experience' or 'fitting personality', into killer questions.

AI Scoring should also not hide disqualifying rules within an opaque number. If a candidate received a low score due to the lack of a legally required certificate, the recruiter should see this clearly. If a form response indicates that a candidate cannot work the required hours, the system should distinguish this information from a lack of data. Transparency of the evaluation logic matters both for the quality of decisions and for the ability to explain the process to the candidate.

The best practice is to assign not only a category and weight to each criterion, but also a method of verification. Can the information come from the CV? Is a response in the form needed? Does it need to be confirmed during an interview? Does it require a document? This way, scoring does not pretend it can resolve everything based on a single source.

9. AI Scoring vs. keyword search

One of the most common simplifications in conversations about ATS is the belief that the system 'rejects CVs if they don't contain the right words'. In some older solutions, exact phrases, simple filters and manually built queries did play a large role. Keyword searching is still useful, but it is not the same as AI Scoring and should not be presented as its full mechanism of action.

Classic search checks for the presence or absence of specific expressions. If a recruiter types 'Salesforce', the system can find people who placed that name in their profile. The problem begins when the experience is described in a different language, the candidate uses a broader category name, or a specific competence is implied by context. A person who 'implemented and administered CRM solutions in a Salesforce environment' should be found easily. A harder case is a candidate describing 'managing a cloud-based sales automation platform' without mentioning the brand. A semantic system can recognise a potential connection, but it should treat it as a conclusion requiring confirmation, not as a certain fact.

AI Scoring based on a language model can match the meaning of requirements with the content of documents. This allows it to spot equivalent experience, role titles specific to an organisation, and skills described in a different order than in the job ad. For example, a requirement for 'managing a customer success team in a SaaS model' can be linked to a description of 'responsibility for an eight-person subscription customer care team'. Such reasoning is useful but not infallible. The more semantically distant the connection, the greater the need to show the recruiter exactly which part of the document the conclusion was based on.

We should not, therefore, pit keywords and AI against each other in an absolute way. Good tools can combine precise rules with contextual analysis. For a 'CISA' certification, the exact phrase matters. For experience in leading an organisational transformation, analysing the meaning is more important than identical words. The choice of mechanism should depend on the nature of the criterion.

It is also worth being cautious about promises that AI 'understands CVs like a human'. This is attractive from a marketing standpoint but is overreaching. A model can analyse text in a way that resembles human language processing, but it does not possess professional experience, intentions or full context. It can correctly link synonyms while completely misinterpreting the scope of responsibility. That is why semantisation increases capabilities but does not remove the need for human verification.


10. What does a score from 0 to 100 mean?

A score from 0 to 100 is primarily an auxiliary scale. It allows you to quickly see how the system evaluated the fit of available information against the project criteria. It is not a probability of hire, a prediction of job success, or a scientifically calculated 'candidate quality'. A score of 86 does not mean the candidate has an 86 per cent chance of succeeding in the role. Nor does it mean they are exactly 14 points better than someone with a score of 72.

The meaning of the number depends on the design of the specific system, how criteria are weighted, and the completeness of the data. In one project, scores might naturally cluster between 60 and 85, in another between 20 and 55 if the requirements are highly specialised. That is why thresholds like 'we reject everyone below 70' are dangerous if they have not been tested for a specific process. Even then, a threshold should not automatically replace human review, especially when the score affects access to employment.

The number can be most useful as a prioritisation tool. A recruiter can start their analysis with highly-rated individuals, but they should also review a sample of profiles from lower brackets. This kind of audit allows you to check whether the system is overlooking non-standard candidates, lowering scores due to a lack of information, or rewarding people who simply wrote a better CV. In practice, it is worth establishing a rule to regularly review the 'tail of the ranking', especially for new types of roles.

A good interface should not display the number in isolation from the justification. If the score is shown in a large font and the explanation is hidden under another click, users will make decisions based on the most visible element. Screen design influences behaviour just as much as the model itself. In a system supporting responsible decisions, the score, strengths, weaker areas and missing data should all be available together.

It is also worth remembering score stability. Generative models can, in some configurations, produce slightly different answers for the same input. The provider should limit this variance, test repeatability and clearly describe what might trigger a recalculation. The user, in turn, should know that changing project requirements or candidate data can legitimately change the score. The result is not a permanently written fact, but the outcome of a specific comparison at a specific time.


[SCREENSHOT: list of candidates or applications with AI Scoring from 0 to 100]


11. Strengths, weaknesses and interview questions

The greatest value of AI Scoring often lies not in the number, but in the contextual layer. A recruiter needs to know more than just that a profile was rated high or low. They need to know why, which elements matter for the project, and what they should do next. That is why a useful scoring tool can prepare a candidate description, highlight strengths and weaknesses, and suggest questions for the interview.

A candidate description should be a brief synthesis, not a summary of the entire CV. Its role is to connect the experience with the specific project. Instead of a generic 'experienced financial manager', it is better to point out that the person managed a team of similar size, was responsible for reporting to an international group and participated in an ERP system implementation, which is one of the challenges of the new role. This information helps the recruiter grasp the profile faster, but remains verifiable in the sources.

Strengths should relate to the criteria and stem from concrete data. It is not about generating compliments. 'Worked as a Sales Manager' is not a strength yet. 'Was responsible for independently acquiring enterprise clients in the DACH market for three years and achieved an annual target over €2m' is information that can be mapped to the project requirements.

Weaknesses require even more caution. In product language, we might call them weaknesses, but in practice, the system should separate at least three types of information. The first is a genuine contradiction with a requirement, such as lacking a legally required licence or availability being later than the project allows. The second is a probable mismatch, when the materials indicate experience significantly different from what is sought. The third is a lack of data, i.e., a situation where the documents simply do not allow a criterion to be resolved. These categories must not be treated the same way.

Suggested questions turn scoring from an organising mechanism into an interview preparation tool. If the CV lacks information on team scale, the system might suggest: 'How many people were in your team and what personnel decisions were you responsible for?'. If a candidate lists a technology without context: 'In which project did you use it and what was your level of independence?'. If LinkedIn and the CV differ on dates: 'Which information is current and what is the reason for the discrepancy?'. A good question is not meant to confirm a preconceived score. It is meant to supply the information that is still missing.

It is also worth avoiding questions that are too general, leading or about areas unrelated to the job. AI can generate a proposal, but the recruiter must check it and tailor it to the stage of the process. Asking about motivation during the first short screening might make sense. Assessing personality based on a stereotypical assumption does not.

The contextual layer has another benefit: it facilitates communication with the Hiring Manager or client. Instead of passing on just a number or a brief 'the profile fits', the recruiter can show arguments in favour, risk areas and a verification plan. This improves shortlists and shifts the discussion from gut feeling to concrete evidence.


[SCREENSHOT: candidate description, strengths, weaknesses and suggested questions in Recruitify]


12. Missing information is not the same as a lack of competence

This is one of the most critical principles of interpreting AI Scoring. Application documents are an incomplete description of a human being. A candidate might not list a skill because they took it for granted, wanted to limit CV length, tailored the document for a different role, or did not know the exact project requirements. A LinkedIn profile might be updated once every few years. An application form might contain a question that is too broad. Therefore, a lack of information is primarily a lack of evidence in the analysed material, not proof of a lack of competence.

The system should distinguish between at least three states:

Criterion confirmed - available data clearly indicates it is met.

Lack of sufficient information - the material does not allow confirming or ruling out that the criterion is met.

Criterion probably unmet or unmet - the data contains information that contradicts the requirement, e.g., the candidate declares availability in six months, and the project requires a start within four weeks.

Lumping the second and third categories together is a frequent source of unfair decisions. If a CV makes no mention of a specific technology, the system can write 'no confirmation in available data'. It should not, without additional basis, state 'the candidate does not know the technology'. If a candidate responds in a form that they have never used it, the situation is different, but even then, the weight of this information depends on whether the technology is a genuine must-have or a skill that can be quickly learned.

The style of the document also matters. A detailed CV provides more text and more potential evidence than the sparse profile of an expert who assumes company names and job titles are enough. As a result, the system may partially reward self-presentation quality. This is not a problem unique to AI, as recruiters also fall for well-written documents, but automation can scale this effect and give it a veneer of mathematical precision.

One way to mitigate this issue is to use the application form to gather critical information in a standardised way. If experience in a given area really matters, it is better to ask a specific question than to hope the candidate happened to describe it in their CV. The second way is generating interview questions. The third is regularly reviewing a portion of low-scoring profiles to check if the system mistook sparse documentation for a lack of qualifications.

This principle should also be part of user training. A recruiter must know that 'not found' does not mean 'does not possess', and 'the system suggests' does not mean 'the fact is confirmed'. Without this awareness, even well-designed explanations can be applied too categorically.


Missing information should lead primarily to a question. Not to an automated rejection.


13. AI Scoring vs. the traditional candidate scorecard

A traditional scorecard relies on pre-established criteria and a scale. A recruiter or Hiring Manager assigns points for experience, competencies, motivation or task results after an interview. A well-used scorecard increases process consistency because it forces evaluators to refer to shared criteria rather than a general impression.

AI Scoring should not replace this scorecard, but rather complement it at an earlier stage. It analyses the data available before the interview and helps prepare an initial assessment. The traditional scorecard can then cover information obtained during the interview, task, test or reference checks. Combining both tools makes sense when the criteria are consistent and it is clear which source can confirm which element.

The difference also relates to accountability. In a scorecard, a human directly assigns the rating. In AI Scoring, part of the interpretation is performed by the system, so the user must be able to trace the logic, check the basis and challenge the score. It is not enough to say that 'the algorithm calculated it'. The organisation remains responsible for how the evaluation is used in the process.

AI can improve the consistency of the initial analysis, but it does not remove human bias. Criteria are created by humans, data is created by humans, and results are interpreted by humans. What is more, the AI score can become a new source of bias - an automatic anchor that evaluators will adjust their own impressions to fit. Therefore, it is sometimes worth hiding the score from some individuals until they make their independent assessment, or at least requiring a brief justification for decisions independent of the score.

In practice, the best model might look like this: AI Scoring structures incoming applications and flags areas to check; the recruiter reviews the profile; the interview is conducted using standardised questions; the scorecard gathers evidence from the interview; the Hiring Manager makes a decision based on the entire set of materials. AI is then one element of the process, not an independent judge.

14. AI Scoring vs. AI Matching - a key distinction

AI Scoring and AI Matching are often lumped into the same category because both features match candidate profiles to project requirements. However, their goals are different, and in articles, product documentation and client communication, they should be consistently separated.

AI Scoring evaluates an application or candidate we are already reviewing within the context of a project. The starting point is a specific profile. The system prepares a score and explanation: to what extent the available information matches the criteria, what the strong and weaker areas are, and what needs verifying. It can apply to a person who applied themselves, or a candidate added to the project by a recruiter.

AI Matching works from the other side. The starting point is the project and the need to find potentially fitting individuals. The feature searches the candidate database, compares profiles with requirements, and suggests to the recruiter who they might want to consider. Its main task is discovery and internal sourcing: finding people the recruiter has not yet added to the project or might have forgotten about.

The simplest distinction is:

AI Scoring: ‘How does this candidate or application stack up against the project requirements?’

AI Matching: ‘Which candidates in the database might fit this project and should be suggested to the recruiter?’

In Recruitify, AI Scoring is an evaluative feature, whereas AI Matching is being developed as a separate mechanism for searching and suggesting candidates from the database. Scoring should not be presented as 'the next step after matching', as the process can run without matching. A candidate can apply via an ad, be added manually, or come from a referral, and then be evaluated by AI Scoring.

Both features can complement each other in the future. Matching finds potential profiles, the recruiter adds selected individuals to the project, and scoring prepares a more detailed assessment against the criteria. However, separate quality standards are still required. Matching should be measured by suggestion relevance and the ability to find valuable profiles. Scoring should be evaluated based on consistency, explainability and utility in analysing a specific person.

15. AI Scoring vs. CV parser, AI summary and killer questions

In a modern ATS, several features might work on similar data but perform completely different tasks. Confusing them leads to unrealistic expectations and makes risk assessment difficult. It is worth distinguishing AI Scoring from a CV parser, candidate summary, search, killer questions, workflow automation and the simple use of a public chatbot.

A CV parser reads a document and converts it into structured data: contact details, employment history, education, skills or languages. Its core task is extraction, not evaluation. A parser can make an error, such as assigning a date to the wrong company, and this error can later affect scoring. That is why extraction quality is one of the pillars of evaluation quality.

An AI summary creates a profile abstract. It can describe experience, key skills and career path, but it does not have to match them to project criteria or award a score. A summary primarily answers 'what is known about this candidate?', whereas scoring asks 'how does this information relate to the requirements of this recruitment?'. In practice, both elements may be presented together, but they are not the same.

Killer questions or disqualifying questions are rules based on specific answers. If a job requires a valid driving licence and the candidate answers 'no', the system can apply a pre-established path. This is not necessarily AI. It is rule-based automation whose logic was defined by a human. It has advantages in unambiguous cases, but can be more ruthless than scoring if a criterion was poorly chosen.

Workflow automation performs actions when a condition is met: sends a message, adds a label, changes a stage, creates a task or notifies the project owner. It can use the scoring result as one of the conditions, but such a connection requires extreme caution. Automatically rejecting everyone below a certain score turns a supporting feature into an actual decision-making mechanism and significantly increases legal and ethical risks.

A separate case is copying a CV and job description into a public chatbot with the prompt 'evaluate this candidate'. Technically, you can get a percentage, summary and questions this way, but it is not equivalent to secure AI Scoring embedded in an ATS. The recruiter may not know where the data goes, how long it is stored, whether it is used to train the service, who has access to the chat history, and whether the organisation has an appropriate agreement with the provider. It is also harder to maintain uniform criteria, log requirement versions, control permissions and document how the result influenced the process.

A feature available directly within the recruitment system should operate within an established environment of data processing, access control, retention, logs and provider agreements. The mere fact of integration with an ATS does not guarantee compliance or security, but it gives the organisation the ability to manage the process in a way that spontaneously copying data into a random tool usually fails to provide.

The most important question is therefore not 'can AI generate an evaluation?'. It can. The question is: is the evaluation generated based on approved criteria, on adequate data, in a controlled environment, and with the ability for a human to review it?

16. AI Scoring in practice - three assessment examples

Theory becomes clearer when we see how the same mechanism can work across different recruitment projects. The following examples do not present a ready-made numerical formula. Rather, they show how requirements, data and contextual interpretation should work together.

Example 1: CFO in a high-growth company

Let’s assume a company is looking for a CFO to prepare the organisation for rapid growth, structure controlling, manage funding and work with investors. A poorly described requirement might read: ‘minimum ten years of experience, prestigious company on the CV, strong leadership and strategic thinking’. Such a brief is difficult to score fairly. ‘Prestige’ is subjective, ‘strategic thinking’ does not stem directly from the document, and the number of years might not correlate with the challenge.

Better criteria relate to results and the working environment: responsibility for full P&L, experience in building controlling structures, involvement in securing funding, managing a finance team, reporting to the board or investors, and working in an organisation of a similar scale of growth. The CV can confirm some of these elements. LinkedIn can show a more recent role. The application form can collect availability, expectations and the answer to a question about the largest funding round managed.

The candidate might get a high score because they ran finance in a company growing from 50 to 300 people, implemented controlling and participated in talks with a fund. A weakness or risk area might be the lack of confirmation of independent responsibility for securing funding. Rather than assuming this competence is missing, the system can suggest a question: ‘What was your personal scope of responsibility in the funding process and which elements did you manage independently?’.

Example 2: Java Developer in a migration project

In the second project, a company is seeking a Java Developer to migrate a legacy application to a new architecture. Must-haves might include commercial experience with Java, working with Spring Boot, relational databases and distributed systems. Nice-to-haves are Kubernetes, public cloud and experience in gradually modernising a monolith.

The candidate lists Java and Spring, but does not use the exact phrase 'microservices' in the CV. They do, however, describe breaking down a monolithic application into independently deployable modules communicating via API. An analysis based solely on keywords might miss this important signal. Semantic analysis can spot the connection, but it should point out the fragment on which the conclusion was based.

There is no mention of Kubernetes in the CV. This should not automatically mean a weakness in the sense of a lack of competence. If Kubernetes is a nice-to-have, the system can mark a lack of confirmation and suggest a question: ‘Have you deployed applications in a containerised environment and what orchestration tools did you use?’. If the candidate responds in the form that they worked with Kubernetes for two years, the information can complete the profile. If they answer 'no', scoring should factor this in as a genuine gap, but in line with the criterion's weight, not as an automatic rejection.

Example 3: B2B salesperson for a new market

The third project concerns a person to develop B2B sales in the German market. The title 'Sales Manager' alone says very little. For the role, independent acquisition of new clients, sales cycle length, contract value, working with decision-makers, knowledge of the DACH market and German language skills allowing for negotiations could all be significant.

Candidate A has impressive sales results, but most of their experience comes from managing existing clients in the Polish market. Candidate B has lower revenues but opened the German market independently, built a pipeline from scratch and led enterprise negotiations. A simple scoring system based on the sum of sales results might reward Candidate A. Scoring anchored in the genuine requirements of the project should notice that the second profile is a better fit for the specific task.

The application form can additionally collect information on willingness to travel, salary expectations and actual language level. The system can highlight market building experience as a strength, and the lack of information on average contract value as an area for verification. An interview question could be: ‘Tell us about the largest client you acquired independently in the DACH market - what did the process look like from first contact to signing the contract?’.

These three examples show why there is no single universal scoring for a 'good candidate'. A CFO, a Java Developer and a salesperson are evaluated against entirely different criteria. Even two candidates with similar job titles might score differently depending on the project's goal. AI can help navigate requirements consistently, but it does not relieve the organisation of the need to answer the hardest question: what do we actually need from a person in this role?

17. Benefits for the recruiter and the recruitment team

The most obvious benefit is time savings, but the value of AI Scoring should not be reduced to reading CVs faster. A well-used feature changes how the initial analysis is organised. It helps spot applications meeting multiple criteria faster, flag areas requiring checks, and prepare a more structured interview. The recruiter still reviews profiles, but they do not start with a blank page every time.

Savings also do not mean a human stops reading documents altogether. It is about a different distribution of attention. Instead of dedicating the same amount of time to every part of every CV, the recruiter can immediately see which requirements have been confirmed, where there is a contradiction and what is missing. Thanks to this, they spend more time on interpretation, contact and verification, and less on manually finding scattered information.

The second benefit is greater consistency. With hundreds of applications, people naturally differ in their attention levels and interpretation of criteria. The system can apply the same analytical framework to every profile. This does not automatically mean greater fairness, because flawed rules can also be applied consistently. However, if criteria are well-prepared and results are monitored, consistency reduces the randomness of the initial review and makes comparing decisions within the team easier.

The third benefit is better justification. The recruiter can quickly prepare arguments for the Hiring Manager or client, instead of sending just the CV and a generic 'worth a talk'. Strengths, weaknesses, information gaps and questions for verification create a shared language of evaluation. This does not mean the generated description should be copied without checks. Its value lies in preparing a draft analysis that a human reviews and refines.

The fourth benefit relates to interview preparation. The recruiter can walk into a meeting with a list of specific areas to dig into, rather than asking everyone the same set of general questions. This allows better use of the candidate's time and faster separation of actual experience from an attractive description.

The fifth benefit is linked to teamwork and onboarding new people. If every consultant uses their own informal way of reading CVs, it is hard to hand over projects, analyse quality and train less experienced recruiters. AI Scoring can introduce a common point. The team can discuss why they disagree with a score, which requirement was poorly written and how to improve the brief. The difference between human and system evaluation itself becomes learning material.

The sixth benefit is faster response times. In an agency, a few hours can decide which firm presents a valuable candidate first. In in-house recruitment, quick contact affects candidate experience and reduces the risk of losing someone to another process. Scoring can shorten the time from application to the first meaningful analysis, but only when embedded in the workflow and leading to action. If it remains just another number in the system, it brings no real value.

18. Benefits for agencies, in-house HR and Hiring Managers

In a recruitment agency, AI Scoring can primarily support delivery. A recruiter often runs several projects simultaneously, and clients' requirements can be incomplete and change during the process. Contextual evaluation helps grasp a profile faster, prepare the screening and build arguments for the shortlist. It can also facilitate project handovers between consultants. A new person does not have to reconstruct all the logic solely from notes if they can see the approved criteria and how candidates map to them.

The speed of working with your own database is also key. Scoring is not AI Matching and does not automatically search for the best people for a project on its own, but it can evaluate candidates whom the recruiter added from previous processes, referrals or sourcing. Thanks to this, the database stops being purely a CV archive. The condition is data relevance and the legal basis for its further use.

In in-house recruitment, the greatest value often appears with high volumes, but is not limited to them. Even with dozens of applications, the tool can help standardise the initial evaluation across locations and recruiters. It can also reveal a briefing problem. If most candidates receive low scores due to a single requirement, it is worth checking whether the market simply does not supply such profiles, or if the criterion is too narrow or poorly communicated.

The Hiring Manager gets more structured material. Instead of reading every CV without context, they can receive a summary, pros and cons, and areas to check. However, they should not see only the AI score. A large number without access to the document and justification can lead to over-reliance. The best model ensures easy access to both the contextual analysis and the full profile, along with the recruiter's decisions.

AI Scoring can also improve the quality of the conversation between HR and the business. Instead of discussions like 'I like this candidate' versus 'this candidate doesn't fit', you have a list of requirements, evidence, gaps and questions. This does not remove subjectivity, but makes it more visible. If a Hiring Manager rejects a high-match person, it is worth asking which criterion was not written down before. If they accept a low-scoring profile, perhaps the system misinterpreted the data or the requirements needed adjustment.

For the organisation, the benefit is better process analysis. You can check whether high scores translate into advancing through subsequent stages, where recruiters most often disagree with the evaluation, which requirements are regularly unconfirmed, and whether lower-bracket candidates are being unfairly overlooked. However, you must be careful not to treat historical decisions as objective truth. If the past process was biased, blindly learning from its results can entrench the problem.

19. Benefits and risks from the candidate’s perspective

From the candidate’s perspective, AI Scoring can bring real benefits. Every application can be subjected to the same analytical framework, even if it arrived as the two-hundredth. The system can notice experience described differently than in the ad, combine form and CV details, and suggest a question to the recruiter that gives the candidate a chance to explain a missing element. Faster pre-selection can also mean a quicker response and shorter waiting times.

However, benefits are not automatic. A candidate might be unfairly downgraded by an incomplete CV, an unconventional career path, a career change, a gap in employment, or language that differs from the ad's standards. They may not know their data is being analysed by AI, fail to understand the role of the score in the decision, and have no easy way to challenge an error. If an organisation automatically rejects people below a threshold, a single misinterpretation can have a significant impact.

Candidate experience therefore depends on transparency and usage. The organisation should clearly state that AI tools are used in the process, describe their role and highlight whether the decision is made by a human. Candidates should have the opportunity to correct inaccurate data and obtain information on how their data is used in compliance with applicable law. In processes with a higher impact, it is worth providing an appeal channel or human review option.

We should also not shift responsibility entirely onto candidates by stating 'if they didn't write it, it's their own fault'. A CV is a communication tool, but it is the organisation that designs the process and decides what data is needed. If a criterion is key, it should be communicated and gathered in a proportionate way. Scoring can support a fairer evaluation only when the entire process is designed fairly.

20. Where does AI Scoring give the most value?

AI Scoring does not deliver the same value in every process. The most obvious application is recruitment with a high volume of applications, where reading every document manually with equal attention is difficult. This includes popular junior roles, remote positions, international recruitment and processes promoted on portals that enable quick applying. In these scenarios, scoring can help establish the review order and flag profiles requiring urgent contact.

The second application is projects with clearly defined, verifiable requirements. If a role requires specific technologies, licences, experience in a specific environment, sales model or operational scale, AI can compare these elements with documents. The more vague and subjective the criteria, the lower the value of automatic scoring. Scoring is much better at answering 'did the candidate run ERP system implementations in organisations over 500 people?' than 'do they have leadership charisma?'.

The third area is reusing the candidate database. A recruiter can add people known from previous processes to a new project and evaluate their profiles against new requirements. However, you must check data relevance and the legal basis for further processing. Scoring on an outdated CV can create a seemingly precise but low-value result.

The fourth application is agency recruitment, where the speed of preparing a shortlist and the quality of candidate presentation matter. Contextual evaluation can help the consultant prepare the screening, spot risks and build arguments for the client. However, it should not be copied directly into the candidate profile without checks. AI-generated phrasing can be too categorical, imprecise or inappropriate for an external recipient.

Scoring may have less value in executive search processes, where candidate numbers are small and key information comes from interviews, market reputation, references and understanding the complex context of the organisation. Even there, it can support structuring data, but it is hard to expect a CV score to capture the ability to lead transformation, build board trust or operate in a specific ownership culture.

Scoring should not be used just because the feature is available. Before implementing, it is worth answering three questions: what specific problem are we solving, how will we check score quality, and what should the user do after receiving it. If there is no clear answer, another number in the ATS can increase complexity instead of productivity.

21. How to implement AI Scoring in your recruitment process?

A good implementation starts with the process, not a button in the system. The first step should be selecting a limited group of roles for which requirements are relatively clear and a sufficient number of applications are available for testing. Running a pilot across all projects simultaneously makes it hard to identify what works and what needs refinement.

Next, you need to define the rules for creating requirements. Who prepares the criteria, who approves them, how are mandatory and nice-to-have conditions distinguished, how often can they be changed, and what happens to scores after a change? It is worth preparing a simple brief template. It should require describing the goal of the role, key results, essential competencies, acceptable alternatives and areas to verify during the interview.

The next step is testing on known profiles. The team can choose sample CVs of candidates they evaluated previously, run the scoring, and compare the result with recruiters' independent assessments. The goal is not for the system to perfectly replicate every historical decision. History can contain errors. The goal is to check whether justifications are logical, whether the system recognises key experiences, how it treats missing data, and whether undesirable patterns emerge.

User training is essential. It should cover not only feature usage, but also score interpretation, automation bias, the difference between missing information and lack of competence, data protection rules and escalation scenarios. A recruiter should know when they can use scoring for prioritisation and when they must not rely on it without additional analysis.

It is worth establishing a human review practice. For instance, you could require that before rejecting candidates from a specific bracket, the recruiter must review the document and note their own reason for the decision. You can also randomly audit a portion of low-scoring applications. It is important that oversight is real, not formal. Clicking 'approve' without reading the profile is not meaningful human involvement.

Finally, you need to define monitoring. Monthly or quarterly, the team can analyse score distribution, differences between roles, score alignment with recruiter decisions, the number of corrections, profiles bypassed by scoring, and candidate feedback. The implementation should have an owner responsible for both efficiency and risk.

22. Most common mistakes when using AI Scoring

The first mistake is running scoring on a weak brief. A generic job title and a copied job ad do not form an evaluation logic. If the Hiring Manager cannot specify which requirements are truly important, AI will not establish this reliably either. It can suggest criteria, but a human must verify them.

The second mistake is equating a high score with a hiring recommendation. Scoring relates to data matching against criteria at a specific stage. It does not cover everything that drives job success. A candidate might have perfectly described experience but low motivation, fail to confirm declared skills, or reject the offer terms.

The third mistake is automatic rejection based on a threshold. This setup is tempting because it maximises time savings, but it can lead to systematically overlooking people with incomplete CVs and increases the impact of a single error. It can also turn a supporting tool into a system actually making decisions, which has legal consequences.

The fourth mistake is failing to distinguish between data and conclusions. If an AI summary sounds convincing, the user might forget it is an interpretation. Every significant conclusion should be linkable to a document fragment or marked as requiring confirmation.

The fifth mistake is analysing too broad a scope of data. More data is not always better. Photos, age, address, marital status, health information or social media activity can be unnecessary and risky. The organisation should consciously define which sources and fields can affect scoring.

The sixth mistake is a lack of testing after changing a model, prompt or logic. An AI feature might behave differently after an update, even if the interface looks the same. The provider should manage changes, and the client should know when scores might shift.

The seventh mistake is assuming a human will automatically correct system issues. Research on human-AI collaboration shows that people can over-rely on recommendations and align their own decisions with the score. Oversight requires competence, time and the right to disagree with the system, not just a recruiter's name in the process.

23. Limitations of AI Scoring - what does the system not know?

AI Scoring sees the information it has been given, and nothing more. It does not know competencies omitted from the CV, does not know if an achievement was described accurately, does not independently confirm declaration truthfulness, and does not observe the candidate at work. It can draw conclusions from text, but should not replace skills tests, behavioural interviews, references or evaluations performed in a real context.

The system may struggle with unconventional career paths. A person transitioning from the military to business, from academia to product, from entrepreneurship to corporate, or from another industry may possess transferable skills that they do not describe in the target role's language. A semantic model can spot some of these, but requirements based on historical job titles might still downgrade the score.

Soft skills are also difficult. A CV might state 'excellent communication skills', but this is not reliable evidence. The system should not award a high score for the declaration alone. It can, however, point to experiences where communication was likely important and suggest interview questions. Assessing actual behaviour, however, requires other methods.

AI also does not know the full context of the organisation. Two companies might use the same job title for completely different scopes. Managing a team of ten in a stable department is different from building a team from scratch in a startup. A budget of one million in one industry might have a different meaning than in another. Therefore, requirements must describe context, not just labels.

Data and document errors are another limitation. A parser can misread a table, the system can mix up dates, and a LinkedIn profile can be outdated. The more precise the numerical score, the easier it is to forget about input uncertainty. The 'garbage in, garbage out' rule remains relevant in the era of language models.

Finally, AI does not bear responsibility for the decision. It can prepare an assessment, but it does not talk to the candidate, does not know all the circumstances, and is not accountable to the organisation, client or regulator. Responsibility remains with the people and entities that design and use the system.

24. Bias, discrimination and the illusion of objectivity

One of the most tempting promises of AI in recruitment is objectivity. A machine is not tired, does not have a favourite university, and does not evaluate a photo like a human - provided the photo and other unnecessary data do not influence the model. However, this does not mean the score is neutral. The system can replicate biases present in the data, criteria, labelling method or the process design itself.

Bias can enter scoring at multiple levels. Criteria might reward paths more accessible to specific groups, such as an unjustified requirement for continuous employment. Historical data might reflect the organisation's past preferences. Company and university names can act as proxy indicators of social status. CV language can differ by culture and gender. Missing data might more often affect people with less conventional career histories.

Language models can also draw conclusions based on indirect traits. Research on CV evaluation by LLMs shows that changing elements suggesting gender or origin can affect the score, though the scale and direction of bias vary between models. The FAIRE benchmark of 2025 demonstrated the presence of some level of bias in tested models during direct CV scoring and ranking [7]. This does not mean every recruitment system will behave identically, but it confirms the need to test specific configurations rather than relying on general model manufacturer claims.

The illusion of objectivity is particularly dangerous because numbers look scientific. A score of 78.4 can give the impression of a measurement similar to temperature, though it is the result of chosen criteria and text interpretation. Precision in display is not proof of accuracy. The organisation should ask: what does the score measure, how was it tested, for which groups does it work less well, and what are the consequences of an error.

Mitigating bias is not just about removing names and photos. Anonymisation helps, but other elements can still reveal or indirectly indicate protected traits. What is needed is score testing, analysing differences between groups where legal and methodologically justified, reviewing criteria, error reporting mechanisms and the ability to withdraw a feature if risk cannot be mitigated.

It is also worth remembering that humans are not a neutral benchmark. AI can, in some situations, limit recruiter randomness and unconscious preferences. The goal should not be to prove that 'AI is more objective than a human' or vice versa. The goal is to design a process where both parties' errors are visible, measured and correctable.

25. Human-in-the-loop - what does real human oversight mean?

The phrase human-in-the-loop appears in almost every description of responsible AI. Simply placing a human in the process is not enough. If a recruiter sees a score, automatically accepts the recommendation and has no time for analysis, their involvement is formal. Real oversight means the ability to understand the system's capabilities and limitations, check the basis of a score, challenge it, and halt or change actions when an issue arises.

The AI Act, in relation to high-risk systems, requires designing solutions so they can be effectively overseen by natural persons. It also highlights the risk of automation bias, i.e., automatic or excessive reliance on system output [4]. The person exercising oversight should have appropriate competence, training, authority and support. In a recruitment context, this means a junior recruiter cannot be the only 'safety layer' if the organisation expects them to process massive volumes and rates them solely on speed.

Real oversight can involve several practices. The recruiter should see the justification and source material. They should be able to change a decision without negative organisational consequences. The system should log corrections and enable analysis of where humans most often disagree with the AI. For automated actions, a kill switch mechanism is needed. Candidates should be able to report inaccuracies, and the organisation needs a procedure to handle such reports.

It is worth separating human review from doing all the work manually. Oversight does not mean a human has to re-analyse every element from scratch, because then the tool loses its purpose. They should, however, check decisive issues, especially before a negative decision. A proportionate approach can be applied: more automation in low-risk administrative tasks, more control where the result affects the candidate's access to the next stage.

The organisation should also train users to recognise situations where the system might fail: unconventional CVs, industry changes, documents in another language, employment gaps, incomplete data, conflicting sources or very narrow criteria. Human-in-the-loop is only effective when the human knows what to look for.

26. AI Scoring and GDPR

The AI Act does not replace GDPR. If AI Scoring processes candidates' personal data, the organisation must still meet obligations under data protection laws. This covers legal basis, transparency, data minimisation, purpose limitation, accuracy of information, security, retention and the exercise of data subject rights.

The first question is: who is the controller and who is the processor. The company running the recruitment usually determines the purposes and means of processing candidates' data, and the ATS provider acts in a specific capacity as a processor. However, the actual split depends on the service design, including whether the provider or its sub-processors use the data for their own purposes, model training, testing or service improvement. Roles, instructions and sub-processors should be clearly described in the contract.

The second question relates to purpose and legal basis. You cannot assume that candidate consent solves every problem. The basis depends on the type of recruitment, stage, national law and whether data is to be used in future projects as well. Special categories of data must be analysed separately. If a model can technically infer origin, health status, views or other sensitive traits, this does not mean the organisation can use such inferences in scoring.

The third area is minimisation. The system should only analyse information relevant to the purpose. Photos, age, home address, family situation or activity unrelated to work should not affect the assessment just because they are in the document or profile. Particular caution is needed when using data from LinkedIn and other external sources. The availability of information on the internet does not automatically mean free rein to process it.

The fourth area is the information obligation. Candidates should receive clear information that their data is analysed using AI, for what purpose, what data categories are used, where they come from, who they are disclosed to, how long they are stored, and what rights they have. Information should not hide key facts under a generic phrase like 'we use modern technologies'. It is also worth explaining the role of the score: whether it only supports the recruiter's work or triggers further actions.

Article 22 of the GDPR is of particular significance, as it concerns decisions based solely on automated processing which produce legal effects or similarly significantly affect individuals [2]. Scoring that supports a recruiter does not automatically mean a solely automated decision. However, if the score leads, without real human involvement, to rejecting an application, blocking the next stage or hiding a profile from the recruiter, the risk of entering this area is much higher. The 'human in the loop' must have a real impact, not just a technical option to approve the system's decision.

The fifth area is data accuracy and the ability to correct it. If a parser misreads a date, LinkedIn is outdated or a form contains an incorrect answer, the score might be wrong. Candidates should be able to correct data, and the organisation should establish whether scoring is recalculated after a correction.

Before implementation, it is worth conducting a Data Protection Impact Assessment (DPIA), especially when technology systematically evaluates individuals and can significantly affect their situation. The UK’s ICO is not an EU law-applying body, but its audits of recruitment tools serve as a practical reference. The regulator highlighted, among other things, the need to perform a DPIA at the procurement stage, clearly divide responsibilities, limit data, test fairness and transparently inform candidates [6].

Retention is also important. Data and scoring results should not be stored indefinitely just because a large database might be useful someday. The organisation should define how long it stores CVs, answers, scores, justifications and logs, and what happens to them after consent is withdrawn or the processing purpose expires.


Legal note: this section is for informational purposes and does not replace a legal analysis of a specific implementation. The scope of obligations depends on how the system operates and is used, the data types, the roles of the parties and national regulations.


27. AI Scoring and the AI Act

The EU AI Act applies a risk-based approach. Annex III lists AI systems intended to be used for recruitment or selection of persons, notably to place job adverts, screen and filter applications, and evaluate candidates [1]. This means AI Scoring used in recruitment is in an area of high regulatory focus. However, we should not simplify this to state that every feature containing AI and CVs automatically has the same legal status.

Classification depends on the intended purpose, design and actual impact on the process. Article 6(3) provides a possibility to deem some systems listed in Annex III as non-high-risk if they do not pose a significant risk of harm to the health, safety or fundamental rights of natural persons, including by not materially influencing the outcome of decision-making, and meet specific conditions. However, a system remains high-risk if it profiles natural persons, and a provider who considers that a system is not high-risk must document its assessment. The Commission published draft detailed guidelines on this classification in 2026 [3].

In practice, a feature that assigns a score to candidates and materially affects review order, advancement to the next stage or rejection requires highly cautious analysis. You cannot resolve classification solely based on names like 'assistant', 'recommendation' or 'scoring'. What matters is the intended use described by the provider and how the organisation actually uses the score.

For systems classified as high-risk, the AI Act outlines extensive requirements for providers. These include a risk management system, quality and data governance, technical documentation, automatic logging of events, transparency for the user, human oversight capabilities, and appropriate levels of accuracy, robustness and cybersecurity, alongside post-market monitoring. Depending on the scenario, a conformity assessment, declaration of conformity, CE marking and registration in the EU database are also required.

Obligations also apply to entities using the system, i.e., deployers. The recruiting organisation should use the solution in accordance with instructions, assign competent persons for oversight, ensure the relevance of input data within its control, monitor operation, and react to issues. In specific situations, individuals subject to a high-risk system must be informed of its use and the type of decisions supported [1]. The AI Act also provides a right to a clear and meaningful explanation of the role of the AI and the main elements of the decision in situations outlined in Article 86.

Human oversight cannot be formal. The person overseeing should understand the system's capabilities and limitations, be aware of automation bias, have access to sufficient information, and possess genuine authority to bypass, reverse or stop operations. A recruiter who automatically accepts every suggestion does not provide real oversight just because they formally clicked a button.

AI literacy is also significant. Organisations should ensure that individuals using AI possess an adequate level of knowledge and competence, taking into account their role, experience, context of use and the persons affected by the system. In practice, recruiter training should cover not only feature usage, but also score interpretation, distinguishing lack of data from lack of competence, discrimination risks, data protection and error reporting procedures [5].

Not every private company using a recruitment system will be required to carry out a fundamental rights impact assessment under Article 27. This obligation applies to specific categories of deployers and uses. Nevertheless, an impact assessment can be a good practice, and a DPIA under GDPR may be required regardless of the AI Act.

The timeline is important. According to current Commission information, following amendments introduced by the AI Omnibus, high-risk system requirements in Annex III, covering employment among others, are to apply from 2 December 2027 [4]. This does not mean organisations can ignore the topic until late 2027. Obligations regarding AI literacy apply from 2 February 2025, and oversight of them starts according to the Commission's timeline in 2026 [5]. GDPR, anti-discrimination laws and employment law apply regardless of this timeline right now.

For organisations, the practical takeaway is simple: do not wait until the last minute. It is worth creating a registry of used AI features, defining their purpose, providers, data sources, impact on decisions, oversight mechanisms and issue reporting procedures. Providers should be asked about classification, documentation, tests, logs, sub-processors, processing locations, model changes and compliance plans. The AI Act is not solely a legal team problem. It impacts product design, procurement, HR, security and the daily work of recruiters.


The safest assumption is: the greater the impact of the score on a candidate's chance of staying in the process, the stronger the documentation, human control and monitoring should be.


28. How to choose an AI Scoring tool?

Comparing tools solely based on demonstration quality is risky. On a demo, you can select perfectly written CVs and roles for which the score looks convincing. Real quality is revealed with incomplete, unusual, multilingual, conflicting documents coming from different industries. Therefore, purchase should involve testing on your own, properly secured data or a prepared set of representative profiles.

The first group of questions relates to features. What is evaluated: the application, the candidate, or both? How are requirements defined? Can the user edit and weight them? How does the system distinguish missing data from unmet criteria? Does it show the justification and source fragments? Does it generate questions? Can the score be recalculated after data changes? Does scoring work in multiple languages?

The second group relates to control. Can automated actions be turned off? Does the system log who triggered the evaluation and which requirement version was used? Can the user report an error? Does the provider inform about model changes? Can the administrator restrict access to the feature and set usage rules?

The third group relates to quality. How does the provider test accuracy, stability and bias? On what roles and languages? Do they share the methodology and known limitations? How do they react to errors? Do they measure differences between groups in a legally compliant way? Does the model generate deterministic or sufficiently repeatable answers?

The fourth group relates to data and security. Where is data processed and stored? Is it sent to third-party model providers? Is it used to train public or proprietary models? How long is it kept in logs? What are the bases for transfers outside the EEA? What do encryption, access control, audit and data deletion look like? Who is the sub-processor?

The fifth group relates to law. How does the provider classify the feature under the AI Act and on what basis? What is the compliance plan? What instructions will the deployer receive? Will documents needed for a DPIA, risk assessment and information obligations be available? Does the contract clearly define roles in data protection?

The sixth group relates to implementation. Does the provider train the team not only on usage, but on responsible interpretation? Do they help calibrate criteria? Do they allow starting with a pilot? Can the client measure effects? The best algorithm won't help if users don't trust the system or trust it too much.

29. How to measure the effectiveness of AI Scoring?

Implementing AI Scoring should have a measurable goal. The simplest indicators relate to efficiency: time from application to first analysis, average profile review time, number of applications analysed in a given timeframe and time to first contact. This data shows whether the tool actually relieves the team.

The second group of indicators relates to utility. You can measure how often the recruiter agrees with the score, how often they correct it, whether they use suggested questions, and whether summaries reduce the need for re-reading. However, agreement with a human is not a measure of truth on its own. It is worth analysing justifications and cases of discrepancy, not just the agreement percentage.

The third group relates to process quality. Do candidates with high scores more often pass screening after competency verification? How many valuable candidates were found in lower brackets? Does the score differ significantly between recruiters using the same criteria? Are fewer candidates waiting long for a response?

The fourth group relates to risk. Complaints, data errors, inaccurate explanations, unequal score distributions and automated actions with negative outcomes should be monitored. Where possible and in compliance with the law, fairness tests should be conducted. It is not enough to check the system once before implementation. Roles, data, models and how users work change.

The fifth group relates to human behaviour. Are recruiters reviewing lower-scored profiles? Are they rejecting candidates faster but without justification? Does the Hiring Manager look at the score first, and only then the CV? You can analyse logs and conduct short qualitative audits. The goal is to detect automation bias and situations where the feature takes on a larger role than planned.

Success should not be evaluated solely by the number of hires with high scores. If recruiters primarily contact this group, such a result is partly a self-fulfilling prophecy. Comparisons, control samples and conscious analysis of candidates outside the top positions are needed.

30. How does AI Scoring work in Recruitify?

AI Scoring in Recruitify is used to evaluate an application or candidate in the context of a specific recruitment project. The feature does not assign a permanent value to a person and does not automatically search for individuals in the database. The starting point is a profile already within the project, and the requirements defined for that recruitment.

The process begins with the Requirements field. The recruiter describes what the project actually requires: what experience is essential, what elements increase fit, which gaps are acceptable, and what should be verified during the interview. The more specific and job-related the requirements, the more useful the assessment will be. Generic phrases like 'good communicator', 'dynamic person' or 'cultural fit' do not give the AI or recruiter a sufficiently precise point of reference.

Recruitify analyses available information from the CV, LinkedIn profile and answers provided in the application form. Thanks to this, the score does not have to rely on a single document. The form can provide data on availability, expectations, working model, licences or experience in an area the candidate did not describe in their CV. LinkedIn can complement history, but is not treated as automatically more reliable than the application document. Discrepancies should lead to verification.

The result has both a numerical and contextual dimension. Recruitify awards a score from 0 to 100, but does not limit itself to the number. The user receives a candidate description in the context of the role, strengths, weaknesses and suggestions for questions to ask during the meeting. The goal is to help prioritise work, grasp the profile faster and prepare the screening better.

A score from 0 to 100 is not a probability of hire. A candidate with a score of 90 is not 'ten per cent better' than someone with a score of 80. The number shows the outcome of comparing available data against the criteria of a specific project. The exact same candidate can receive a different score across two recruitments because requirements change. The score can also change after data updates or refining the Requirements.

After generating the assessment, the recruiter should read the justification and check it against the source material. Particular attention should be paid to weaknesses and areas the system could not confirm. A lack of mention in the CV does not have to mean a lack of competence. In many cases, the best next action is to use a suggested question or ask the candidate to supplement information.

Recruitify clearly separates AI Scoring from AI Matching. Scoring evaluates an application or candidate already being analysed in a project. AI Matching, which is under development, is designed to search the database and suggest individuals potentially fitting the project to the recruiter. These are two different tasks: evaluating a specific profile versus finding and recommending profiles from the database.

The best practice is to start with tests on a few projects with clear requirements. The team can compare scores with their own evaluation, check justification quality, review a portion of low-scoring profiles, and refine how Requirements are written. AI Scoring is meant to support the recruiter's decision, not automatically reject candidates.





31. The future of AI Scoring

AI Scoring will shift from a simple percentage towards more complex, but also more controlled, decision support. The most valuable solutions will not try to create a single 'truth about the candidate'. They will combine different sources, show uncertainty, point out the basis of conclusions and tailor the evaluation method to the stage of the process.

We can expect a greater role for structured rubrics. Instead of a general prompt like 'evaluate the candidate', the system will work on criteria with a definition, weight, source of proof and level of certainty. Development may involve comparing candidates, but should avoid situations where a difference of a few points is treated as an objective advantage of one human over another.

The second direction will be integration with subsequent stages. Scoring before the interview can generate questions, and after the interview, the system can help organise notes according to the same scorecard. It will be important to separate declared from confirmed information and ensure that interview data is processed in compliance with the law.

The third direction will be testing and auditability. Regulations, client requirements and market maturity will increase pressure to document models, logic, data, accuracy and bias. Providers will have to show not just an attractive interface, but also a risk and change management process.

The fourth direction will be combining scoring with AI Matching and talent intelligence. The system can first suggest profiles from the database, then evaluate them in the project, and later support communication and process analysis. This combination increases productivity but also the risk of creating a closed loop of automated recommendations. Every stage should have a clear goal, separate metrics and the possibility of human intervention.

The most significant shift, however, will be organisational. Companies will stop asking simply 'do we have an AI feature?', and start asking 'what decision does this feature support, on what data, with what risk, and who is responsible for control?'. It is the answers to these questions that will determine whether AI Scoring improves recruitment or merely gives old problems a modern interface.

32. AI Scoring FAQ - Frequently Asked Questions

What is AI Candidate Scoring?

AI Scoring is an AI-supported evaluation of a specific application or candidate against the requirements of a specific project. It can take the form of a numerical score, a contextual description, or both elements simultaneously. It should not be understood as a universal assessment of human worth.

Does AI Scoring evaluate the candidate or the application?

It can evaluate both. An application is a specific entry to a project, often alongside form responses. A candidate can be added to the project by a recruiter without a new application. In both cases, the result should relate to the fit with the given project's requirements.

Are AI Scoring and AI Matching the same thing?

No. AI Scoring evaluates an application or candidate already being analysed in a project. AI Matching searches the database and suggests candidates potentially fitting the project. The first feature answers 'how does this profile stack up?', and the second 'who is worth finding and considering?'.

Does a score of 85 mean an 85 per cent chance of being hired?

No. A score from 0 to 100 is a fit scale used by a specific system. It is not a probability of hiring or job success. Its meaning depends on criteria, data and calculation methods.

Is a candidate with a score of 90 better than a candidate with a score of 80?

This cannot be determined solely on the basis of scoring. The difference might stem from CV completeness, how experience was described or a single heavily-weighted criterion. The score helps set analysis priority but does not replace evaluating the entire profile.

Can AI Scoring automatically reject candidates?

Technically, the score can be linked with automation, but automatically rejecting solely based on a threshold significantly increases the risk of error, discrimination and entering the territory of solely automated decisions. A safer approach is using scoring for prioritisation and ensuring a real human review.

Does the recruiter have to read the whole CV after receiving a score?

Scoring can shorten the analysis and point out key fragments, but before a significant decision, the recruiter should check the source material. The review scope can depend on the stage and risk. The score and generated summary alone should not be the sole basis for rejection.

What data does AI Scoring analyse in Recruitify?

In Recruitify, scoring can factor in project requirements and information from the CV, LinkedIn profile and application form responses. The scope of information used depends on the data available in the profile and project.

Is LinkedIn more reliable than a CV?

Not automatically. A LinkedIn profile can be newer or broader, but it is also created by the candidate and can contain gaps. It is best to treat sources as complementary and flag inconsistencies for verification.

What happens if important information is missing from the CV?

A good system should indicate a lack of confirmation, rather than automatically assuming a lack of competence. The recruiter can check other sources or ask a question during the interview. Critical data is best collected in the application form.

Does AI Scoring recognise synonyms and context?

Modern language models can connect similar meanings and recognise experience described in different words. However, they are not infallible. Semantic connections should be explained and checkable in the document.

Does AI Scoring detect if a CV was written by AI?

This is not the primary task of scoring, and there is no foolproof method to detect every AI-generated text. It is more important to verify whether the declared competencies are real. Suggested questions can help verify them during the interview.

Can AI Scoring detect lies in a CV?

No. It can spot inconsistencies between sources or vague fragments, but it does not confirm the truth of experience. Verification requires a conversation, test, references or documents.

Does AI Scoring evaluate soft skills?

It can analyse information suggesting certain experiences, but a CV is not a sufficient source for a reliable evaluation of empathy, communication, resilience or leadership style. The system should suggest questions rather than issue categorical evaluations.

Can the score change?

Yes. The score can change after updating the CV, supplementing the form, changing project requirements or updating features. Therefore, it should be treated as the result of a specific analysis at a given time.

Can you compare scores across different projects?

Usually, this makes no sense. Every project has different criteria, weights and score distributions. A score of 70 in one recruitment does not necessarily mean the same as 70 in another.

Can AI Scoring reduce recruiter bias?

It can increase consistency and limit some random evaluations, but it can also introduce or reinforce bias stemming from criteria, data and the model. Testing, monitoring and the ability to correct are needed.

Does removing names and photos solve the discrimination problem?

Not fully. Anonymisation limits some signals, but other information can indirectly indicate gender, age, origin or social status. Bias must be analysed across the entire process.

Do you need to inform candidates about using AI?

In many situations, information obligations arise under GDPR, and for high-risk systems, the AI Act outlines additional rules for informing individuals subject to the system. The scope of information depends on the specific implementation, but transparency should be a standard even when the law only requires a minimum.

Is AI Scoring subject to the AI Act?

Systems intended to be used for screening and filtering applications and evaluating candidates are listed in Annex III of the AI Act as employment-related uses. The final classification of a specific feature depends on its purpose, impact, usage and exceptions in Article 6(3). This requires an individual analysis.

When do AI Act rules for high-risk systems in recruitment start to apply?

According to the current timeline, following the entry into force of the AI Omnibus, rules for high-risk systems in Annex III are to apply from 2 December 2027. Some other provisions, including those on AI literacy, apply earlier.

Is AI Scoring compliant with GDPR?

The feature name alone does not determine compliance. Compliance depends on the legal basis, data scope, transparency, retention, security, division of roles, the ability to realise candidate rights, and how decisions are made. A tool can support a compliant process, but it does not 'solve GDPR' for the organisation.

Is a DPIA required?

In many implementations of candidate scoring, a Data Protection Impact Assessment may be required or at least highly justified due to systematic evaluation of individuals and potentially significant impact. The decision should be made based on the specific process and consultation with a DPO or lawyer.

What does human-in-the-loop mean?

It means real, not formal human involvement. The recruiter should understand limitations, see the score basis, be able to challenge it, and have actual authority to make a different decision. Merely clicking confirmation is not enough.

Will AI Scoring replace recruiters?

It should not. It automates part of the analysis and structuring of information, but it does not replace conversations, understanding context, evaluating motivation, negotiating or responsibility for the decision. It shifts the weight of a recruiter's work from manual screening to interpretation, verification and human connection.

For what type of recruitment is scoring not a good solution?

It can deliver less value in processes with very few candidates, unclear criteria, a key role for confidential market knowledge or a strong emphasis on competencies evaluated only in interaction. It should not be forced just because it is available.

How to start testing AI Scoring?

It is best to select a few roles with clear requirements, prepare criteria, test the feature on representative profiles and compare results with recruiters' independent evaluations. Then, establish human review rules, train users, and measure both time, quality and risk.

33. Summary - AI Scoring does not choose the person, it supports the evaluation

AI Candidate Scoring is one of the most practical applications of artificial intelligence in modern recruitment systems. It can help teams structure incoming applications faster, match profiles to requirements more consistently, prepare interviews better and collaborate more smoothly with Hiring Managers or clients. Its value, however, does not come from the score of 0 to 100 itself. What matters most is the context, the explanation, the quality of criteria, and the ability to verify and challenge the evaluation.

Scoring is not AI Matching. It is not primarily used to find candidates in the database, but to evaluate an application or candidate in a specific project. It is not a CV parser, though it uses extracted data. Nor is it a simple killer question or an automated rejection decision. It can link with these features, but each has a different task and a different risk profile.

The biggest mistake would be to assume the number closes the discussion. The score can help set the order of work, but it should not define a candidate's value or replace a conversation. Documents are incomplete, criteria can be flawed, and models can generate inaccurate or biased conclusions. That is why a good system highlights not only fit, but also information gaps, limitations and questions for verification.

In 2026, a responsible implementation of AI Scoring requires combining three perspectives. The first is productivity: whether the feature genuinely saves time and improves process quality. The second is recruitment practice: whether criteria are job-related and the recruiter retains their own judgment. The third is governance: whether the organisation controls data, bias, security, transparency and compliance with GDPR and the AI Act.


The best AI Scoring does not tell a recruiter whom to hire. It helps them see faster what is already known, what is still unknown and what is worth asking before they make a decision.


Sources and Further Reading

[1] European Parliament and Council of the EU, Regulation (EU) 2024/1689 - Artificial Intelligence Act, notably Annex III point 4 on employment, recruiting and candidate evaluation: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

[2] European Parliament and Council of the EU, Regulation (EU) 2016/679 - GDPR, notably Article 22 on automated individual decision-making: https://eur-lex.europa.eu/eli/reg/2016/679/oj

[3] European Commission, Draft Commission guidelines on the classification of high-risk AI systems, including materials on Annex III AI Act, 2026: https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems

[4] European Commission, Guidelines for providers and deployers of AI high-risk systems - current timeline of high-risk rules application, 2026: https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems

[5] European Commission, AI Literacy - Questions & Answers, current information on Article 4 AI Act: https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers

[6] Information Commissioner’s Office, Thinking of using AI to assist recruitment? Our key data protection considerations, 2024: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/11/thinking-of-using-ai-to-assist-recruitment-our-key-data-protection-considerations/

[7] Wen A. et al., FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations, 2025: https://arxiv.org/abs/2504.01420

[8] Hunkenschroer A. L., Luetge C., Ethics of AI-Enabled Recruiting and Selection: A Review and Research Agenda, Journal of Business Ethics, 2022: https://link.springer.com/article/10.1007/s10551-022-05049-6

[9] Information Commissioner’s Office, AI tools used in recruitment - audit outcomes and recommendations, 2024: https://ico.org.uk/action-weve-taken/audits-and-overview-reports/2024/11/ai-tools-used-in-recruitment/

[10] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0): https://www.nist.gov/itl/ai-risk-management-framework

[11] Information Commissioner’s Office, Automated decisions can streamline the hiring process with the right safeguards in place, 2026: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2026/03/automated-decisions-can-streamline-the-hiring-process-with-the-right-safeguards-in-place/

[12] Recruitify, What is an ATS? A complete guide to applicant tracking systems (2026): https://www.recruitify.ai/blog/what-is-an-ats-complete-guide-to-applicant-tracking-systems-(2026)/


Editorial Note: this article describes principles and good practices at a general level. It does not constitute legal advice. Classification and obligations concerning a specific AI system depend on its purpose, design, data, implementation method and actual impact on decisions.

News & Updates

Stay up-to-date with the latest innovations, features, and tips about Recruitify!

First Name
Email

By providing your email address within the newsletter sign-up form, you confirm its processing to send marketing information regarding the Administrator’s products and services. The Administrator of your personal data processed for the abovementioned purposes is Recruitify Spółka z o.o., based in Warsaw, Poland (KRS 0000709889). For more information on the principles of personal data processing and the rights of data subjects, please check the Privacy Policy.

Share

Published

Category

Applicant Tracking System

Author

Iwo Paliszewski

AI candidate scoring

Last updated:

AI Candidate Scoring: What is it, how does it work, and how to use it responsibly? The ultimate guide (2026)

Innovations

Iwo Paliszewski

Iwo Paliszewski

You search Google for ‘AI Candidate Scoring’ and, more often than not, you are presented with one of two answers. The first goes: artificial intelligence reads CVs, assigns a matching percentage to the candidate and highlights the top talent. The second warns that the algorithm takes over the recruiter’s decisions, creates soulless rankings and can automatically shut someone out of a job opportunity. Both answers simplify the topic so much that instead of clarifying it, they merely build further misunderstandings.

AI Scoring does not have to be a magical autopilot selecting the ‘best people’, nor a black box passing judgment. In a well-designed process, it is a tool supporting the assessment of a specific application or candidate against the requirements of a specific project. It can present a numerical score, for example from 0 to 100, but it can also provide a contextual assessment: a profile summary, strengths, weaker areas, information gaps and questions worth asking during an interview. It delivers the greatest value when it does not replace human judgment, but helps structure, speed up and better justify it.

This guide answers the question ‘what is AI Candidate Scoring’ as it should be answered in 2026: comprehensively, practically and without marketing shortcuts. We explain exactly what is assessed, where the score comes from, how scoring differs from AI Matching, CV parsers and disqualifying questions, how to set up solid criteria, what data can be analysed, where the risks lie, and what GDPR and the EU AI Act say about AI used in recruitment.


Quick Definition
AI Candidate Scoring is an AI-supported assessment of a specific application or candidate against the requirements of a given recruitment project. The result can be numerical, contextual, or a combination of both. A good scoring system does not just show a number; it also explains what influenced the rating, highlights the strengths and weaker areas of the profile, and suggests questions for further verification. It is not the same as AI Matching, which is designed to search the database and suggest candidates who fit a project.


Table of Contents

1. What is AI Candidate Scoring?

2. What exactly does AI Scoring assess: the human, the candidate or the application?

3. Why has AI Scoring become important right now?

4. Numerical and contextual assessment - two dimensions of good scoring

5. How does AI Scoring work step-by-step?

6. What data can AI Scoring use?

7. Why are project requirements more important than the AI model itself?

8. Must-haves, nice-to-haves and disqualifying criteria

9. AI Scoring vs. keyword search

10. What does a score from 0 to 100 mean?

11. Strengths, weaknesses and interview questions

12. Missing information is not the same as a lack of competence

13. AI Scoring vs. the traditional candidate scorecard

14. AI Scoring vs. AI Matching - a key distinction

15. AI Scoring vs. CV parser, AI summary and killer questions

16. AI Scoring in practice - three assessment examples

17. Benefits for the recruiter and the recruitment team

18. Benefits for agencies, in-house HR and Hiring Managers

19. Benefits and risks from the candidate’s perspective

20. Where does AI Scoring deliver the most value?

21. How to implement AI Scoring in your recruitment process?

22. Most common mistakes when using AI Scoring

23. Limitations of AI Scoring - what does the system not know?

24. Bias, discrimination and the illusion of objectivity

25. Human-in-the-loop - what does real human oversight mean?

26. AI Scoring and GDPR

27. AI Scoring and the AI Act

28. How to choose an AI Scoring tool?

29. How to measure the effectiveness of AI Scoring?

30. How does AI Scoring work in Recruitify?

31. The future of AI Scoring

32. FAQ - Frequently Asked Questions

33. Summary - AI Scoring does not choose the person, it supports the evaluation


1. What is AI Candidate Scoring?

AI Candidate Scoring is a method of supporting the assessment of an application or a candidate's profile using artificial intelligence. The system analyses the available information about the candidate, compares it to the requirements of a specific project and prepares a result designed to help the recruiter in their next steps. Depending on the solution, the result can take the form of a number, percentage, category, description or a set of several elements. In Recruitify, this includes a score from 0 to 100, a contextual description of the candidate, a comparison of strengths and weaker areas, and suggestions for questions to ask during the interview.

The key phrase here is 'against the requirements of a specific project'. AI Scoring should not evaluate whether someone is a good or bad candidate in general. There is no universal professional value of a human being that can be honestly wrapped up in a single number. The exact same candidate might get a high score in a project looking for someone to scale B2B sales in the Polish market, and a significantly lower score in a project requiring experience in international enterprise sales. This does not mean their skills changed between the two measurements. The context and objective of the evaluation changed.

That is why it is safest to think of AI Scoring not as an evaluation of a person, but as a structured analysis of how well the currently available information fits. This phrasing is longer, but much more precise. The system does not see the entire professional history, potential, character, motivation or behaviour of the candidate in their future job. It sees a specific set of data and compares it against a specific set of criteria. Its utility therefore depends entirely on the quality of both sides of this comparison.

It is also worth distinguishing scoring from ranking itself. Scoring is the process of evaluating and explaining a result. Ranking is one of the possible ways to use this result, for example, by sorting applications from the highest to the lowest score. You can have scoring without an automatic ranking, and you can also create a ranking based on simple rules that have nothing to do with AI. In practice, ATS systems often combine both elements, but from the perspective of a responsible process, they should not be treated as synonyms.


The Golden Rule
AI Scoring should answer the question: ‘What in the available data suggests this candidate fits this role, what raises doubts, and what still needs to be verified?’, rather than: ‘Does this person deserve the job?’.


2. What exactly does AI Scoring assess: the human, the candidate or the application?

In everyday language, people usually talk about 'candidate scoring', but within a recruitment system, the object of evaluation can be both an application and a candidate assigned to a project. This distinction has practical significance. An application is a specific submission to a specific job opening. It can include the CV sent in response to the advert, answers from the application form, the application source, and additional documents or information provided during submission. A candidate, on the other hand, is a person who already exists in the database, who can be added to a project by a recruiter, sourced, or reconsidered in a new recruitment round without submitting a new application.

AI Scoring can therefore assess an incoming application, but it can also be run for a person found previously in the database or added to a project by a recruiter. In both cases, the evaluation should be anchored in the requirements of that specific recruitment. However, the scope of available data may differ. For an application, the system often has a fresh CV and answers to application questions at its disposal. For a database candidate, it can use the profile, previous documents, LinkedIn data, and information gathered in the system. If some of the data is outdated, the score may also be less up-to-date.

This leads to an important conclusion: a score is not a permanent label attached to a person. It should not be logged in anyone's mind or in the organisation's culture that 'Anna is a 64-point candidate' or 'Piotr is a 91 per cent candidate'. The correct phrasing is: 'Based on currently available data, in relation to the requirements of Project X, the system assigned Anna's application a score of 64/100'. The exact same profile in another project might receive a completely different score. Even within the same recruitment, the score can change if the recruiter refines the requirements, the candidate updates their details, or a new CV is uploaded.

This approach also guards against one of the biggest dangers of automation: turning a supporting score into a label that defines a person. In recruitment, the anchoring effect is very powerful. When a recruiter sees 92 points, they may start looking for confirmation of that high score. When they see 48, they might read the profile more critically. That is why explanations are needed alongside the numerical score, and users should be trained to treat the score as the starting point of their analysis, not the conclusion.


3. Why has AI Scoring become important right now?

Candidate scoring is not an invention of the generative AI era. Recruiters have been using scorecards, Excel spreadsheets, killer questions, weighted criteria and simple point systems for years. In many companies, the person managing the recruitment was already assigning points to candidates for experience, language skills, availability or knowledge of a specific technology. The problem was that such a process was time-consuming, highly dependent on user discipline, and often ended up as a paper-based methodology that nobody consistently used after receiving the hundredth application.

The shift today is not so much the existence of a point system itself, but rather the ability to automatically analyse massive amounts of unstructured text. Classic rules are good at dealing with 'yes' or 'no' answers, a specific number of years, or an exact certificate name. Modern language models can additionally compare the meaning of requirements with project descriptions, responsibilities and achievements written in different words. Thanks to this, scoring does not have to be limited to matching exact phrases. It can prepare a draft summary, highlight potential strengths and weaknesses, and suggest questions to help clarify ambiguities.

At the same time, the recruitment environment has changed dramatically. Easy applying, one-click apply forms, automatic alerts and generative tools that help build CVs mean that submitting an application takes much less effort than it did a few years ago. In many processes, the number of applications grows faster than a team's capacity to read them thoroughly. However, a larger volume does not mean more people who actually fit the role. Recruiters must find valuable profiles among documents that are random, mass-sent, linguistically very similar or tailored to match the job advert's keywords.

In such an environment, time is not the only issue. A human does not read the two-hundredth CV with the same freshness as the first. Concentration levels drop, pressure rises, and initial evaluation is increasingly based on a few of the most visible elements. A candidate who applied in the morning might get more attention than an equally good person whose document landed in the system at the end of the day. AI Scoring can help apply the same analytical framework to all submissions, but this does not guarantee automatic objectivity. If the criteria are flawed, the system will replicate the error more consistently than a human.

Scoring has therefore become important right now because three trends have converged: a growing volume of applications, increasing language analysis capabilities, and pressure for a faster, more measurable process. This does not mean every recruitment needs AI. It does mean, however, that teams must find a way to maintain evaluation quality in a scenario where manual document screening increasingly becomes the bottleneck.


AI Scoring is not the answer to a lack of recruitment strategy. It is an attempt to bring structure to evaluation where scale, pace and the volume of data make consistent human work difficult.


4. Numerical and contextual assessment - two dimensions of good scoring

The most visible element of scoring is usually the number. A score of 84/100 allows you to spot an application quickly, compare it with others and set the review order. It is particularly convenient when there are dozens or hundreds of profiles in a project. However, the number alone says surprisingly little. It does not explain which criteria were met, which ones carried the most weight, what was missing from the documents, or whether a lower score is due to a genuine mismatch or simply a lack of information.

That is why a good AI Scoring system should have two complementary dimensions. The first is numerical and structuring. It helps set the order of work and quickly flag profiles that require attention based on the adopted criteria. The second is contextual and explanatory. It shows what the system found in the data, how it linked that information to the role, which elements it considered strengths, which ones were weaknesses or unconfirmed, and what still needs checking.

In Recruitify, users get a score from 0 to 100, a candidate description in the context of the project, strengths, weaknesses and suggested interview questions. These elements are not just decoration for the percentage. They are precisely what allows the recruiter to judge whether the score makes sense. If someone gets 88 points but the justification does not confirm a critical must-have, the recruiter should spot the issue. If a profile has 58 points mainly because it lacks information on team scale or salary expectations, the correct response might be a quick screening question, not a rejection.

Contextual evaluation also helps reduce the anchoring effect. A large number displayed on the screen easily becomes the first and most dominant piece of information about a candidate. A human starts looking for arguments to back up the score instead of analysing the profile independently. When they see the sources of the score, unconfirmed areas and questions next to the number, it is much easier to treat the scoring as a working hypothesis rather than a final verdict.

However, contextual results must use cautious language. ‘The candidate managed a team of ten’ can be a fact derived from the CV. ‘The candidate is an excellent leader’ is a conclusion that the document does not prove. ‘No information on budget management was found in the available materials’ means a lack of evidence, not a proven lack of competence. A good tool and a well-trained user distinguish between facts, interpretations and information gaps.

It is possible to have purely numerical scoring, but it will lack transparency. It is also possible to have a purely descriptive evaluation, but with a high volume of applications, it is harder to use for prioritisation. The greatest value comes from combining both layers.


The number organises. The explanation allows you to evaluate. The question allows you to verify.


5. How does AI Scoring work step-by-step?

While different systems may use different models, algorithms and interfaces, the practical AI Scoring process can be presented as a sequence of several linked stages. Crucially, the evaluation does not start with the CV. It starts with defining the role and deciding what information matters for this specific recruitment.

1. The recruiter creates a project and describes the requirements. The system needs a point of reference. A job title alone or the full text of a job ad is usually not enough. Requirements should specify genuine must-haves, nice-to-haves, acceptable alternatives and questions that cannot be resolved solely on the basis of documents.

2. Criteria are checked and approved by a human. AI can help structure the description, but it should not decide on its own that exactly five years of experience, working at a specific company or graduating from a particular university is essential. The recruiter and the Hiring Manager must confirm that the criteria are job-related, proportionate and verifiable.

3. The system gathers available data about the application or candidate. In Recruitify, this can include information from the CV, LinkedIn profile and answers provided in the application form. The scope of data may vary between individuals, so the score must always be read in the context of the material's completeness.

4. Data is structured and mapped to the criteria. The system can recognise not only identical words but also semantic connections. A requirement for managing a sales team can be linked to a description of running a business development department. However, such a conclusion should remain auditable by the recruiter.

5. The AI distinguishes between what is confirmed and what is unknown. A good solution should not turn silence into a negative answer. A lack of information about Kubernetes is not the same as a declared lack of knowledge of Kubernetes. On the other hand, information from a form stating that a candidate can only start work in six months is a concrete signal that contradicts a requirement to start within a month.

6. A numerical score and contextual evaluation are generated. The system can award a score from 0 to 100, prepare a profile description, point out strengths and weaknesses, and suggest questions for further discussion. The result should be a trail of analysis, not an automated decision to advance or reject.

7. The recruiter checks the justification and source material. If a score seems too high or too low, the cause must be determined. The issue could lie in the requirements, data extraction from the CV, model interpretation or simply the incompleteness of the profile. A human must have the ability to make a different decision than suggested by the scoring order.

8. The result leads to action. The recruiter can start with high-match profiles, plan a call, ask a follow-up question or go back to the Hiring Manager to refine the brief. Scoring creates no value if it ends with a colourful number that nobody acts on.

9. The organisation monitors quality and impact. The team should track where scoring works well, when it generates inaccurate conclusions, how often recruiters disagree with the assessment, and whether lower scores are leading to valuable people being automatically overlooked. Implementation does not end when you toggle the feature on. That is when the responsible management of its impact begins.

In practice, this process may take a few seconds on the system's side, but its quality is the result of prior human work. The better the brief, the more adequate the data and the more conscious the user, the greater the utility of the result.


[SCREENSHOT: example of AI Scoring flow in Recruitify - project requirements, score 0-100, summary and questions]


6. What data can AI Scoring use?

The quality of the scoring depends not only on the AI model, but primarily on the quality, relevance and scope of the input data. In Recruitify, the evaluation can take into account information from the CV, LinkedIn profile and answers provided in the application form. These sources are not equivalent and should not be mechanically lumped together. Each shows a different piece of the profile and has its own limitations.

The CV usually provides the most structured professional history: job titles, companies, dates, responsibilities, education, skills and achievements. However, it is a selective document. The candidate decides what to fit on one or two pages, may shorten older experience, omit a project deemed less relevant, or use role titles specific to their previous organisation. A CV can also be outdated or generic, prepared without knowledge of a specific project's criteria.

A LinkedIn profile can contain a newer employment history, broader project descriptions, recommendations or skills omitted from the CV. However, it should not automatically be treated as more reliable. It is also created by the user, can be incomplete, outdated or written in a marketing style. Its value lies primarily in adding context and spotting discrepancies, not in serving as external proof of truth.

The application form allows you to collect information that is often missing from both the CV and LinkedIn. It can cover availability, salary expectations, preferred working model, right to work in a given country, willingness to relocate, required licences or experience in a very specific area. A candidate usually does not write in their CV whether they can start work in August, whether they accept two days a week in the office, or whether they have worked with a specific version of a system before. However, they can answer such questions in the form.

It is precisely the combination of sources that allows you to build a fuller picture. Imagine a candidate for a Finance Manager role. Her CV describes team leadership and month-end closing, LinkedIn shows a more recent promotion, and in the form, she confirms knowledge of a specific ERP system and availability in the required timeframe. Scoring based solely on the CV would miss the last two elements. At the same time, the system should not extrapolate anything that no source confirms.

Data should be interpreted according to three basic states: criterion confirmed, criterion probably unmet, and lack of sufficient information. This distinction is more important than the technical number of sources. Two outdated profiles do not create better proof than one fresh and specific answer.

It is also worth establishing a hierarchy of relevance and rules for resolving contradictions. What should the system do if a CV indicates employment ended in May, but LinkedIn shows ongoing employment? Should a form response submitted yesterday take precedence over a CV prepared six months ago? A good process does not hide such discrepancies under a uniform score. It signals them to the recruiter and turns them into questions for verification.

Combining multiple sources does not mean you should analyse everything that is technically available. The data minimisation principle requires using only the information necessary for a specific purpose. Age, photos, family situation, health status, origin or other sensitive areas should not affect scoring just because they can be found in a document or profile. More data does not always mean a better evaluation. Sometimes it just means more noise and a higher risk of biased conclusions.


The best scoring does not analyse the largest amount of data. It analyses data that is adequate, up-to-date and needed to evaluate the requirements of a specific project.


7. Why are project requirements more important than the AI model itself?

The most advanced model will not fix a poorly defined recruitment. If project requirements are vague, contradictory, unrealistic or contain biases, AI will only apply them faster and more consistently. It is precisely this consistency, often presented as an advantage of automation, that can become a problem: a human reading a few CVs might eventually notice that a criterion makes no sense. The system will keep repeating it for every person until someone changes it.

A good requirement should be specific, related to actual work and verifiable in the available data or during a later stage of the process. ‘A minimum of three years of experience in managing ERP implementation projects’ is more useful than ‘extensive project experience’. ‘Independently acquiring B2B clients with a contract value over €100k’ gives more context than ‘strong sales skills’. ‘Knowledge of English at a level allowing for negotiation’ describes a business need better than just ‘C1’, if the organisation does not intend to check the level with a formal test.

Requirements should not simply copy the entire job advert. An advert also serves a communication and employer branding function. It contains a description of the company, benefits, scope of responsibility and engaging language. Scoring, on the other hand, needs clear evaluation logic: what is a mandatory condition, what increases fit, what experience is equivalent, what can be learned after hiring, and what is absolutely non-negotiable.

The weight of criteria is also critical. Five years of experience is not always twice as good as two and a half years. A lack of one technology might be easy to catch up on, while a lack of a legally required licence actually prevents starting work. Scoring should reflect the business significance of criteria, not the number of words dedicated to them in the ad. If the ‘nice-to-have’ list is longer than the mandatory requirements, the system should not automatically allow the sum of minor additions to overshadow the lack of a core competence.

Before launching scoring, it is worth conducting a quick calibration with the Hiring Manager. Instead of only asking ‘who are we looking for?’, it is better to establish: what is this person supposed to deliver in the first six months, what experience genuinely increases the chance of success, which gaps are acceptable, and which ones make hiring impossible. Only on this basis can you create criteria that hold value for both AI and people.

8. Must-have, nice-to-have and disqualifying criteria

One of the most important stages of preparing scoring is separating three categories of requirements. Must-haves are conditions absolutely necessary to perform the job or start it within the set timeframe. Nice-to-haves increase fit, but their absence should not close the door to the process. Disqualifying criteria, however, are binary conditions where failure to meet them can justify automatically or semi-automatically directing an application to a separate pipeline - provided they are legal, proportionate and unambiguous.

In practice, organisations overuse the must-have category. A Hiring Manager might deem experience at a specific company, exactly five years of work, knowledge of an internal tool or graduation from a preferred university as essential, even though none of these are conditions for successfully performing the role. If AI Scoring receives such a list uncritically, it will reinforce a narrow 'ideal candidate' profile and lower the scores of people with diverse but highly valuable career paths.

Disqualifying criteria should be applied with extreme caution. They make sense where the answer is binary and directly related to the ability to do the work: a legally required licence, the right to work in a given location, availability during required hours, or willingness to work in a model the organisation cannot change. Do not turn elements requiring interpretation, such as 'good cultural fit', 'sufficient experience' or 'fitting personality', into killer questions.

AI Scoring should also not hide disqualifying rules within an opaque number. If a candidate received a low score due to the lack of a legally required certificate, the recruiter should see this clearly. If a form response indicates that a candidate cannot work the required hours, the system should distinguish this information from a lack of data. Transparency of the evaluation logic matters both for the quality of decisions and for the ability to explain the process to the candidate.

The best practice is to assign not only a category and weight to each criterion, but also a method of verification. Can the information come from the CV? Is a response in the form needed? Does it need to be confirmed during an interview? Does it require a document? This way, scoring does not pretend it can resolve everything based on a single source.

9. AI Scoring vs. keyword search

One of the most common simplifications in conversations about ATS is the belief that the system 'rejects CVs if they don't contain the right words'. In some older solutions, exact phrases, simple filters and manually built queries did play a large role. Keyword searching is still useful, but it is not the same as AI Scoring and should not be presented as its full mechanism of action.

Classic search checks for the presence or absence of specific expressions. If a recruiter types 'Salesforce', the system can find people who placed that name in their profile. The problem begins when the experience is described in a different language, the candidate uses a broader category name, or a specific competence is implied by context. A person who 'implemented and administered CRM solutions in a Salesforce environment' should be found easily. A harder case is a candidate describing 'managing a cloud-based sales automation platform' without mentioning the brand. A semantic system can recognise a potential connection, but it should treat it as a conclusion requiring confirmation, not as a certain fact.

AI Scoring based on a language model can match the meaning of requirements with the content of documents. This allows it to spot equivalent experience, role titles specific to an organisation, and skills described in a different order than in the job ad. For example, a requirement for 'managing a customer success team in a SaaS model' can be linked to a description of 'responsibility for an eight-person subscription customer care team'. Such reasoning is useful but not infallible. The more semantically distant the connection, the greater the need to show the recruiter exactly which part of the document the conclusion was based on.

We should not, therefore, pit keywords and AI against each other in an absolute way. Good tools can combine precise rules with contextual analysis. For a 'CISA' certification, the exact phrase matters. For experience in leading an organisational transformation, analysing the meaning is more important than identical words. The choice of mechanism should depend on the nature of the criterion.

It is also worth being cautious about promises that AI 'understands CVs like a human'. This is attractive from a marketing standpoint but is overreaching. A model can analyse text in a way that resembles human language processing, but it does not possess professional experience, intentions or full context. It can correctly link synonyms while completely misinterpreting the scope of responsibility. That is why semantisation increases capabilities but does not remove the need for human verification.


10. What does a score from 0 to 100 mean?

A score from 0 to 100 is primarily an auxiliary scale. It allows you to quickly see how the system evaluated the fit of available information against the project criteria. It is not a probability of hire, a prediction of job success, or a scientifically calculated 'candidate quality'. A score of 86 does not mean the candidate has an 86 per cent chance of succeeding in the role. Nor does it mean they are exactly 14 points better than someone with a score of 72.

The meaning of the number depends on the design of the specific system, how criteria are weighted, and the completeness of the data. In one project, scores might naturally cluster between 60 and 85, in another between 20 and 55 if the requirements are highly specialised. That is why thresholds like 'we reject everyone below 70' are dangerous if they have not been tested for a specific process. Even then, a threshold should not automatically replace human review, especially when the score affects access to employment.

The number can be most useful as a prioritisation tool. A recruiter can start their analysis with highly-rated individuals, but they should also review a sample of profiles from lower brackets. This kind of audit allows you to check whether the system is overlooking non-standard candidates, lowering scores due to a lack of information, or rewarding people who simply wrote a better CV. In practice, it is worth establishing a rule to regularly review the 'tail of the ranking', especially for new types of roles.

A good interface should not display the number in isolation from the justification. If the score is shown in a large font and the explanation is hidden under another click, users will make decisions based on the most visible element. Screen design influences behaviour just as much as the model itself. In a system supporting responsible decisions, the score, strengths, weaker areas and missing data should all be available together.

It is also worth remembering score stability. Generative models can, in some configurations, produce slightly different answers for the same input. The provider should limit this variance, test repeatability and clearly describe what might trigger a recalculation. The user, in turn, should know that changing project requirements or candidate data can legitimately change the score. The result is not a permanently written fact, but the outcome of a specific comparison at a specific time.


[SCREENSHOT: list of candidates or applications with AI Scoring from 0 to 100]


11. Strengths, weaknesses and interview questions

The greatest value of AI Scoring often lies not in the number, but in the contextual layer. A recruiter needs to know more than just that a profile was rated high or low. They need to know why, which elements matter for the project, and what they should do next. That is why a useful scoring tool can prepare a candidate description, highlight strengths and weaknesses, and suggest questions for the interview.

A candidate description should be a brief synthesis, not a summary of the entire CV. Its role is to connect the experience with the specific project. Instead of a generic 'experienced financial manager', it is better to point out that the person managed a team of similar size, was responsible for reporting to an international group and participated in an ERP system implementation, which is one of the challenges of the new role. This information helps the recruiter grasp the profile faster, but remains verifiable in the sources.

Strengths should relate to the criteria and stem from concrete data. It is not about generating compliments. 'Worked as a Sales Manager' is not a strength yet. 'Was responsible for independently acquiring enterprise clients in the DACH market for three years and achieved an annual target over €2m' is information that can be mapped to the project requirements.

Weaknesses require even more caution. In product language, we might call them weaknesses, but in practice, the system should separate at least three types of information. The first is a genuine contradiction with a requirement, such as lacking a legally required licence or availability being later than the project allows. The second is a probable mismatch, when the materials indicate experience significantly different from what is sought. The third is a lack of data, i.e., a situation where the documents simply do not allow a criterion to be resolved. These categories must not be treated the same way.

Suggested questions turn scoring from an organising mechanism into an interview preparation tool. If the CV lacks information on team scale, the system might suggest: 'How many people were in your team and what personnel decisions were you responsible for?'. If a candidate lists a technology without context: 'In which project did you use it and what was your level of independence?'. If LinkedIn and the CV differ on dates: 'Which information is current and what is the reason for the discrepancy?'. A good question is not meant to confirm a preconceived score. It is meant to supply the information that is still missing.

It is also worth avoiding questions that are too general, leading or about areas unrelated to the job. AI can generate a proposal, but the recruiter must check it and tailor it to the stage of the process. Asking about motivation during the first short screening might make sense. Assessing personality based on a stereotypical assumption does not.

The contextual layer has another benefit: it facilitates communication with the Hiring Manager or client. Instead of passing on just a number or a brief 'the profile fits', the recruiter can show arguments in favour, risk areas and a verification plan. This improves shortlists and shifts the discussion from gut feeling to concrete evidence.


[SCREENSHOT: candidate description, strengths, weaknesses and suggested questions in Recruitify]


12. Missing information is not the same as a lack of competence

This is one of the most critical principles of interpreting AI Scoring. Application documents are an incomplete description of a human being. A candidate might not list a skill because they took it for granted, wanted to limit CV length, tailored the document for a different role, or did not know the exact project requirements. A LinkedIn profile might be updated once every few years. An application form might contain a question that is too broad. Therefore, a lack of information is primarily a lack of evidence in the analysed material, not proof of a lack of competence.

The system should distinguish between at least three states:

Criterion confirmed - available data clearly indicates it is met.

Lack of sufficient information - the material does not allow confirming or ruling out that the criterion is met.

Criterion probably unmet or unmet - the data contains information that contradicts the requirement, e.g., the candidate declares availability in six months, and the project requires a start within four weeks.

Lumping the second and third categories together is a frequent source of unfair decisions. If a CV makes no mention of a specific technology, the system can write 'no confirmation in available data'. It should not, without additional basis, state 'the candidate does not know the technology'. If a candidate responds in a form that they have never used it, the situation is different, but even then, the weight of this information depends on whether the technology is a genuine must-have or a skill that can be quickly learned.

The style of the document also matters. A detailed CV provides more text and more potential evidence than the sparse profile of an expert who assumes company names and job titles are enough. As a result, the system may partially reward self-presentation quality. This is not a problem unique to AI, as recruiters also fall for well-written documents, but automation can scale this effect and give it a veneer of mathematical precision.

One way to mitigate this issue is to use the application form to gather critical information in a standardised way. If experience in a given area really matters, it is better to ask a specific question than to hope the candidate happened to describe it in their CV. The second way is generating interview questions. The third is regularly reviewing a portion of low-scoring profiles to check if the system mistook sparse documentation for a lack of qualifications.

This principle should also be part of user training. A recruiter must know that 'not found' does not mean 'does not possess', and 'the system suggests' does not mean 'the fact is confirmed'. Without this awareness, even well-designed explanations can be applied too categorically.


Missing information should lead primarily to a question. Not to an automated rejection.


13. AI Scoring vs. the traditional candidate scorecard

A traditional scorecard relies on pre-established criteria and a scale. A recruiter or Hiring Manager assigns points for experience, competencies, motivation or task results after an interview. A well-used scorecard increases process consistency because it forces evaluators to refer to shared criteria rather than a general impression.

AI Scoring should not replace this scorecard, but rather complement it at an earlier stage. It analyses the data available before the interview and helps prepare an initial assessment. The traditional scorecard can then cover information obtained during the interview, task, test or reference checks. Combining both tools makes sense when the criteria are consistent and it is clear which source can confirm which element.

The difference also relates to accountability. In a scorecard, a human directly assigns the rating. In AI Scoring, part of the interpretation is performed by the system, so the user must be able to trace the logic, check the basis and challenge the score. It is not enough to say that 'the algorithm calculated it'. The organisation remains responsible for how the evaluation is used in the process.

AI can improve the consistency of the initial analysis, but it does not remove human bias. Criteria are created by humans, data is created by humans, and results are interpreted by humans. What is more, the AI score can become a new source of bias - an automatic anchor that evaluators will adjust their own impressions to fit. Therefore, it is sometimes worth hiding the score from some individuals until they make their independent assessment, or at least requiring a brief justification for decisions independent of the score.

In practice, the best model might look like this: AI Scoring structures incoming applications and flags areas to check; the recruiter reviews the profile; the interview is conducted using standardised questions; the scorecard gathers evidence from the interview; the Hiring Manager makes a decision based on the entire set of materials. AI is then one element of the process, not an independent judge.

14. AI Scoring vs. AI Matching - a key distinction

AI Scoring and AI Matching are often lumped into the same category because both features match candidate profiles to project requirements. However, their goals are different, and in articles, product documentation and client communication, they should be consistently separated.

AI Scoring evaluates an application or candidate we are already reviewing within the context of a project. The starting point is a specific profile. The system prepares a score and explanation: to what extent the available information matches the criteria, what the strong and weaker areas are, and what needs verifying. It can apply to a person who applied themselves, or a candidate added to the project by a recruiter.

AI Matching works from the other side. The starting point is the project and the need to find potentially fitting individuals. The feature searches the candidate database, compares profiles with requirements, and suggests to the recruiter who they might want to consider. Its main task is discovery and internal sourcing: finding people the recruiter has not yet added to the project or might have forgotten about.

The simplest distinction is:

AI Scoring: ‘How does this candidate or application stack up against the project requirements?’

AI Matching: ‘Which candidates in the database might fit this project and should be suggested to the recruiter?’

In Recruitify, AI Scoring is an evaluative feature, whereas AI Matching is being developed as a separate mechanism for searching and suggesting candidates from the database. Scoring should not be presented as 'the next step after matching', as the process can run without matching. A candidate can apply via an ad, be added manually, or come from a referral, and then be evaluated by AI Scoring.

Both features can complement each other in the future. Matching finds potential profiles, the recruiter adds selected individuals to the project, and scoring prepares a more detailed assessment against the criteria. However, separate quality standards are still required. Matching should be measured by suggestion relevance and the ability to find valuable profiles. Scoring should be evaluated based on consistency, explainability and utility in analysing a specific person.

15. AI Scoring vs. CV parser, AI summary and killer questions

In a modern ATS, several features might work on similar data but perform completely different tasks. Confusing them leads to unrealistic expectations and makes risk assessment difficult. It is worth distinguishing AI Scoring from a CV parser, candidate summary, search, killer questions, workflow automation and the simple use of a public chatbot.

A CV parser reads a document and converts it into structured data: contact details, employment history, education, skills or languages. Its core task is extraction, not evaluation. A parser can make an error, such as assigning a date to the wrong company, and this error can later affect scoring. That is why extraction quality is one of the pillars of evaluation quality.

An AI summary creates a profile abstract. It can describe experience, key skills and career path, but it does not have to match them to project criteria or award a score. A summary primarily answers 'what is known about this candidate?', whereas scoring asks 'how does this information relate to the requirements of this recruitment?'. In practice, both elements may be presented together, but they are not the same.

Killer questions or disqualifying questions are rules based on specific answers. If a job requires a valid driving licence and the candidate answers 'no', the system can apply a pre-established path. This is not necessarily AI. It is rule-based automation whose logic was defined by a human. It has advantages in unambiguous cases, but can be more ruthless than scoring if a criterion was poorly chosen.

Workflow automation performs actions when a condition is met: sends a message, adds a label, changes a stage, creates a task or notifies the project owner. It can use the scoring result as one of the conditions, but such a connection requires extreme caution. Automatically rejecting everyone below a certain score turns a supporting feature into an actual decision-making mechanism and significantly increases legal and ethical risks.

A separate case is copying a CV and job description into a public chatbot with the prompt 'evaluate this candidate'. Technically, you can get a percentage, summary and questions this way, but it is not equivalent to secure AI Scoring embedded in an ATS. The recruiter may not know where the data goes, how long it is stored, whether it is used to train the service, who has access to the chat history, and whether the organisation has an appropriate agreement with the provider. It is also harder to maintain uniform criteria, log requirement versions, control permissions and document how the result influenced the process.

A feature available directly within the recruitment system should operate within an established environment of data processing, access control, retention, logs and provider agreements. The mere fact of integration with an ATS does not guarantee compliance or security, but it gives the organisation the ability to manage the process in a way that spontaneously copying data into a random tool usually fails to provide.

The most important question is therefore not 'can AI generate an evaluation?'. It can. The question is: is the evaluation generated based on approved criteria, on adequate data, in a controlled environment, and with the ability for a human to review it?

16. AI Scoring in practice - three assessment examples

Theory becomes clearer when we see how the same mechanism can work across different recruitment projects. The following examples do not present a ready-made numerical formula. Rather, they show how requirements, data and contextual interpretation should work together.

Example 1: CFO in a high-growth company

Let’s assume a company is looking for a CFO to prepare the organisation for rapid growth, structure controlling, manage funding and work with investors. A poorly described requirement might read: ‘minimum ten years of experience, prestigious company on the CV, strong leadership and strategic thinking’. Such a brief is difficult to score fairly. ‘Prestige’ is subjective, ‘strategic thinking’ does not stem directly from the document, and the number of years might not correlate with the challenge.

Better criteria relate to results and the working environment: responsibility for full P&L, experience in building controlling structures, involvement in securing funding, managing a finance team, reporting to the board or investors, and working in an organisation of a similar scale of growth. The CV can confirm some of these elements. LinkedIn can show a more recent role. The application form can collect availability, expectations and the answer to a question about the largest funding round managed.

The candidate might get a high score because they ran finance in a company growing from 50 to 300 people, implemented controlling and participated in talks with a fund. A weakness or risk area might be the lack of confirmation of independent responsibility for securing funding. Rather than assuming this competence is missing, the system can suggest a question: ‘What was your personal scope of responsibility in the funding process and which elements did you manage independently?’.

Example 2: Java Developer in a migration project

In the second project, a company is seeking a Java Developer to migrate a legacy application to a new architecture. Must-haves might include commercial experience with Java, working with Spring Boot, relational databases and distributed systems. Nice-to-haves are Kubernetes, public cloud and experience in gradually modernising a monolith.

The candidate lists Java and Spring, but does not use the exact phrase 'microservices' in the CV. They do, however, describe breaking down a monolithic application into independently deployable modules communicating via API. An analysis based solely on keywords might miss this important signal. Semantic analysis can spot the connection, but it should point out the fragment on which the conclusion was based.

There is no mention of Kubernetes in the CV. This should not automatically mean a weakness in the sense of a lack of competence. If Kubernetes is a nice-to-have, the system can mark a lack of confirmation and suggest a question: ‘Have you deployed applications in a containerised environment and what orchestration tools did you use?’. If the candidate responds in the form that they worked with Kubernetes for two years, the information can complete the profile. If they answer 'no', scoring should factor this in as a genuine gap, but in line with the criterion's weight, not as an automatic rejection.

Example 3: B2B salesperson for a new market

The third project concerns a person to develop B2B sales in the German market. The title 'Sales Manager' alone says very little. For the role, independent acquisition of new clients, sales cycle length, contract value, working with decision-makers, knowledge of the DACH market and German language skills allowing for negotiations could all be significant.

Candidate A has impressive sales results, but most of their experience comes from managing existing clients in the Polish market. Candidate B has lower revenues but opened the German market independently, built a pipeline from scratch and led enterprise negotiations. A simple scoring system based on the sum of sales results might reward Candidate A. Scoring anchored in the genuine requirements of the project should notice that the second profile is a better fit for the specific task.

The application form can additionally collect information on willingness to travel, salary expectations and actual language level. The system can highlight market building experience as a strength, and the lack of information on average contract value as an area for verification. An interview question could be: ‘Tell us about the largest client you acquired independently in the DACH market - what did the process look like from first contact to signing the contract?’.

These three examples show why there is no single universal scoring for a 'good candidate'. A CFO, a Java Developer and a salesperson are evaluated against entirely different criteria. Even two candidates with similar job titles might score differently depending on the project's goal. AI can help navigate requirements consistently, but it does not relieve the organisation of the need to answer the hardest question: what do we actually need from a person in this role?

17. Benefits for the recruiter and the recruitment team

The most obvious benefit is time savings, but the value of AI Scoring should not be reduced to reading CVs faster. A well-used feature changes how the initial analysis is organised. It helps spot applications meeting multiple criteria faster, flag areas requiring checks, and prepare a more structured interview. The recruiter still reviews profiles, but they do not start with a blank page every time.

Savings also do not mean a human stops reading documents altogether. It is about a different distribution of attention. Instead of dedicating the same amount of time to every part of every CV, the recruiter can immediately see which requirements have been confirmed, where there is a contradiction and what is missing. Thanks to this, they spend more time on interpretation, contact and verification, and less on manually finding scattered information.

The second benefit is greater consistency. With hundreds of applications, people naturally differ in their attention levels and interpretation of criteria. The system can apply the same analytical framework to every profile. This does not automatically mean greater fairness, because flawed rules can also be applied consistently. However, if criteria are well-prepared and results are monitored, consistency reduces the randomness of the initial review and makes comparing decisions within the team easier.

The third benefit is better justification. The recruiter can quickly prepare arguments for the Hiring Manager or client, instead of sending just the CV and a generic 'worth a talk'. Strengths, weaknesses, information gaps and questions for verification create a shared language of evaluation. This does not mean the generated description should be copied without checks. Its value lies in preparing a draft analysis that a human reviews and refines.

The fourth benefit relates to interview preparation. The recruiter can walk into a meeting with a list of specific areas to dig into, rather than asking everyone the same set of general questions. This allows better use of the candidate's time and faster separation of actual experience from an attractive description.

The fifth benefit is linked to teamwork and onboarding new people. If every consultant uses their own informal way of reading CVs, it is hard to hand over projects, analyse quality and train less experienced recruiters. AI Scoring can introduce a common point. The team can discuss why they disagree with a score, which requirement was poorly written and how to improve the brief. The difference between human and system evaluation itself becomes learning material.

The sixth benefit is faster response times. In an agency, a few hours can decide which firm presents a valuable candidate first. In in-house recruitment, quick contact affects candidate experience and reduces the risk of losing someone to another process. Scoring can shorten the time from application to the first meaningful analysis, but only when embedded in the workflow and leading to action. If it remains just another number in the system, it brings no real value.

18. Benefits for agencies, in-house HR and Hiring Managers

In a recruitment agency, AI Scoring can primarily support delivery. A recruiter often runs several projects simultaneously, and clients' requirements can be incomplete and change during the process. Contextual evaluation helps grasp a profile faster, prepare the screening and build arguments for the shortlist. It can also facilitate project handovers between consultants. A new person does not have to reconstruct all the logic solely from notes if they can see the approved criteria and how candidates map to them.

The speed of working with your own database is also key. Scoring is not AI Matching and does not automatically search for the best people for a project on its own, but it can evaluate candidates whom the recruiter added from previous processes, referrals or sourcing. Thanks to this, the database stops being purely a CV archive. The condition is data relevance and the legal basis for its further use.

In in-house recruitment, the greatest value often appears with high volumes, but is not limited to them. Even with dozens of applications, the tool can help standardise the initial evaluation across locations and recruiters. It can also reveal a briefing problem. If most candidates receive low scores due to a single requirement, it is worth checking whether the market simply does not supply such profiles, or if the criterion is too narrow or poorly communicated.

The Hiring Manager gets more structured material. Instead of reading every CV without context, they can receive a summary, pros and cons, and areas to check. However, they should not see only the AI score. A large number without access to the document and justification can lead to over-reliance. The best model ensures easy access to both the contextual analysis and the full profile, along with the recruiter's decisions.

AI Scoring can also improve the quality of the conversation between HR and the business. Instead of discussions like 'I like this candidate' versus 'this candidate doesn't fit', you have a list of requirements, evidence, gaps and questions. This does not remove subjectivity, but makes it more visible. If a Hiring Manager rejects a high-match person, it is worth asking which criterion was not written down before. If they accept a low-scoring profile, perhaps the system misinterpreted the data or the requirements needed adjustment.

For the organisation, the benefit is better process analysis. You can check whether high scores translate into advancing through subsequent stages, where recruiters most often disagree with the evaluation, which requirements are regularly unconfirmed, and whether lower-bracket candidates are being unfairly overlooked. However, you must be careful not to treat historical decisions as objective truth. If the past process was biased, blindly learning from its results can entrench the problem.

19. Benefits and risks from the candidate’s perspective

From the candidate’s perspective, AI Scoring can bring real benefits. Every application can be subjected to the same analytical framework, even if it arrived as the two-hundredth. The system can notice experience described differently than in the ad, combine form and CV details, and suggest a question to the recruiter that gives the candidate a chance to explain a missing element. Faster pre-selection can also mean a quicker response and shorter waiting times.

However, benefits are not automatic. A candidate might be unfairly downgraded by an incomplete CV, an unconventional career path, a career change, a gap in employment, or language that differs from the ad's standards. They may not know their data is being analysed by AI, fail to understand the role of the score in the decision, and have no easy way to challenge an error. If an organisation automatically rejects people below a threshold, a single misinterpretation can have a significant impact.

Candidate experience therefore depends on transparency and usage. The organisation should clearly state that AI tools are used in the process, describe their role and highlight whether the decision is made by a human. Candidates should have the opportunity to correct inaccurate data and obtain information on how their data is used in compliance with applicable law. In processes with a higher impact, it is worth providing an appeal channel or human review option.

We should also not shift responsibility entirely onto candidates by stating 'if they didn't write it, it's their own fault'. A CV is a communication tool, but it is the organisation that designs the process and decides what data is needed. If a criterion is key, it should be communicated and gathered in a proportionate way. Scoring can support a fairer evaluation only when the entire process is designed fairly.

20. Where does AI Scoring give the most value?

AI Scoring does not deliver the same value in every process. The most obvious application is recruitment with a high volume of applications, where reading every document manually with equal attention is difficult. This includes popular junior roles, remote positions, international recruitment and processes promoted on portals that enable quick applying. In these scenarios, scoring can help establish the review order and flag profiles requiring urgent contact.

The second application is projects with clearly defined, verifiable requirements. If a role requires specific technologies, licences, experience in a specific environment, sales model or operational scale, AI can compare these elements with documents. The more vague and subjective the criteria, the lower the value of automatic scoring. Scoring is much better at answering 'did the candidate run ERP system implementations in organisations over 500 people?' than 'do they have leadership charisma?'.

The third area is reusing the candidate database. A recruiter can add people known from previous processes to a new project and evaluate their profiles against new requirements. However, you must check data relevance and the legal basis for further processing. Scoring on an outdated CV can create a seemingly precise but low-value result.

The fourth application is agency recruitment, where the speed of preparing a shortlist and the quality of candidate presentation matter. Contextual evaluation can help the consultant prepare the screening, spot risks and build arguments for the client. However, it should not be copied directly into the candidate profile without checks. AI-generated phrasing can be too categorical, imprecise or inappropriate for an external recipient.

Scoring may have less value in executive search processes, where candidate numbers are small and key information comes from interviews, market reputation, references and understanding the complex context of the organisation. Even there, it can support structuring data, but it is hard to expect a CV score to capture the ability to lead transformation, build board trust or operate in a specific ownership culture.

Scoring should not be used just because the feature is available. Before implementing, it is worth answering three questions: what specific problem are we solving, how will we check score quality, and what should the user do after receiving it. If there is no clear answer, another number in the ATS can increase complexity instead of productivity.

21. How to implement AI Scoring in your recruitment process?

A good implementation starts with the process, not a button in the system. The first step should be selecting a limited group of roles for which requirements are relatively clear and a sufficient number of applications are available for testing. Running a pilot across all projects simultaneously makes it hard to identify what works and what needs refinement.

Next, you need to define the rules for creating requirements. Who prepares the criteria, who approves them, how are mandatory and nice-to-have conditions distinguished, how often can they be changed, and what happens to scores after a change? It is worth preparing a simple brief template. It should require describing the goal of the role, key results, essential competencies, acceptable alternatives and areas to verify during the interview.

The next step is testing on known profiles. The team can choose sample CVs of candidates they evaluated previously, run the scoring, and compare the result with recruiters' independent assessments. The goal is not for the system to perfectly replicate every historical decision. History can contain errors. The goal is to check whether justifications are logical, whether the system recognises key experiences, how it treats missing data, and whether undesirable patterns emerge.

User training is essential. It should cover not only feature usage, but also score interpretation, automation bias, the difference between missing information and lack of competence, data protection rules and escalation scenarios. A recruiter should know when they can use scoring for prioritisation and when they must not rely on it without additional analysis.

It is worth establishing a human review practice. For instance, you could require that before rejecting candidates from a specific bracket, the recruiter must review the document and note their own reason for the decision. You can also randomly audit a portion of low-scoring applications. It is important that oversight is real, not formal. Clicking 'approve' without reading the profile is not meaningful human involvement.

Finally, you need to define monitoring. Monthly or quarterly, the team can analyse score distribution, differences between roles, score alignment with recruiter decisions, the number of corrections, profiles bypassed by scoring, and candidate feedback. The implementation should have an owner responsible for both efficiency and risk.

22. Most common mistakes when using AI Scoring

The first mistake is running scoring on a weak brief. A generic job title and a copied job ad do not form an evaluation logic. If the Hiring Manager cannot specify which requirements are truly important, AI will not establish this reliably either. It can suggest criteria, but a human must verify them.

The second mistake is equating a high score with a hiring recommendation. Scoring relates to data matching against criteria at a specific stage. It does not cover everything that drives job success. A candidate might have perfectly described experience but low motivation, fail to confirm declared skills, or reject the offer terms.

The third mistake is automatic rejection based on a threshold. This setup is tempting because it maximises time savings, but it can lead to systematically overlooking people with incomplete CVs and increases the impact of a single error. It can also turn a supporting tool into a system actually making decisions, which has legal consequences.

The fourth mistake is failing to distinguish between data and conclusions. If an AI summary sounds convincing, the user might forget it is an interpretation. Every significant conclusion should be linkable to a document fragment or marked as requiring confirmation.

The fifth mistake is analysing too broad a scope of data. More data is not always better. Photos, age, address, marital status, health information or social media activity can be unnecessary and risky. The organisation should consciously define which sources and fields can affect scoring.

The sixth mistake is a lack of testing after changing a model, prompt or logic. An AI feature might behave differently after an update, even if the interface looks the same. The provider should manage changes, and the client should know when scores might shift.

The seventh mistake is assuming a human will automatically correct system issues. Research on human-AI collaboration shows that people can over-rely on recommendations and align their own decisions with the score. Oversight requires competence, time and the right to disagree with the system, not just a recruiter's name in the process.

23. Limitations of AI Scoring - what does the system not know?

AI Scoring sees the information it has been given, and nothing more. It does not know competencies omitted from the CV, does not know if an achievement was described accurately, does not independently confirm declaration truthfulness, and does not observe the candidate at work. It can draw conclusions from text, but should not replace skills tests, behavioural interviews, references or evaluations performed in a real context.

The system may struggle with unconventional career paths. A person transitioning from the military to business, from academia to product, from entrepreneurship to corporate, or from another industry may possess transferable skills that they do not describe in the target role's language. A semantic model can spot some of these, but requirements based on historical job titles might still downgrade the score.

Soft skills are also difficult. A CV might state 'excellent communication skills', but this is not reliable evidence. The system should not award a high score for the declaration alone. It can, however, point to experiences where communication was likely important and suggest interview questions. Assessing actual behaviour, however, requires other methods.

AI also does not know the full context of the organisation. Two companies might use the same job title for completely different scopes. Managing a team of ten in a stable department is different from building a team from scratch in a startup. A budget of one million in one industry might have a different meaning than in another. Therefore, requirements must describe context, not just labels.

Data and document errors are another limitation. A parser can misread a table, the system can mix up dates, and a LinkedIn profile can be outdated. The more precise the numerical score, the easier it is to forget about input uncertainty. The 'garbage in, garbage out' rule remains relevant in the era of language models.

Finally, AI does not bear responsibility for the decision. It can prepare an assessment, but it does not talk to the candidate, does not know all the circumstances, and is not accountable to the organisation, client or regulator. Responsibility remains with the people and entities that design and use the system.

24. Bias, discrimination and the illusion of objectivity

One of the most tempting promises of AI in recruitment is objectivity. A machine is not tired, does not have a favourite university, and does not evaluate a photo like a human - provided the photo and other unnecessary data do not influence the model. However, this does not mean the score is neutral. The system can replicate biases present in the data, criteria, labelling method or the process design itself.

Bias can enter scoring at multiple levels. Criteria might reward paths more accessible to specific groups, such as an unjustified requirement for continuous employment. Historical data might reflect the organisation's past preferences. Company and university names can act as proxy indicators of social status. CV language can differ by culture and gender. Missing data might more often affect people with less conventional career histories.

Language models can also draw conclusions based on indirect traits. Research on CV evaluation by LLMs shows that changing elements suggesting gender or origin can affect the score, though the scale and direction of bias vary between models. The FAIRE benchmark of 2025 demonstrated the presence of some level of bias in tested models during direct CV scoring and ranking [7]. This does not mean every recruitment system will behave identically, but it confirms the need to test specific configurations rather than relying on general model manufacturer claims.

The illusion of objectivity is particularly dangerous because numbers look scientific. A score of 78.4 can give the impression of a measurement similar to temperature, though it is the result of chosen criteria and text interpretation. Precision in display is not proof of accuracy. The organisation should ask: what does the score measure, how was it tested, for which groups does it work less well, and what are the consequences of an error.

Mitigating bias is not just about removing names and photos. Anonymisation helps, but other elements can still reveal or indirectly indicate protected traits. What is needed is score testing, analysing differences between groups where legal and methodologically justified, reviewing criteria, error reporting mechanisms and the ability to withdraw a feature if risk cannot be mitigated.

It is also worth remembering that humans are not a neutral benchmark. AI can, in some situations, limit recruiter randomness and unconscious preferences. The goal should not be to prove that 'AI is more objective than a human' or vice versa. The goal is to design a process where both parties' errors are visible, measured and correctable.

25. Human-in-the-loop - what does real human oversight mean?

The phrase human-in-the-loop appears in almost every description of responsible AI. Simply placing a human in the process is not enough. If a recruiter sees a score, automatically accepts the recommendation and has no time for analysis, their involvement is formal. Real oversight means the ability to understand the system's capabilities and limitations, check the basis of a score, challenge it, and halt or change actions when an issue arises.

The AI Act, in relation to high-risk systems, requires designing solutions so they can be effectively overseen by natural persons. It also highlights the risk of automation bias, i.e., automatic or excessive reliance on system output [4]. The person exercising oversight should have appropriate competence, training, authority and support. In a recruitment context, this means a junior recruiter cannot be the only 'safety layer' if the organisation expects them to process massive volumes and rates them solely on speed.

Real oversight can involve several practices. The recruiter should see the justification and source material. They should be able to change a decision without negative organisational consequences. The system should log corrections and enable analysis of where humans most often disagree with the AI. For automated actions, a kill switch mechanism is needed. Candidates should be able to report inaccuracies, and the organisation needs a procedure to handle such reports.

It is worth separating human review from doing all the work manually. Oversight does not mean a human has to re-analyse every element from scratch, because then the tool loses its purpose. They should, however, check decisive issues, especially before a negative decision. A proportionate approach can be applied: more automation in low-risk administrative tasks, more control where the result affects the candidate's access to the next stage.

The organisation should also train users to recognise situations where the system might fail: unconventional CVs, industry changes, documents in another language, employment gaps, incomplete data, conflicting sources or very narrow criteria. Human-in-the-loop is only effective when the human knows what to look for.

26. AI Scoring and GDPR

The AI Act does not replace GDPR. If AI Scoring processes candidates' personal data, the organisation must still meet obligations under data protection laws. This covers legal basis, transparency, data minimisation, purpose limitation, accuracy of information, security, retention and the exercise of data subject rights.

The first question is: who is the controller and who is the processor. The company running the recruitment usually determines the purposes and means of processing candidates' data, and the ATS provider acts in a specific capacity as a processor. However, the actual split depends on the service design, including whether the provider or its sub-processors use the data for their own purposes, model training, testing or service improvement. Roles, instructions and sub-processors should be clearly described in the contract.

The second question relates to purpose and legal basis. You cannot assume that candidate consent solves every problem. The basis depends on the type of recruitment, stage, national law and whether data is to be used in future projects as well. Special categories of data must be analysed separately. If a model can technically infer origin, health status, views or other sensitive traits, this does not mean the organisation can use such inferences in scoring.

The third area is minimisation. The system should only analyse information relevant to the purpose. Photos, age, home address, family situation or activity unrelated to work should not affect the assessment just because they are in the document or profile. Particular caution is needed when using data from LinkedIn and other external sources. The availability of information on the internet does not automatically mean free rein to process it.

The fourth area is the information obligation. Candidates should receive clear information that their data is analysed using AI, for what purpose, what data categories are used, where they come from, who they are disclosed to, how long they are stored, and what rights they have. Information should not hide key facts under a generic phrase like 'we use modern technologies'. It is also worth explaining the role of the score: whether it only supports the recruiter's work or triggers further actions.

Article 22 of the GDPR is of particular significance, as it concerns decisions based solely on automated processing which produce legal effects or similarly significantly affect individuals [2]. Scoring that supports a recruiter does not automatically mean a solely automated decision. However, if the score leads, without real human involvement, to rejecting an application, blocking the next stage or hiding a profile from the recruiter, the risk of entering this area is much higher. The 'human in the loop' must have a real impact, not just a technical option to approve the system's decision.

The fifth area is data accuracy and the ability to correct it. If a parser misreads a date, LinkedIn is outdated or a form contains an incorrect answer, the score might be wrong. Candidates should be able to correct data, and the organisation should establish whether scoring is recalculated after a correction.

Before implementation, it is worth conducting a Data Protection Impact Assessment (DPIA), especially when technology systematically evaluates individuals and can significantly affect their situation. The UK’s ICO is not an EU law-applying body, but its audits of recruitment tools serve as a practical reference. The regulator highlighted, among other things, the need to perform a DPIA at the procurement stage, clearly divide responsibilities, limit data, test fairness and transparently inform candidates [6].

Retention is also important. Data and scoring results should not be stored indefinitely just because a large database might be useful someday. The organisation should define how long it stores CVs, answers, scores, justifications and logs, and what happens to them after consent is withdrawn or the processing purpose expires.


Legal note: this section is for informational purposes and does not replace a legal analysis of a specific implementation. The scope of obligations depends on how the system operates and is used, the data types, the roles of the parties and national regulations.


27. AI Scoring and the AI Act

The EU AI Act applies a risk-based approach. Annex III lists AI systems intended to be used for recruitment or selection of persons, notably to place job adverts, screen and filter applications, and evaluate candidates [1]. This means AI Scoring used in recruitment is in an area of high regulatory focus. However, we should not simplify this to state that every feature containing AI and CVs automatically has the same legal status.

Classification depends on the intended purpose, design and actual impact on the process. Article 6(3) provides a possibility to deem some systems listed in Annex III as non-high-risk if they do not pose a significant risk of harm to the health, safety or fundamental rights of natural persons, including by not materially influencing the outcome of decision-making, and meet specific conditions. However, a system remains high-risk if it profiles natural persons, and a provider who considers that a system is not high-risk must document its assessment. The Commission published draft detailed guidelines on this classification in 2026 [3].

In practice, a feature that assigns a score to candidates and materially affects review order, advancement to the next stage or rejection requires highly cautious analysis. You cannot resolve classification solely based on names like 'assistant', 'recommendation' or 'scoring'. What matters is the intended use described by the provider and how the organisation actually uses the score.

For systems classified as high-risk, the AI Act outlines extensive requirements for providers. These include a risk management system, quality and data governance, technical documentation, automatic logging of events, transparency for the user, human oversight capabilities, and appropriate levels of accuracy, robustness and cybersecurity, alongside post-market monitoring. Depending on the scenario, a conformity assessment, declaration of conformity, CE marking and registration in the EU database are also required.

Obligations also apply to entities using the system, i.e., deployers. The recruiting organisation should use the solution in accordance with instructions, assign competent persons for oversight, ensure the relevance of input data within its control, monitor operation, and react to issues. In specific situations, individuals subject to a high-risk system must be informed of its use and the type of decisions supported [1]. The AI Act also provides a right to a clear and meaningful explanation of the role of the AI and the main elements of the decision in situations outlined in Article 86.

Human oversight cannot be formal. The person overseeing should understand the system's capabilities and limitations, be aware of automation bias, have access to sufficient information, and possess genuine authority to bypass, reverse or stop operations. A recruiter who automatically accepts every suggestion does not provide real oversight just because they formally clicked a button.

AI literacy is also significant. Organisations should ensure that individuals using AI possess an adequate level of knowledge and competence, taking into account their role, experience, context of use and the persons affected by the system. In practice, recruiter training should cover not only feature usage, but also score interpretation, distinguishing lack of data from lack of competence, discrimination risks, data protection and error reporting procedures [5].

Not every private company using a recruitment system will be required to carry out a fundamental rights impact assessment under Article 27. This obligation applies to specific categories of deployers and uses. Nevertheless, an impact assessment can be a good practice, and a DPIA under GDPR may be required regardless of the AI Act.

The timeline is important. According to current Commission information, following amendments introduced by the AI Omnibus, high-risk system requirements in Annex III, covering employment among others, are to apply from 2 December 2027 [4]. This does not mean organisations can ignore the topic until late 2027. Obligations regarding AI literacy apply from 2 February 2025, and oversight of them starts according to the Commission's timeline in 2026 [5]. GDPR, anti-discrimination laws and employment law apply regardless of this timeline right now.

For organisations, the practical takeaway is simple: do not wait until the last minute. It is worth creating a registry of used AI features, defining their purpose, providers, data sources, impact on decisions, oversight mechanisms and issue reporting procedures. Providers should be asked about classification, documentation, tests, logs, sub-processors, processing locations, model changes and compliance plans. The AI Act is not solely a legal team problem. It impacts product design, procurement, HR, security and the daily work of recruiters.


The safest assumption is: the greater the impact of the score on a candidate's chance of staying in the process, the stronger the documentation, human control and monitoring should be.


28. How to choose an AI Scoring tool?

Comparing tools solely based on demonstration quality is risky. On a demo, you can select perfectly written CVs and roles for which the score looks convincing. Real quality is revealed with incomplete, unusual, multilingual, conflicting documents coming from different industries. Therefore, purchase should involve testing on your own, properly secured data or a prepared set of representative profiles.

The first group of questions relates to features. What is evaluated: the application, the candidate, or both? How are requirements defined? Can the user edit and weight them? How does the system distinguish missing data from unmet criteria? Does it show the justification and source fragments? Does it generate questions? Can the score be recalculated after data changes? Does scoring work in multiple languages?

The second group relates to control. Can automated actions be turned off? Does the system log who triggered the evaluation and which requirement version was used? Can the user report an error? Does the provider inform about model changes? Can the administrator restrict access to the feature and set usage rules?

The third group relates to quality. How does the provider test accuracy, stability and bias? On what roles and languages? Do they share the methodology and known limitations? How do they react to errors? Do they measure differences between groups in a legally compliant way? Does the model generate deterministic or sufficiently repeatable answers?

The fourth group relates to data and security. Where is data processed and stored? Is it sent to third-party model providers? Is it used to train public or proprietary models? How long is it kept in logs? What are the bases for transfers outside the EEA? What do encryption, access control, audit and data deletion look like? Who is the sub-processor?

The fifth group relates to law. How does the provider classify the feature under the AI Act and on what basis? What is the compliance plan? What instructions will the deployer receive? Will documents needed for a DPIA, risk assessment and information obligations be available? Does the contract clearly define roles in data protection?

The sixth group relates to implementation. Does the provider train the team not only on usage, but on responsible interpretation? Do they help calibrate criteria? Do they allow starting with a pilot? Can the client measure effects? The best algorithm won't help if users don't trust the system or trust it too much.

29. How to measure the effectiveness of AI Scoring?

Implementing AI Scoring should have a measurable goal. The simplest indicators relate to efficiency: time from application to first analysis, average profile review time, number of applications analysed in a given timeframe and time to first contact. This data shows whether the tool actually relieves the team.

The second group of indicators relates to utility. You can measure how often the recruiter agrees with the score, how often they correct it, whether they use suggested questions, and whether summaries reduce the need for re-reading. However, agreement with a human is not a measure of truth on its own. It is worth analysing justifications and cases of discrepancy, not just the agreement percentage.

The third group relates to process quality. Do candidates with high scores more often pass screening after competency verification? How many valuable candidates were found in lower brackets? Does the score differ significantly between recruiters using the same criteria? Are fewer candidates waiting long for a response?

The fourth group relates to risk. Complaints, data errors, inaccurate explanations, unequal score distributions and automated actions with negative outcomes should be monitored. Where possible and in compliance with the law, fairness tests should be conducted. It is not enough to check the system once before implementation. Roles, data, models and how users work change.

The fifth group relates to human behaviour. Are recruiters reviewing lower-scored profiles? Are they rejecting candidates faster but without justification? Does the Hiring Manager look at the score first, and only then the CV? You can analyse logs and conduct short qualitative audits. The goal is to detect automation bias and situations where the feature takes on a larger role than planned.

Success should not be evaluated solely by the number of hires with high scores. If recruiters primarily contact this group, such a result is partly a self-fulfilling prophecy. Comparisons, control samples and conscious analysis of candidates outside the top positions are needed.

30. How does AI Scoring work in Recruitify?

AI Scoring in Recruitify is used to evaluate an application or candidate in the context of a specific recruitment project. The feature does not assign a permanent value to a person and does not automatically search for individuals in the database. The starting point is a profile already within the project, and the requirements defined for that recruitment.

The process begins with the Requirements field. The recruiter describes what the project actually requires: what experience is essential, what elements increase fit, which gaps are acceptable, and what should be verified during the interview. The more specific and job-related the requirements, the more useful the assessment will be. Generic phrases like 'good communicator', 'dynamic person' or 'cultural fit' do not give the AI or recruiter a sufficiently precise point of reference.

Recruitify analyses available information from the CV, LinkedIn profile and answers provided in the application form. Thanks to this, the score does not have to rely on a single document. The form can provide data on availability, expectations, working model, licences or experience in an area the candidate did not describe in their CV. LinkedIn can complement history, but is not treated as automatically more reliable than the application document. Discrepancies should lead to verification.

The result has both a numerical and contextual dimension. Recruitify awards a score from 0 to 100, but does not limit itself to the number. The user receives a candidate description in the context of the role, strengths, weaknesses and suggestions for questions to ask during the meeting. The goal is to help prioritise work, grasp the profile faster and prepare the screening better.

A score from 0 to 100 is not a probability of hire. A candidate with a score of 90 is not 'ten per cent better' than someone with a score of 80. The number shows the outcome of comparing available data against the criteria of a specific project. The exact same candidate can receive a different score across two recruitments because requirements change. The score can also change after data updates or refining the Requirements.

After generating the assessment, the recruiter should read the justification and check it against the source material. Particular attention should be paid to weaknesses and areas the system could not confirm. A lack of mention in the CV does not have to mean a lack of competence. In many cases, the best next action is to use a suggested question or ask the candidate to supplement information.

Recruitify clearly separates AI Scoring from AI Matching. Scoring evaluates an application or candidate already being analysed in a project. AI Matching, which is under development, is designed to search the database and suggest individuals potentially fitting the project to the recruiter. These are two different tasks: evaluating a specific profile versus finding and recommending profiles from the database.

The best practice is to start with tests on a few projects with clear requirements. The team can compare scores with their own evaluation, check justification quality, review a portion of low-scoring profiles, and refine how Requirements are written. AI Scoring is meant to support the recruiter's decision, not automatically reject candidates.





31. The future of AI Scoring

AI Scoring will shift from a simple percentage towards more complex, but also more controlled, decision support. The most valuable solutions will not try to create a single 'truth about the candidate'. They will combine different sources, show uncertainty, point out the basis of conclusions and tailor the evaluation method to the stage of the process.

We can expect a greater role for structured rubrics. Instead of a general prompt like 'evaluate the candidate', the system will work on criteria with a definition, weight, source of proof and level of certainty. Development may involve comparing candidates, but should avoid situations where a difference of a few points is treated as an objective advantage of one human over another.

The second direction will be integration with subsequent stages. Scoring before the interview can generate questions, and after the interview, the system can help organise notes according to the same scorecard. It will be important to separate declared from confirmed information and ensure that interview data is processed in compliance with the law.

The third direction will be testing and auditability. Regulations, client requirements and market maturity will increase pressure to document models, logic, data, accuracy and bias. Providers will have to show not just an attractive interface, but also a risk and change management process.

The fourth direction will be combining scoring with AI Matching and talent intelligence. The system can first suggest profiles from the database, then evaluate them in the project, and later support communication and process analysis. This combination increases productivity but also the risk of creating a closed loop of automated recommendations. Every stage should have a clear goal, separate metrics and the possibility of human intervention.

The most significant shift, however, will be organisational. Companies will stop asking simply 'do we have an AI feature?', and start asking 'what decision does this feature support, on what data, with what risk, and who is responsible for control?'. It is the answers to these questions that will determine whether AI Scoring improves recruitment or merely gives old problems a modern interface.

32. AI Scoring FAQ - Frequently Asked Questions

What is AI Candidate Scoring?

AI Scoring is an AI-supported evaluation of a specific application or candidate against the requirements of a specific project. It can take the form of a numerical score, a contextual description, or both elements simultaneously. It should not be understood as a universal assessment of human worth.

Does AI Scoring evaluate the candidate or the application?

It can evaluate both. An application is a specific entry to a project, often alongside form responses. A candidate can be added to the project by a recruiter without a new application. In both cases, the result should relate to the fit with the given project's requirements.

Are AI Scoring and AI Matching the same thing?

No. AI Scoring evaluates an application or candidate already being analysed in a project. AI Matching searches the database and suggests candidates potentially fitting the project. The first feature answers 'how does this profile stack up?', and the second 'who is worth finding and considering?'.

Does a score of 85 mean an 85 per cent chance of being hired?

No. A score from 0 to 100 is a fit scale used by a specific system. It is not a probability of hiring or job success. Its meaning depends on criteria, data and calculation methods.

Is a candidate with a score of 90 better than a candidate with a score of 80?

This cannot be determined solely on the basis of scoring. The difference might stem from CV completeness, how experience was described or a single heavily-weighted criterion. The score helps set analysis priority but does not replace evaluating the entire profile.

Can AI Scoring automatically reject candidates?

Technically, the score can be linked with automation, but automatically rejecting solely based on a threshold significantly increases the risk of error, discrimination and entering the territory of solely automated decisions. A safer approach is using scoring for prioritisation and ensuring a real human review.

Does the recruiter have to read the whole CV after receiving a score?

Scoring can shorten the analysis and point out key fragments, but before a significant decision, the recruiter should check the source material. The review scope can depend on the stage and risk. The score and generated summary alone should not be the sole basis for rejection.

What data does AI Scoring analyse in Recruitify?

In Recruitify, scoring can factor in project requirements and information from the CV, LinkedIn profile and application form responses. The scope of information used depends on the data available in the profile and project.

Is LinkedIn more reliable than a CV?

Not automatically. A LinkedIn profile can be newer or broader, but it is also created by the candidate and can contain gaps. It is best to treat sources as complementary and flag inconsistencies for verification.

What happens if important information is missing from the CV?

A good system should indicate a lack of confirmation, rather than automatically assuming a lack of competence. The recruiter can check other sources or ask a question during the interview. Critical data is best collected in the application form.

Does AI Scoring recognise synonyms and context?

Modern language models can connect similar meanings and recognise experience described in different words. However, they are not infallible. Semantic connections should be explained and checkable in the document.

Does AI Scoring detect if a CV was written by AI?

This is not the primary task of scoring, and there is no foolproof method to detect every AI-generated text. It is more important to verify whether the declared competencies are real. Suggested questions can help verify them during the interview.

Can AI Scoring detect lies in a CV?

No. It can spot inconsistencies between sources or vague fragments, but it does not confirm the truth of experience. Verification requires a conversation, test, references or documents.

Does AI Scoring evaluate soft skills?

It can analyse information suggesting certain experiences, but a CV is not a sufficient source for a reliable evaluation of empathy, communication, resilience or leadership style. The system should suggest questions rather than issue categorical evaluations.

Can the score change?

Yes. The score can change after updating the CV, supplementing the form, changing project requirements or updating features. Therefore, it should be treated as the result of a specific analysis at a given time.

Can you compare scores across different projects?

Usually, this makes no sense. Every project has different criteria, weights and score distributions. A score of 70 in one recruitment does not necessarily mean the same as 70 in another.

Can AI Scoring reduce recruiter bias?

It can increase consistency and limit some random evaluations, but it can also introduce or reinforce bias stemming from criteria, data and the model. Testing, monitoring and the ability to correct are needed.

Does removing names and photos solve the discrimination problem?

Not fully. Anonymisation limits some signals, but other information can indirectly indicate gender, age, origin or social status. Bias must be analysed across the entire process.

Do you need to inform candidates about using AI?

In many situations, information obligations arise under GDPR, and for high-risk systems, the AI Act outlines additional rules for informing individuals subject to the system. The scope of information depends on the specific implementation, but transparency should be a standard even when the law only requires a minimum.

Is AI Scoring subject to the AI Act?

Systems intended to be used for screening and filtering applications and evaluating candidates are listed in Annex III of the AI Act as employment-related uses. The final classification of a specific feature depends on its purpose, impact, usage and exceptions in Article 6(3). This requires an individual analysis.

When do AI Act rules for high-risk systems in recruitment start to apply?

According to the current timeline, following the entry into force of the AI Omnibus, rules for high-risk systems in Annex III are to apply from 2 December 2027. Some other provisions, including those on AI literacy, apply earlier.

Is AI Scoring compliant with GDPR?

The feature name alone does not determine compliance. Compliance depends on the legal basis, data scope, transparency, retention, security, division of roles, the ability to realise candidate rights, and how decisions are made. A tool can support a compliant process, but it does not 'solve GDPR' for the organisation.

Is a DPIA required?

In many implementations of candidate scoring, a Data Protection Impact Assessment may be required or at least highly justified due to systematic evaluation of individuals and potentially significant impact. The decision should be made based on the specific process and consultation with a DPO or lawyer.

What does human-in-the-loop mean?

It means real, not formal human involvement. The recruiter should understand limitations, see the score basis, be able to challenge it, and have actual authority to make a different decision. Merely clicking confirmation is not enough.

Will AI Scoring replace recruiters?

It should not. It automates part of the analysis and structuring of information, but it does not replace conversations, understanding context, evaluating motivation, negotiating or responsibility for the decision. It shifts the weight of a recruiter's work from manual screening to interpretation, verification and human connection.

For what type of recruitment is scoring not a good solution?

It can deliver less value in processes with very few candidates, unclear criteria, a key role for confidential market knowledge or a strong emphasis on competencies evaluated only in interaction. It should not be forced just because it is available.

How to start testing AI Scoring?

It is best to select a few roles with clear requirements, prepare criteria, test the feature on representative profiles and compare results with recruiters' independent evaluations. Then, establish human review rules, train users, and measure both time, quality and risk.

33. Summary - AI Scoring does not choose the person, it supports the evaluation

AI Candidate Scoring is one of the most practical applications of artificial intelligence in modern recruitment systems. It can help teams structure incoming applications faster, match profiles to requirements more consistently, prepare interviews better and collaborate more smoothly with Hiring Managers or clients. Its value, however, does not come from the score of 0 to 100 itself. What matters most is the context, the explanation, the quality of criteria, and the ability to verify and challenge the evaluation.

Scoring is not AI Matching. It is not primarily used to find candidates in the database, but to evaluate an application or candidate in a specific project. It is not a CV parser, though it uses extracted data. Nor is it a simple killer question or an automated rejection decision. It can link with these features, but each has a different task and a different risk profile.

The biggest mistake would be to assume the number closes the discussion. The score can help set the order of work, but it should not define a candidate's value or replace a conversation. Documents are incomplete, criteria can be flawed, and models can generate inaccurate or biased conclusions. That is why a good system highlights not only fit, but also information gaps, limitations and questions for verification.

In 2026, a responsible implementation of AI Scoring requires combining three perspectives. The first is productivity: whether the feature genuinely saves time and improves process quality. The second is recruitment practice: whether criteria are job-related and the recruiter retains their own judgment. The third is governance: whether the organisation controls data, bias, security, transparency and compliance with GDPR and the AI Act.


The best AI Scoring does not tell a recruiter whom to hire. It helps them see faster what is already known, what is still unknown and what is worth asking before they make a decision.


Sources and Further Reading

[1] European Parliament and Council of the EU, Regulation (EU) 2024/1689 - Artificial Intelligence Act, notably Annex III point 4 on employment, recruiting and candidate evaluation: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

[2] European Parliament and Council of the EU, Regulation (EU) 2016/679 - GDPR, notably Article 22 on automated individual decision-making: https://eur-lex.europa.eu/eli/reg/2016/679/oj

[3] European Commission, Draft Commission guidelines on the classification of high-risk AI systems, including materials on Annex III AI Act, 2026: https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems

[4] European Commission, Guidelines for providers and deployers of AI high-risk systems - current timeline of high-risk rules application, 2026: https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems

[5] European Commission, AI Literacy - Questions & Answers, current information on Article 4 AI Act: https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers

[6] Information Commissioner’s Office, Thinking of using AI to assist recruitment? Our key data protection considerations, 2024: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/11/thinking-of-using-ai-to-assist-recruitment-our-key-data-protection-considerations/

[7] Wen A. et al., FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations, 2025: https://arxiv.org/abs/2504.01420

[8] Hunkenschroer A. L., Luetge C., Ethics of AI-Enabled Recruiting and Selection: A Review and Research Agenda, Journal of Business Ethics, 2022: https://link.springer.com/article/10.1007/s10551-022-05049-6

[9] Information Commissioner’s Office, AI tools used in recruitment - audit outcomes and recommendations, 2024: https://ico.org.uk/action-weve-taken/audits-and-overview-reports/2024/11/ai-tools-used-in-recruitment/

[10] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0): https://www.nist.gov/itl/ai-risk-management-framework

[11] Information Commissioner’s Office, Automated decisions can streamline the hiring process with the right safeguards in place, 2026: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2026/03/automated-decisions-can-streamline-the-hiring-process-with-the-right-safeguards-in-place/

[12] Recruitify, What is an ATS? A complete guide to applicant tracking systems (2026): https://www.recruitify.ai/blog/what-is-an-ats-complete-guide-to-applicant-tracking-systems-(2026)/


Editorial Note: this article describes principles and good practices at a general level. It does not constitute legal advice. Classification and obligations concerning a specific AI system depend on its purpose, design, data, implementation method and actual impact on decisions.

News & Updates

Stay up-to-date with the latest innovations, features, and tips about Recruitify!

First Name
Email

By providing your email address within the newsletter sign-up form, you confirm its processing to send marketing information regarding the Administrator’s products and services. The Administrator of your personal data processed for the abovementioned purposes is Recruitify Spółka z o.o., based in Warsaw, Poland (KRS 0000709889). For more information on the principles of personal data processing and the rights of data subjects, please check the Privacy Policy.

Share

Published

Category

Applicant Tracking System

Author

Iwo Paliszewski

AI candidate scoring

Last updated:

AI Candidate Scoring: What is it, how does it work, and how to use it responsibly? The ultimate guide (2026)

Innovations

Iwo Paliszewski

Iwo Paliszewski

You search Google for ‘AI Candidate Scoring’ and, more often than not, you are presented with one of two answers. The first goes: artificial intelligence reads CVs, assigns a matching percentage to the candidate and highlights the top talent. The second warns that the algorithm takes over the recruiter’s decisions, creates soulless rankings and can automatically shut someone out of a job opportunity. Both answers simplify the topic so much that instead of clarifying it, they merely build further misunderstandings.

AI Scoring does not have to be a magical autopilot selecting the ‘best people’, nor a black box passing judgment. In a well-designed process, it is a tool supporting the assessment of a specific application or candidate against the requirements of a specific project. It can present a numerical score, for example from 0 to 100, but it can also provide a contextual assessment: a profile summary, strengths, weaker areas, information gaps and questions worth asking during an interview. It delivers the greatest value when it does not replace human judgment, but helps structure, speed up and better justify it.

This guide answers the question ‘what is AI Candidate Scoring’ as it should be answered in 2026: comprehensively, practically and without marketing shortcuts. We explain exactly what is assessed, where the score comes from, how scoring differs from AI Matching, CV parsers and disqualifying questions, how to set up solid criteria, what data can be analysed, where the risks lie, and what GDPR and the EU AI Act say about AI used in recruitment.


Quick Definition
AI Candidate Scoring is an AI-supported assessment of a specific application or candidate against the requirements of a given recruitment project. The result can be numerical, contextual, or a combination of both. A good scoring system does not just show a number; it also explains what influenced the rating, highlights the strengths and weaker areas of the profile, and suggests questions for further verification. It is not the same as AI Matching, which is designed to search the database and suggest candidates who fit a project.


Table of Contents

1. What is AI Candidate Scoring?

2. What exactly does AI Scoring assess: the human, the candidate or the application?

3. Why has AI Scoring become important right now?

4. Numerical and contextual assessment - two dimensions of good scoring

5. How does AI Scoring work step-by-step?

6. What data can AI Scoring use?

7. Why are project requirements more important than the AI model itself?

8. Must-haves, nice-to-haves and disqualifying criteria

9. AI Scoring vs. keyword search

10. What does a score from 0 to 100 mean?

11. Strengths, weaknesses and interview questions

12. Missing information is not the same as a lack of competence

13. AI Scoring vs. the traditional candidate scorecard

14. AI Scoring vs. AI Matching - a key distinction

15. AI Scoring vs. CV parser, AI summary and killer questions

16. AI Scoring in practice - three assessment examples

17. Benefits for the recruiter and the recruitment team

18. Benefits for agencies, in-house HR and Hiring Managers

19. Benefits and risks from the candidate’s perspective

20. Where does AI Scoring deliver the most value?

21. How to implement AI Scoring in your recruitment process?

22. Most common mistakes when using AI Scoring

23. Limitations of AI Scoring - what does the system not know?

24. Bias, discrimination and the illusion of objectivity

25. Human-in-the-loop - what does real human oversight mean?

26. AI Scoring and GDPR

27. AI Scoring and the AI Act

28. How to choose an AI Scoring tool?

29. How to measure the effectiveness of AI Scoring?

30. How does AI Scoring work in Recruitify?

31. The future of AI Scoring

32. FAQ - Frequently Asked Questions

33. Summary - AI Scoring does not choose the person, it supports the evaluation


1. What is AI Candidate Scoring?

AI Candidate Scoring is a method of supporting the assessment of an application or a candidate's profile using artificial intelligence. The system analyses the available information about the candidate, compares it to the requirements of a specific project and prepares a result designed to help the recruiter in their next steps. Depending on the solution, the result can take the form of a number, percentage, category, description or a set of several elements. In Recruitify, this includes a score from 0 to 100, a contextual description of the candidate, a comparison of strengths and weaker areas, and suggestions for questions to ask during the interview.

The key phrase here is 'against the requirements of a specific project'. AI Scoring should not evaluate whether someone is a good or bad candidate in general. There is no universal professional value of a human being that can be honestly wrapped up in a single number. The exact same candidate might get a high score in a project looking for someone to scale B2B sales in the Polish market, and a significantly lower score in a project requiring experience in international enterprise sales. This does not mean their skills changed between the two measurements. The context and objective of the evaluation changed.

That is why it is safest to think of AI Scoring not as an evaluation of a person, but as a structured analysis of how well the currently available information fits. This phrasing is longer, but much more precise. The system does not see the entire professional history, potential, character, motivation or behaviour of the candidate in their future job. It sees a specific set of data and compares it against a specific set of criteria. Its utility therefore depends entirely on the quality of both sides of this comparison.

It is also worth distinguishing scoring from ranking itself. Scoring is the process of evaluating and explaining a result. Ranking is one of the possible ways to use this result, for example, by sorting applications from the highest to the lowest score. You can have scoring without an automatic ranking, and you can also create a ranking based on simple rules that have nothing to do with AI. In practice, ATS systems often combine both elements, but from the perspective of a responsible process, they should not be treated as synonyms.


The Golden Rule
AI Scoring should answer the question: ‘What in the available data suggests this candidate fits this role, what raises doubts, and what still needs to be verified?’, rather than: ‘Does this person deserve the job?’.


2. What exactly does AI Scoring assess: the human, the candidate or the application?

In everyday language, people usually talk about 'candidate scoring', but within a recruitment system, the object of evaluation can be both an application and a candidate assigned to a project. This distinction has practical significance. An application is a specific submission to a specific job opening. It can include the CV sent in response to the advert, answers from the application form, the application source, and additional documents or information provided during submission. A candidate, on the other hand, is a person who already exists in the database, who can be added to a project by a recruiter, sourced, or reconsidered in a new recruitment round without submitting a new application.

AI Scoring can therefore assess an incoming application, but it can also be run for a person found previously in the database or added to a project by a recruiter. In both cases, the evaluation should be anchored in the requirements of that specific recruitment. However, the scope of available data may differ. For an application, the system often has a fresh CV and answers to application questions at its disposal. For a database candidate, it can use the profile, previous documents, LinkedIn data, and information gathered in the system. If some of the data is outdated, the score may also be less up-to-date.

This leads to an important conclusion: a score is not a permanent label attached to a person. It should not be logged in anyone's mind or in the organisation's culture that 'Anna is a 64-point candidate' or 'Piotr is a 91 per cent candidate'. The correct phrasing is: 'Based on currently available data, in relation to the requirements of Project X, the system assigned Anna's application a score of 64/100'. The exact same profile in another project might receive a completely different score. Even within the same recruitment, the score can change if the recruiter refines the requirements, the candidate updates their details, or a new CV is uploaded.

This approach also guards against one of the biggest dangers of automation: turning a supporting score into a label that defines a person. In recruitment, the anchoring effect is very powerful. When a recruiter sees 92 points, they may start looking for confirmation of that high score. When they see 48, they might read the profile more critically. That is why explanations are needed alongside the numerical score, and users should be trained to treat the score as the starting point of their analysis, not the conclusion.


3. Why has AI Scoring become important right now?

Candidate scoring is not an invention of the generative AI era. Recruiters have been using scorecards, Excel spreadsheets, killer questions, weighted criteria and simple point systems for years. In many companies, the person managing the recruitment was already assigning points to candidates for experience, language skills, availability or knowledge of a specific technology. The problem was that such a process was time-consuming, highly dependent on user discipline, and often ended up as a paper-based methodology that nobody consistently used after receiving the hundredth application.

The shift today is not so much the existence of a point system itself, but rather the ability to automatically analyse massive amounts of unstructured text. Classic rules are good at dealing with 'yes' or 'no' answers, a specific number of years, or an exact certificate name. Modern language models can additionally compare the meaning of requirements with project descriptions, responsibilities and achievements written in different words. Thanks to this, scoring does not have to be limited to matching exact phrases. It can prepare a draft summary, highlight potential strengths and weaknesses, and suggest questions to help clarify ambiguities.

At the same time, the recruitment environment has changed dramatically. Easy applying, one-click apply forms, automatic alerts and generative tools that help build CVs mean that submitting an application takes much less effort than it did a few years ago. In many processes, the number of applications grows faster than a team's capacity to read them thoroughly. However, a larger volume does not mean more people who actually fit the role. Recruiters must find valuable profiles among documents that are random, mass-sent, linguistically very similar or tailored to match the job advert's keywords.

In such an environment, time is not the only issue. A human does not read the two-hundredth CV with the same freshness as the first. Concentration levels drop, pressure rises, and initial evaluation is increasingly based on a few of the most visible elements. A candidate who applied in the morning might get more attention than an equally good person whose document landed in the system at the end of the day. AI Scoring can help apply the same analytical framework to all submissions, but this does not guarantee automatic objectivity. If the criteria are flawed, the system will replicate the error more consistently than a human.

Scoring has therefore become important right now because three trends have converged: a growing volume of applications, increasing language analysis capabilities, and pressure for a faster, more measurable process. This does not mean every recruitment needs AI. It does mean, however, that teams must find a way to maintain evaluation quality in a scenario where manual document screening increasingly becomes the bottleneck.


AI Scoring is not the answer to a lack of recruitment strategy. It is an attempt to bring structure to evaluation where scale, pace and the volume of data make consistent human work difficult.


4. Numerical and contextual assessment - two dimensions of good scoring

The most visible element of scoring is usually the number. A score of 84/100 allows you to spot an application quickly, compare it with others and set the review order. It is particularly convenient when there are dozens or hundreds of profiles in a project. However, the number alone says surprisingly little. It does not explain which criteria were met, which ones carried the most weight, what was missing from the documents, or whether a lower score is due to a genuine mismatch or simply a lack of information.

That is why a good AI Scoring system should have two complementary dimensions. The first is numerical and structuring. It helps set the order of work and quickly flag profiles that require attention based on the adopted criteria. The second is contextual and explanatory. It shows what the system found in the data, how it linked that information to the role, which elements it considered strengths, which ones were weaknesses or unconfirmed, and what still needs checking.

In Recruitify, users get a score from 0 to 100, a candidate description in the context of the project, strengths, weaknesses and suggested interview questions. These elements are not just decoration for the percentage. They are precisely what allows the recruiter to judge whether the score makes sense. If someone gets 88 points but the justification does not confirm a critical must-have, the recruiter should spot the issue. If a profile has 58 points mainly because it lacks information on team scale or salary expectations, the correct response might be a quick screening question, not a rejection.

Contextual evaluation also helps reduce the anchoring effect. A large number displayed on the screen easily becomes the first and most dominant piece of information about a candidate. A human starts looking for arguments to back up the score instead of analysing the profile independently. When they see the sources of the score, unconfirmed areas and questions next to the number, it is much easier to treat the scoring as a working hypothesis rather than a final verdict.

However, contextual results must use cautious language. ‘The candidate managed a team of ten’ can be a fact derived from the CV. ‘The candidate is an excellent leader’ is a conclusion that the document does not prove. ‘No information on budget management was found in the available materials’ means a lack of evidence, not a proven lack of competence. A good tool and a well-trained user distinguish between facts, interpretations and information gaps.

It is possible to have purely numerical scoring, but it will lack transparency. It is also possible to have a purely descriptive evaluation, but with a high volume of applications, it is harder to use for prioritisation. The greatest value comes from combining both layers.


The number organises. The explanation allows you to evaluate. The question allows you to verify.


5. How does AI Scoring work step-by-step?

While different systems may use different models, algorithms and interfaces, the practical AI Scoring process can be presented as a sequence of several linked stages. Crucially, the evaluation does not start with the CV. It starts with defining the role and deciding what information matters for this specific recruitment.

1. The recruiter creates a project and describes the requirements. The system needs a point of reference. A job title alone or the full text of a job ad is usually not enough. Requirements should specify genuine must-haves, nice-to-haves, acceptable alternatives and questions that cannot be resolved solely on the basis of documents.

2. Criteria are checked and approved by a human. AI can help structure the description, but it should not decide on its own that exactly five years of experience, working at a specific company or graduating from a particular university is essential. The recruiter and the Hiring Manager must confirm that the criteria are job-related, proportionate and verifiable.

3. The system gathers available data about the application or candidate. In Recruitify, this can include information from the CV, LinkedIn profile and answers provided in the application form. The scope of data may vary between individuals, so the score must always be read in the context of the material's completeness.

4. Data is structured and mapped to the criteria. The system can recognise not only identical words but also semantic connections. A requirement for managing a sales team can be linked to a description of running a business development department. However, such a conclusion should remain auditable by the recruiter.

5. The AI distinguishes between what is confirmed and what is unknown. A good solution should not turn silence into a negative answer. A lack of information about Kubernetes is not the same as a declared lack of knowledge of Kubernetes. On the other hand, information from a form stating that a candidate can only start work in six months is a concrete signal that contradicts a requirement to start within a month.

6. A numerical score and contextual evaluation are generated. The system can award a score from 0 to 100, prepare a profile description, point out strengths and weaknesses, and suggest questions for further discussion. The result should be a trail of analysis, not an automated decision to advance or reject.

7. The recruiter checks the justification and source material. If a score seems too high or too low, the cause must be determined. The issue could lie in the requirements, data extraction from the CV, model interpretation or simply the incompleteness of the profile. A human must have the ability to make a different decision than suggested by the scoring order.

8. The result leads to action. The recruiter can start with high-match profiles, plan a call, ask a follow-up question or go back to the Hiring Manager to refine the brief. Scoring creates no value if it ends with a colourful number that nobody acts on.

9. The organisation monitors quality and impact. The team should track where scoring works well, when it generates inaccurate conclusions, how often recruiters disagree with the assessment, and whether lower scores are leading to valuable people being automatically overlooked. Implementation does not end when you toggle the feature on. That is when the responsible management of its impact begins.

In practice, this process may take a few seconds on the system's side, but its quality is the result of prior human work. The better the brief, the more adequate the data and the more conscious the user, the greater the utility of the result.


[SCREENSHOT: example of AI Scoring flow in Recruitify - project requirements, score 0-100, summary and questions]


6. What data can AI Scoring use?

The quality of the scoring depends not only on the AI model, but primarily on the quality, relevance and scope of the input data. In Recruitify, the evaluation can take into account information from the CV, LinkedIn profile and answers provided in the application form. These sources are not equivalent and should not be mechanically lumped together. Each shows a different piece of the profile and has its own limitations.

The CV usually provides the most structured professional history: job titles, companies, dates, responsibilities, education, skills and achievements. However, it is a selective document. The candidate decides what to fit on one or two pages, may shorten older experience, omit a project deemed less relevant, or use role titles specific to their previous organisation. A CV can also be outdated or generic, prepared without knowledge of a specific project's criteria.

A LinkedIn profile can contain a newer employment history, broader project descriptions, recommendations or skills omitted from the CV. However, it should not automatically be treated as more reliable. It is also created by the user, can be incomplete, outdated or written in a marketing style. Its value lies primarily in adding context and spotting discrepancies, not in serving as external proof of truth.

The application form allows you to collect information that is often missing from both the CV and LinkedIn. It can cover availability, salary expectations, preferred working model, right to work in a given country, willingness to relocate, required licences or experience in a very specific area. A candidate usually does not write in their CV whether they can start work in August, whether they accept two days a week in the office, or whether they have worked with a specific version of a system before. However, they can answer such questions in the form.

It is precisely the combination of sources that allows you to build a fuller picture. Imagine a candidate for a Finance Manager role. Her CV describes team leadership and month-end closing, LinkedIn shows a more recent promotion, and in the form, she confirms knowledge of a specific ERP system and availability in the required timeframe. Scoring based solely on the CV would miss the last two elements. At the same time, the system should not extrapolate anything that no source confirms.

Data should be interpreted according to three basic states: criterion confirmed, criterion probably unmet, and lack of sufficient information. This distinction is more important than the technical number of sources. Two outdated profiles do not create better proof than one fresh and specific answer.

It is also worth establishing a hierarchy of relevance and rules for resolving contradictions. What should the system do if a CV indicates employment ended in May, but LinkedIn shows ongoing employment? Should a form response submitted yesterday take precedence over a CV prepared six months ago? A good process does not hide such discrepancies under a uniform score. It signals them to the recruiter and turns them into questions for verification.

Combining multiple sources does not mean you should analyse everything that is technically available. The data minimisation principle requires using only the information necessary for a specific purpose. Age, photos, family situation, health status, origin or other sensitive areas should not affect scoring just because they can be found in a document or profile. More data does not always mean a better evaluation. Sometimes it just means more noise and a higher risk of biased conclusions.


The best scoring does not analyse the largest amount of data. It analyses data that is adequate, up-to-date and needed to evaluate the requirements of a specific project.


7. Why are project requirements more important than the AI model itself?

The most advanced model will not fix a poorly defined recruitment. If project requirements are vague, contradictory, unrealistic or contain biases, AI will only apply them faster and more consistently. It is precisely this consistency, often presented as an advantage of automation, that can become a problem: a human reading a few CVs might eventually notice that a criterion makes no sense. The system will keep repeating it for every person until someone changes it.

A good requirement should be specific, related to actual work and verifiable in the available data or during a later stage of the process. ‘A minimum of three years of experience in managing ERP implementation projects’ is more useful than ‘extensive project experience’. ‘Independently acquiring B2B clients with a contract value over €100k’ gives more context than ‘strong sales skills’. ‘Knowledge of English at a level allowing for negotiation’ describes a business need better than just ‘C1’, if the organisation does not intend to check the level with a formal test.

Requirements should not simply copy the entire job advert. An advert also serves a communication and employer branding function. It contains a description of the company, benefits, scope of responsibility and engaging language. Scoring, on the other hand, needs clear evaluation logic: what is a mandatory condition, what increases fit, what experience is equivalent, what can be learned after hiring, and what is absolutely non-negotiable.

The weight of criteria is also critical. Five years of experience is not always twice as good as two and a half years. A lack of one technology might be easy to catch up on, while a lack of a legally required licence actually prevents starting work. Scoring should reflect the business significance of criteria, not the number of words dedicated to them in the ad. If the ‘nice-to-have’ list is longer than the mandatory requirements, the system should not automatically allow the sum of minor additions to overshadow the lack of a core competence.

Before launching scoring, it is worth conducting a quick calibration with the Hiring Manager. Instead of only asking ‘who are we looking for?’, it is better to establish: what is this person supposed to deliver in the first six months, what experience genuinely increases the chance of success, which gaps are acceptable, and which ones make hiring impossible. Only on this basis can you create criteria that hold value for both AI and people.

8. Must-have, nice-to-have and disqualifying criteria

One of the most important stages of preparing scoring is separating three categories of requirements. Must-haves are conditions absolutely necessary to perform the job or start it within the set timeframe. Nice-to-haves increase fit, but their absence should not close the door to the process. Disqualifying criteria, however, are binary conditions where failure to meet them can justify automatically or semi-automatically directing an application to a separate pipeline - provided they are legal, proportionate and unambiguous.

In practice, organisations overuse the must-have category. A Hiring Manager might deem experience at a specific company, exactly five years of work, knowledge of an internal tool or graduation from a preferred university as essential, even though none of these are conditions for successfully performing the role. If AI Scoring receives such a list uncritically, it will reinforce a narrow 'ideal candidate' profile and lower the scores of people with diverse but highly valuable career paths.

Disqualifying criteria should be applied with extreme caution. They make sense where the answer is binary and directly related to the ability to do the work: a legally required licence, the right to work in a given location, availability during required hours, or willingness to work in a model the organisation cannot change. Do not turn elements requiring interpretation, such as 'good cultural fit', 'sufficient experience' or 'fitting personality', into killer questions.

AI Scoring should also not hide disqualifying rules within an opaque number. If a candidate received a low score due to the lack of a legally required certificate, the recruiter should see this clearly. If a form response indicates that a candidate cannot work the required hours, the system should distinguish this information from a lack of data. Transparency of the evaluation logic matters both for the quality of decisions and for the ability to explain the process to the candidate.

The best practice is to assign not only a category and weight to each criterion, but also a method of verification. Can the information come from the CV? Is a response in the form needed? Does it need to be confirmed during an interview? Does it require a document? This way, scoring does not pretend it can resolve everything based on a single source.

9. AI Scoring vs. keyword search

One of the most common simplifications in conversations about ATS is the belief that the system 'rejects CVs if they don't contain the right words'. In some older solutions, exact phrases, simple filters and manually built queries did play a large role. Keyword searching is still useful, but it is not the same as AI Scoring and should not be presented as its full mechanism of action.

Classic search checks for the presence or absence of specific expressions. If a recruiter types 'Salesforce', the system can find people who placed that name in their profile. The problem begins when the experience is described in a different language, the candidate uses a broader category name, or a specific competence is implied by context. A person who 'implemented and administered CRM solutions in a Salesforce environment' should be found easily. A harder case is a candidate describing 'managing a cloud-based sales automation platform' without mentioning the brand. A semantic system can recognise a potential connection, but it should treat it as a conclusion requiring confirmation, not as a certain fact.

AI Scoring based on a language model can match the meaning of requirements with the content of documents. This allows it to spot equivalent experience, role titles specific to an organisation, and skills described in a different order than in the job ad. For example, a requirement for 'managing a customer success team in a SaaS model' can be linked to a description of 'responsibility for an eight-person subscription customer care team'. Such reasoning is useful but not infallible. The more semantically distant the connection, the greater the need to show the recruiter exactly which part of the document the conclusion was based on.

We should not, therefore, pit keywords and AI against each other in an absolute way. Good tools can combine precise rules with contextual analysis. For a 'CISA' certification, the exact phrase matters. For experience in leading an organisational transformation, analysing the meaning is more important than identical words. The choice of mechanism should depend on the nature of the criterion.

It is also worth being cautious about promises that AI 'understands CVs like a human'. This is attractive from a marketing standpoint but is overreaching. A model can analyse text in a way that resembles human language processing, but it does not possess professional experience, intentions or full context. It can correctly link synonyms while completely misinterpreting the scope of responsibility. That is why semantisation increases capabilities but does not remove the need for human verification.


10. What does a score from 0 to 100 mean?

A score from 0 to 100 is primarily an auxiliary scale. It allows you to quickly see how the system evaluated the fit of available information against the project criteria. It is not a probability of hire, a prediction of job success, or a scientifically calculated 'candidate quality'. A score of 86 does not mean the candidate has an 86 per cent chance of succeeding in the role. Nor does it mean they are exactly 14 points better than someone with a score of 72.

The meaning of the number depends on the design of the specific system, how criteria are weighted, and the completeness of the data. In one project, scores might naturally cluster between 60 and 85, in another between 20 and 55 if the requirements are highly specialised. That is why thresholds like 'we reject everyone below 70' are dangerous if they have not been tested for a specific process. Even then, a threshold should not automatically replace human review, especially when the score affects access to employment.

The number can be most useful as a prioritisation tool. A recruiter can start their analysis with highly-rated individuals, but they should also review a sample of profiles from lower brackets. This kind of audit allows you to check whether the system is overlooking non-standard candidates, lowering scores due to a lack of information, or rewarding people who simply wrote a better CV. In practice, it is worth establishing a rule to regularly review the 'tail of the ranking', especially for new types of roles.

A good interface should not display the number in isolation from the justification. If the score is shown in a large font and the explanation is hidden under another click, users will make decisions based on the most visible element. Screen design influences behaviour just as much as the model itself. In a system supporting responsible decisions, the score, strengths, weaker areas and missing data should all be available together.

It is also worth remembering score stability. Generative models can, in some configurations, produce slightly different answers for the same input. The provider should limit this variance, test repeatability and clearly describe what might trigger a recalculation. The user, in turn, should know that changing project requirements or candidate data can legitimately change the score. The result is not a permanently written fact, but the outcome of a specific comparison at a specific time.


[SCREENSHOT: list of candidates or applications with AI Scoring from 0 to 100]


11. Strengths, weaknesses and interview questions

The greatest value of AI Scoring often lies not in the number, but in the contextual layer. A recruiter needs to know more than just that a profile was rated high or low. They need to know why, which elements matter for the project, and what they should do next. That is why a useful scoring tool can prepare a candidate description, highlight strengths and weaknesses, and suggest questions for the interview.

A candidate description should be a brief synthesis, not a summary of the entire CV. Its role is to connect the experience with the specific project. Instead of a generic 'experienced financial manager', it is better to point out that the person managed a team of similar size, was responsible for reporting to an international group and participated in an ERP system implementation, which is one of the challenges of the new role. This information helps the recruiter grasp the profile faster, but remains verifiable in the sources.

Strengths should relate to the criteria and stem from concrete data. It is not about generating compliments. 'Worked as a Sales Manager' is not a strength yet. 'Was responsible for independently acquiring enterprise clients in the DACH market for three years and achieved an annual target over €2m' is information that can be mapped to the project requirements.

Weaknesses require even more caution. In product language, we might call them weaknesses, but in practice, the system should separate at least three types of information. The first is a genuine contradiction with a requirement, such as lacking a legally required licence or availability being later than the project allows. The second is a probable mismatch, when the materials indicate experience significantly different from what is sought. The third is a lack of data, i.e., a situation where the documents simply do not allow a criterion to be resolved. These categories must not be treated the same way.

Suggested questions turn scoring from an organising mechanism into an interview preparation tool. If the CV lacks information on team scale, the system might suggest: 'How many people were in your team and what personnel decisions were you responsible for?'. If a candidate lists a technology without context: 'In which project did you use it and what was your level of independence?'. If LinkedIn and the CV differ on dates: 'Which information is current and what is the reason for the discrepancy?'. A good question is not meant to confirm a preconceived score. It is meant to supply the information that is still missing.

It is also worth avoiding questions that are too general, leading or about areas unrelated to the job. AI can generate a proposal, but the recruiter must check it and tailor it to the stage of the process. Asking about motivation during the first short screening might make sense. Assessing personality based on a stereotypical assumption does not.

The contextual layer has another benefit: it facilitates communication with the Hiring Manager or client. Instead of passing on just a number or a brief 'the profile fits', the recruiter can show arguments in favour, risk areas and a verification plan. This improves shortlists and shifts the discussion from gut feeling to concrete evidence.


[SCREENSHOT: candidate description, strengths, weaknesses and suggested questions in Recruitify]


12. Missing information is not the same as a lack of competence

This is one of the most critical principles of interpreting AI Scoring. Application documents are an incomplete description of a human being. A candidate might not list a skill because they took it for granted, wanted to limit CV length, tailored the document for a different role, or did not know the exact project requirements. A LinkedIn profile might be updated once every few years. An application form might contain a question that is too broad. Therefore, a lack of information is primarily a lack of evidence in the analysed material, not proof of a lack of competence.

The system should distinguish between at least three states:

Criterion confirmed - available data clearly indicates it is met.

Lack of sufficient information - the material does not allow confirming or ruling out that the criterion is met.

Criterion probably unmet or unmet - the data contains information that contradicts the requirement, e.g., the candidate declares availability in six months, and the project requires a start within four weeks.

Lumping the second and third categories together is a frequent source of unfair decisions. If a CV makes no mention of a specific technology, the system can write 'no confirmation in available data'. It should not, without additional basis, state 'the candidate does not know the technology'. If a candidate responds in a form that they have never used it, the situation is different, but even then, the weight of this information depends on whether the technology is a genuine must-have or a skill that can be quickly learned.

The style of the document also matters. A detailed CV provides more text and more potential evidence than the sparse profile of an expert who assumes company names and job titles are enough. As a result, the system may partially reward self-presentation quality. This is not a problem unique to AI, as recruiters also fall for well-written documents, but automation can scale this effect and give it a veneer of mathematical precision.

One way to mitigate this issue is to use the application form to gather critical information in a standardised way. If experience in a given area really matters, it is better to ask a specific question than to hope the candidate happened to describe it in their CV. The second way is generating interview questions. The third is regularly reviewing a portion of low-scoring profiles to check if the system mistook sparse documentation for a lack of qualifications.

This principle should also be part of user training. A recruiter must know that 'not found' does not mean 'does not possess', and 'the system suggests' does not mean 'the fact is confirmed'. Without this awareness, even well-designed explanations can be applied too categorically.


Missing information should lead primarily to a question. Not to an automated rejection.


13. AI Scoring vs. the traditional candidate scorecard

A traditional scorecard relies on pre-established criteria and a scale. A recruiter or Hiring Manager assigns points for experience, competencies, motivation or task results after an interview. A well-used scorecard increases process consistency because it forces evaluators to refer to shared criteria rather than a general impression.

AI Scoring should not replace this scorecard, but rather complement it at an earlier stage. It analyses the data available before the interview and helps prepare an initial assessment. The traditional scorecard can then cover information obtained during the interview, task, test or reference checks. Combining both tools makes sense when the criteria are consistent and it is clear which source can confirm which element.

The difference also relates to accountability. In a scorecard, a human directly assigns the rating. In AI Scoring, part of the interpretation is performed by the system, so the user must be able to trace the logic, check the basis and challenge the score. It is not enough to say that 'the algorithm calculated it'. The organisation remains responsible for how the evaluation is used in the process.

AI can improve the consistency of the initial analysis, but it does not remove human bias. Criteria are created by humans, data is created by humans, and results are interpreted by humans. What is more, the AI score can become a new source of bias - an automatic anchor that evaluators will adjust their own impressions to fit. Therefore, it is sometimes worth hiding the score from some individuals until they make their independent assessment, or at least requiring a brief justification for decisions independent of the score.

In practice, the best model might look like this: AI Scoring structures incoming applications and flags areas to check; the recruiter reviews the profile; the interview is conducted using standardised questions; the scorecard gathers evidence from the interview; the Hiring Manager makes a decision based on the entire set of materials. AI is then one element of the process, not an independent judge.

14. AI Scoring vs. AI Matching - a key distinction

AI Scoring and AI Matching are often lumped into the same category because both features match candidate profiles to project requirements. However, their goals are different, and in articles, product documentation and client communication, they should be consistently separated.

AI Scoring evaluates an application or candidate we are already reviewing within the context of a project. The starting point is a specific profile. The system prepares a score and explanation: to what extent the available information matches the criteria, what the strong and weaker areas are, and what needs verifying. It can apply to a person who applied themselves, or a candidate added to the project by a recruiter.

AI Matching works from the other side. The starting point is the project and the need to find potentially fitting individuals. The feature searches the candidate database, compares profiles with requirements, and suggests to the recruiter who they might want to consider. Its main task is discovery and internal sourcing: finding people the recruiter has not yet added to the project or might have forgotten about.

The simplest distinction is:

AI Scoring: ‘How does this candidate or application stack up against the project requirements?’

AI Matching: ‘Which candidates in the database might fit this project and should be suggested to the recruiter?’

In Recruitify, AI Scoring is an evaluative feature, whereas AI Matching is being developed as a separate mechanism for searching and suggesting candidates from the database. Scoring should not be presented as 'the next step after matching', as the process can run without matching. A candidate can apply via an ad, be added manually, or come from a referral, and then be evaluated by AI Scoring.

Both features can complement each other in the future. Matching finds potential profiles, the recruiter adds selected individuals to the project, and scoring prepares a more detailed assessment against the criteria. However, separate quality standards are still required. Matching should be measured by suggestion relevance and the ability to find valuable profiles. Scoring should be evaluated based on consistency, explainability and utility in analysing a specific person.

15. AI Scoring vs. CV parser, AI summary and killer questions

In a modern ATS, several features might work on similar data but perform completely different tasks. Confusing them leads to unrealistic expectations and makes risk assessment difficult. It is worth distinguishing AI Scoring from a CV parser, candidate summary, search, killer questions, workflow automation and the simple use of a public chatbot.

A CV parser reads a document and converts it into structured data: contact details, employment history, education, skills or languages. Its core task is extraction, not evaluation. A parser can make an error, such as assigning a date to the wrong company, and this error can later affect scoring. That is why extraction quality is one of the pillars of evaluation quality.

An AI summary creates a profile abstract. It can describe experience, key skills and career path, but it does not have to match them to project criteria or award a score. A summary primarily answers 'what is known about this candidate?', whereas scoring asks 'how does this information relate to the requirements of this recruitment?'. In practice, both elements may be presented together, but they are not the same.

Killer questions or disqualifying questions are rules based on specific answers. If a job requires a valid driving licence and the candidate answers 'no', the system can apply a pre-established path. This is not necessarily AI. It is rule-based automation whose logic was defined by a human. It has advantages in unambiguous cases, but can be more ruthless than scoring if a criterion was poorly chosen.

Workflow automation performs actions when a condition is met: sends a message, adds a label, changes a stage, creates a task or notifies the project owner. It can use the scoring result as one of the conditions, but such a connection requires extreme caution. Automatically rejecting everyone below a certain score turns a supporting feature into an actual decision-making mechanism and significantly increases legal and ethical risks.

A separate case is copying a CV and job description into a public chatbot with the prompt 'evaluate this candidate'. Technically, you can get a percentage, summary and questions this way, but it is not equivalent to secure AI Scoring embedded in an ATS. The recruiter may not know where the data goes, how long it is stored, whether it is used to train the service, who has access to the chat history, and whether the organisation has an appropriate agreement with the provider. It is also harder to maintain uniform criteria, log requirement versions, control permissions and document how the result influenced the process.

A feature available directly within the recruitment system should operate within an established environment of data processing, access control, retention, logs and provider agreements. The mere fact of integration with an ATS does not guarantee compliance or security, but it gives the organisation the ability to manage the process in a way that spontaneously copying data into a random tool usually fails to provide.

The most important question is therefore not 'can AI generate an evaluation?'. It can. The question is: is the evaluation generated based on approved criteria, on adequate data, in a controlled environment, and with the ability for a human to review it?

16. AI Scoring in practice - three assessment examples

Theory becomes clearer when we see how the same mechanism can work across different recruitment projects. The following examples do not present a ready-made numerical formula. Rather, they show how requirements, data and contextual interpretation should work together.

Example 1: CFO in a high-growth company

Let’s assume a company is looking for a CFO to prepare the organisation for rapid growth, structure controlling, manage funding and work with investors. A poorly described requirement might read: ‘minimum ten years of experience, prestigious company on the CV, strong leadership and strategic thinking’. Such a brief is difficult to score fairly. ‘Prestige’ is subjective, ‘strategic thinking’ does not stem directly from the document, and the number of years might not correlate with the challenge.

Better criteria relate to results and the working environment: responsibility for full P&L, experience in building controlling structures, involvement in securing funding, managing a finance team, reporting to the board or investors, and working in an organisation of a similar scale of growth. The CV can confirm some of these elements. LinkedIn can show a more recent role. The application form can collect availability, expectations and the answer to a question about the largest funding round managed.

The candidate might get a high score because they ran finance in a company growing from 50 to 300 people, implemented controlling and participated in talks with a fund. A weakness or risk area might be the lack of confirmation of independent responsibility for securing funding. Rather than assuming this competence is missing, the system can suggest a question: ‘What was your personal scope of responsibility in the funding process and which elements did you manage independently?’.

Example 2: Java Developer in a migration project

In the second project, a company is seeking a Java Developer to migrate a legacy application to a new architecture. Must-haves might include commercial experience with Java, working with Spring Boot, relational databases and distributed systems. Nice-to-haves are Kubernetes, public cloud and experience in gradually modernising a monolith.

The candidate lists Java and Spring, but does not use the exact phrase 'microservices' in the CV. They do, however, describe breaking down a monolithic application into independently deployable modules communicating via API. An analysis based solely on keywords might miss this important signal. Semantic analysis can spot the connection, but it should point out the fragment on which the conclusion was based.

There is no mention of Kubernetes in the CV. This should not automatically mean a weakness in the sense of a lack of competence. If Kubernetes is a nice-to-have, the system can mark a lack of confirmation and suggest a question: ‘Have you deployed applications in a containerised environment and what orchestration tools did you use?’. If the candidate responds in the form that they worked with Kubernetes for two years, the information can complete the profile. If they answer 'no', scoring should factor this in as a genuine gap, but in line with the criterion's weight, not as an automatic rejection.

Example 3: B2B salesperson for a new market

The third project concerns a person to develop B2B sales in the German market. The title 'Sales Manager' alone says very little. For the role, independent acquisition of new clients, sales cycle length, contract value, working with decision-makers, knowledge of the DACH market and German language skills allowing for negotiations could all be significant.

Candidate A has impressive sales results, but most of their experience comes from managing existing clients in the Polish market. Candidate B has lower revenues but opened the German market independently, built a pipeline from scratch and led enterprise negotiations. A simple scoring system based on the sum of sales results might reward Candidate A. Scoring anchored in the genuine requirements of the project should notice that the second profile is a better fit for the specific task.

The application form can additionally collect information on willingness to travel, salary expectations and actual language level. The system can highlight market building experience as a strength, and the lack of information on average contract value as an area for verification. An interview question could be: ‘Tell us about the largest client you acquired independently in the DACH market - what did the process look like from first contact to signing the contract?’.

These three examples show why there is no single universal scoring for a 'good candidate'. A CFO, a Java Developer and a salesperson are evaluated against entirely different criteria. Even two candidates with similar job titles might score differently depending on the project's goal. AI can help navigate requirements consistently, but it does not relieve the organisation of the need to answer the hardest question: what do we actually need from a person in this role?

17. Benefits for the recruiter and the recruitment team

The most obvious benefit is time savings, but the value of AI Scoring should not be reduced to reading CVs faster. A well-used feature changes how the initial analysis is organised. It helps spot applications meeting multiple criteria faster, flag areas requiring checks, and prepare a more structured interview. The recruiter still reviews profiles, but they do not start with a blank page every time.

Savings also do not mean a human stops reading documents altogether. It is about a different distribution of attention. Instead of dedicating the same amount of time to every part of every CV, the recruiter can immediately see which requirements have been confirmed, where there is a contradiction and what is missing. Thanks to this, they spend more time on interpretation, contact and verification, and less on manually finding scattered information.

The second benefit is greater consistency. With hundreds of applications, people naturally differ in their attention levels and interpretation of criteria. The system can apply the same analytical framework to every profile. This does not automatically mean greater fairness, because flawed rules can also be applied consistently. However, if criteria are well-prepared and results are monitored, consistency reduces the randomness of the initial review and makes comparing decisions within the team easier.

The third benefit is better justification. The recruiter can quickly prepare arguments for the Hiring Manager or client, instead of sending just the CV and a generic 'worth a talk'. Strengths, weaknesses, information gaps and questions for verification create a shared language of evaluation. This does not mean the generated description should be copied without checks. Its value lies in preparing a draft analysis that a human reviews and refines.

The fourth benefit relates to interview preparation. The recruiter can walk into a meeting with a list of specific areas to dig into, rather than asking everyone the same set of general questions. This allows better use of the candidate's time and faster separation of actual experience from an attractive description.

The fifth benefit is linked to teamwork and onboarding new people. If every consultant uses their own informal way of reading CVs, it is hard to hand over projects, analyse quality and train less experienced recruiters. AI Scoring can introduce a common point. The team can discuss why they disagree with a score, which requirement was poorly written and how to improve the brief. The difference between human and system evaluation itself becomes learning material.

The sixth benefit is faster response times. In an agency, a few hours can decide which firm presents a valuable candidate first. In in-house recruitment, quick contact affects candidate experience and reduces the risk of losing someone to another process. Scoring can shorten the time from application to the first meaningful analysis, but only when embedded in the workflow and leading to action. If it remains just another number in the system, it brings no real value.

18. Benefits for agencies, in-house HR and Hiring Managers

In a recruitment agency, AI Scoring can primarily support delivery. A recruiter often runs several projects simultaneously, and clients' requirements can be incomplete and change during the process. Contextual evaluation helps grasp a profile faster, prepare the screening and build arguments for the shortlist. It can also facilitate project handovers between consultants. A new person does not have to reconstruct all the logic solely from notes if they can see the approved criteria and how candidates map to them.

The speed of working with your own database is also key. Scoring is not AI Matching and does not automatically search for the best people for a project on its own, but it can evaluate candidates whom the recruiter added from previous processes, referrals or sourcing. Thanks to this, the database stops being purely a CV archive. The condition is data relevance and the legal basis for its further use.

In in-house recruitment, the greatest value often appears with high volumes, but is not limited to them. Even with dozens of applications, the tool can help standardise the initial evaluation across locations and recruiters. It can also reveal a briefing problem. If most candidates receive low scores due to a single requirement, it is worth checking whether the market simply does not supply such profiles, or if the criterion is too narrow or poorly communicated.

The Hiring Manager gets more structured material. Instead of reading every CV without context, they can receive a summary, pros and cons, and areas to check. However, they should not see only the AI score. A large number without access to the document and justification can lead to over-reliance. The best model ensures easy access to both the contextual analysis and the full profile, along with the recruiter's decisions.

AI Scoring can also improve the quality of the conversation between HR and the business. Instead of discussions like 'I like this candidate' versus 'this candidate doesn't fit', you have a list of requirements, evidence, gaps and questions. This does not remove subjectivity, but makes it more visible. If a Hiring Manager rejects a high-match person, it is worth asking which criterion was not written down before. If they accept a low-scoring profile, perhaps the system misinterpreted the data or the requirements needed adjustment.

For the organisation, the benefit is better process analysis. You can check whether high scores translate into advancing through subsequent stages, where recruiters most often disagree with the evaluation, which requirements are regularly unconfirmed, and whether lower-bracket candidates are being unfairly overlooked. However, you must be careful not to treat historical decisions as objective truth. If the past process was biased, blindly learning from its results can entrench the problem.

19. Benefits and risks from the candidate’s perspective

From the candidate’s perspective, AI Scoring can bring real benefits. Every application can be subjected to the same analytical framework, even if it arrived as the two-hundredth. The system can notice experience described differently than in the ad, combine form and CV details, and suggest a question to the recruiter that gives the candidate a chance to explain a missing element. Faster pre-selection can also mean a quicker response and shorter waiting times.

However, benefits are not automatic. A candidate might be unfairly downgraded by an incomplete CV, an unconventional career path, a career change, a gap in employment, or language that differs from the ad's standards. They may not know their data is being analysed by AI, fail to understand the role of the score in the decision, and have no easy way to challenge an error. If an organisation automatically rejects people below a threshold, a single misinterpretation can have a significant impact.

Candidate experience therefore depends on transparency and usage. The organisation should clearly state that AI tools are used in the process, describe their role and highlight whether the decision is made by a human. Candidates should have the opportunity to correct inaccurate data and obtain information on how their data is used in compliance with applicable law. In processes with a higher impact, it is worth providing an appeal channel or human review option.

We should also not shift responsibility entirely onto candidates by stating 'if they didn't write it, it's their own fault'. A CV is a communication tool, but it is the organisation that designs the process and decides what data is needed. If a criterion is key, it should be communicated and gathered in a proportionate way. Scoring can support a fairer evaluation only when the entire process is designed fairly.

20. Where does AI Scoring give the most value?

AI Scoring does not deliver the same value in every process. The most obvious application is recruitment with a high volume of applications, where reading every document manually with equal attention is difficult. This includes popular junior roles, remote positions, international recruitment and processes promoted on portals that enable quick applying. In these scenarios, scoring can help establish the review order and flag profiles requiring urgent contact.

The second application is projects with clearly defined, verifiable requirements. If a role requires specific technologies, licences, experience in a specific environment, sales model or operational scale, AI can compare these elements with documents. The more vague and subjective the criteria, the lower the value of automatic scoring. Scoring is much better at answering 'did the candidate run ERP system implementations in organisations over 500 people?' than 'do they have leadership charisma?'.

The third area is reusing the candidate database. A recruiter can add people known from previous processes to a new project and evaluate their profiles against new requirements. However, you must check data relevance and the legal basis for further processing. Scoring on an outdated CV can create a seemingly precise but low-value result.

The fourth application is agency recruitment, where the speed of preparing a shortlist and the quality of candidate presentation matter. Contextual evaluation can help the consultant prepare the screening, spot risks and build arguments for the client. However, it should not be copied directly into the candidate profile without checks. AI-generated phrasing can be too categorical, imprecise or inappropriate for an external recipient.

Scoring may have less value in executive search processes, where candidate numbers are small and key information comes from interviews, market reputation, references and understanding the complex context of the organisation. Even there, it can support structuring data, but it is hard to expect a CV score to capture the ability to lead transformation, build board trust or operate in a specific ownership culture.

Scoring should not be used just because the feature is available. Before implementing, it is worth answering three questions: what specific problem are we solving, how will we check score quality, and what should the user do after receiving it. If there is no clear answer, another number in the ATS can increase complexity instead of productivity.

21. How to implement AI Scoring in your recruitment process?

A good implementation starts with the process, not a button in the system. The first step should be selecting a limited group of roles for which requirements are relatively clear and a sufficient number of applications are available for testing. Running a pilot across all projects simultaneously makes it hard to identify what works and what needs refinement.

Next, you need to define the rules for creating requirements. Who prepares the criteria, who approves them, how are mandatory and nice-to-have conditions distinguished, how often can they be changed, and what happens to scores after a change? It is worth preparing a simple brief template. It should require describing the goal of the role, key results, essential competencies, acceptable alternatives and areas to verify during the interview.

The next step is testing on known profiles. The team can choose sample CVs of candidates they evaluated previously, run the scoring, and compare the result with recruiters' independent assessments. The goal is not for the system to perfectly replicate every historical decision. History can contain errors. The goal is to check whether justifications are logical, whether the system recognises key experiences, how it treats missing data, and whether undesirable patterns emerge.

User training is essential. It should cover not only feature usage, but also score interpretation, automation bias, the difference between missing information and lack of competence, data protection rules and escalation scenarios. A recruiter should know when they can use scoring for prioritisation and when they must not rely on it without additional analysis.

It is worth establishing a human review practice. For instance, you could require that before rejecting candidates from a specific bracket, the recruiter must review the document and note their own reason for the decision. You can also randomly audit a portion of low-scoring applications. It is important that oversight is real, not formal. Clicking 'approve' without reading the profile is not meaningful human involvement.

Finally, you need to define monitoring. Monthly or quarterly, the team can analyse score distribution, differences between roles, score alignment with recruiter decisions, the number of corrections, profiles bypassed by scoring, and candidate feedback. The implementation should have an owner responsible for both efficiency and risk.

22. Most common mistakes when using AI Scoring

The first mistake is running scoring on a weak brief. A generic job title and a copied job ad do not form an evaluation logic. If the Hiring Manager cannot specify which requirements are truly important, AI will not establish this reliably either. It can suggest criteria, but a human must verify them.

The second mistake is equating a high score with a hiring recommendation. Scoring relates to data matching against criteria at a specific stage. It does not cover everything that drives job success. A candidate might have perfectly described experience but low motivation, fail to confirm declared skills, or reject the offer terms.

The third mistake is automatic rejection based on a threshold. This setup is tempting because it maximises time savings, but it can lead to systematically overlooking people with incomplete CVs and increases the impact of a single error. It can also turn a supporting tool into a system actually making decisions, which has legal consequences.

The fourth mistake is failing to distinguish between data and conclusions. If an AI summary sounds convincing, the user might forget it is an interpretation. Every significant conclusion should be linkable to a document fragment or marked as requiring confirmation.

The fifth mistake is analysing too broad a scope of data. More data is not always better. Photos, age, address, marital status, health information or social media activity can be unnecessary and risky. The organisation should consciously define which sources and fields can affect scoring.

The sixth mistake is a lack of testing after changing a model, prompt or logic. An AI feature might behave differently after an update, even if the interface looks the same. The provider should manage changes, and the client should know when scores might shift.

The seventh mistake is assuming a human will automatically correct system issues. Research on human-AI collaboration shows that people can over-rely on recommendations and align their own decisions with the score. Oversight requires competence, time and the right to disagree with the system, not just a recruiter's name in the process.

23. Limitations of AI Scoring - what does the system not know?

AI Scoring sees the information it has been given, and nothing more. It does not know competencies omitted from the CV, does not know if an achievement was described accurately, does not independently confirm declaration truthfulness, and does not observe the candidate at work. It can draw conclusions from text, but should not replace skills tests, behavioural interviews, references or evaluations performed in a real context.

The system may struggle with unconventional career paths. A person transitioning from the military to business, from academia to product, from entrepreneurship to corporate, or from another industry may possess transferable skills that they do not describe in the target role's language. A semantic model can spot some of these, but requirements based on historical job titles might still downgrade the score.

Soft skills are also difficult. A CV might state 'excellent communication skills', but this is not reliable evidence. The system should not award a high score for the declaration alone. It can, however, point to experiences where communication was likely important and suggest interview questions. Assessing actual behaviour, however, requires other methods.

AI also does not know the full context of the organisation. Two companies might use the same job title for completely different scopes. Managing a team of ten in a stable department is different from building a team from scratch in a startup. A budget of one million in one industry might have a different meaning than in another. Therefore, requirements must describe context, not just labels.

Data and document errors are another limitation. A parser can misread a table, the system can mix up dates, and a LinkedIn profile can be outdated. The more precise the numerical score, the easier it is to forget about input uncertainty. The 'garbage in, garbage out' rule remains relevant in the era of language models.

Finally, AI does not bear responsibility for the decision. It can prepare an assessment, but it does not talk to the candidate, does not know all the circumstances, and is not accountable to the organisation, client or regulator. Responsibility remains with the people and entities that design and use the system.

24. Bias, discrimination and the illusion of objectivity

One of the most tempting promises of AI in recruitment is objectivity. A machine is not tired, does not have a favourite university, and does not evaluate a photo like a human - provided the photo and other unnecessary data do not influence the model. However, this does not mean the score is neutral. The system can replicate biases present in the data, criteria, labelling method or the process design itself.

Bias can enter scoring at multiple levels. Criteria might reward paths more accessible to specific groups, such as an unjustified requirement for continuous employment. Historical data might reflect the organisation's past preferences. Company and university names can act as proxy indicators of social status. CV language can differ by culture and gender. Missing data might more often affect people with less conventional career histories.

Language models can also draw conclusions based on indirect traits. Research on CV evaluation by LLMs shows that changing elements suggesting gender or origin can affect the score, though the scale and direction of bias vary between models. The FAIRE benchmark of 2025 demonstrated the presence of some level of bias in tested models during direct CV scoring and ranking [7]. This does not mean every recruitment system will behave identically, but it confirms the need to test specific configurations rather than relying on general model manufacturer claims.

The illusion of objectivity is particularly dangerous because numbers look scientific. A score of 78.4 can give the impression of a measurement similar to temperature, though it is the result of chosen criteria and text interpretation. Precision in display is not proof of accuracy. The organisation should ask: what does the score measure, how was it tested, for which groups does it work less well, and what are the consequences of an error.

Mitigating bias is not just about removing names and photos. Anonymisation helps, but other elements can still reveal or indirectly indicate protected traits. What is needed is score testing, analysing differences between groups where legal and methodologically justified, reviewing criteria, error reporting mechanisms and the ability to withdraw a feature if risk cannot be mitigated.

It is also worth remembering that humans are not a neutral benchmark. AI can, in some situations, limit recruiter randomness and unconscious preferences. The goal should not be to prove that 'AI is more objective than a human' or vice versa. The goal is to design a process where both parties' errors are visible, measured and correctable.

25. Human-in-the-loop - what does real human oversight mean?

The phrase human-in-the-loop appears in almost every description of responsible AI. Simply placing a human in the process is not enough. If a recruiter sees a score, automatically accepts the recommendation and has no time for analysis, their involvement is formal. Real oversight means the ability to understand the system's capabilities and limitations, check the basis of a score, challenge it, and halt or change actions when an issue arises.

The AI Act, in relation to high-risk systems, requires designing solutions so they can be effectively overseen by natural persons. It also highlights the risk of automation bias, i.e., automatic or excessive reliance on system output [4]. The person exercising oversight should have appropriate competence, training, authority and support. In a recruitment context, this means a junior recruiter cannot be the only 'safety layer' if the organisation expects them to process massive volumes and rates them solely on speed.

Real oversight can involve several practices. The recruiter should see the justification and source material. They should be able to change a decision without negative organisational consequences. The system should log corrections and enable analysis of where humans most often disagree with the AI. For automated actions, a kill switch mechanism is needed. Candidates should be able to report inaccuracies, and the organisation needs a procedure to handle such reports.

It is worth separating human review from doing all the work manually. Oversight does not mean a human has to re-analyse every element from scratch, because then the tool loses its purpose. They should, however, check decisive issues, especially before a negative decision. A proportionate approach can be applied: more automation in low-risk administrative tasks, more control where the result affects the candidate's access to the next stage.

The organisation should also train users to recognise situations where the system might fail: unconventional CVs, industry changes, documents in another language, employment gaps, incomplete data, conflicting sources or very narrow criteria. Human-in-the-loop is only effective when the human knows what to look for.

26. AI Scoring and GDPR

The AI Act does not replace GDPR. If AI Scoring processes candidates' personal data, the organisation must still meet obligations under data protection laws. This covers legal basis, transparency, data minimisation, purpose limitation, accuracy of information, security, retention and the exercise of data subject rights.

The first question is: who is the controller and who is the processor. The company running the recruitment usually determines the purposes and means of processing candidates' data, and the ATS provider acts in a specific capacity as a processor. However, the actual split depends on the service design, including whether the provider or its sub-processors use the data for their own purposes, model training, testing or service improvement. Roles, instructions and sub-processors should be clearly described in the contract.

The second question relates to purpose and legal basis. You cannot assume that candidate consent solves every problem. The basis depends on the type of recruitment, stage, national law and whether data is to be used in future projects as well. Special categories of data must be analysed separately. If a model can technically infer origin, health status, views or other sensitive traits, this does not mean the organisation can use such inferences in scoring.

The third area is minimisation. The system should only analyse information relevant to the purpose. Photos, age, home address, family situation or activity unrelated to work should not affect the assessment just because they are in the document or profile. Particular caution is needed when using data from LinkedIn and other external sources. The availability of information on the internet does not automatically mean free rein to process it.

The fourth area is the information obligation. Candidates should receive clear information that their data is analysed using AI, for what purpose, what data categories are used, where they come from, who they are disclosed to, how long they are stored, and what rights they have. Information should not hide key facts under a generic phrase like 'we use modern technologies'. It is also worth explaining the role of the score: whether it only supports the recruiter's work or triggers further actions.

Article 22 of the GDPR is of particular significance, as it concerns decisions based solely on automated processing which produce legal effects or similarly significantly affect individuals [2]. Scoring that supports a recruiter does not automatically mean a solely automated decision. However, if the score leads, without real human involvement, to rejecting an application, blocking the next stage or hiding a profile from the recruiter, the risk of entering this area is much higher. The 'human in the loop' must have a real impact, not just a technical option to approve the system's decision.

The fifth area is data accuracy and the ability to correct it. If a parser misreads a date, LinkedIn is outdated or a form contains an incorrect answer, the score might be wrong. Candidates should be able to correct data, and the organisation should establish whether scoring is recalculated after a correction.

Before implementation, it is worth conducting a Data Protection Impact Assessment (DPIA), especially when technology systematically evaluates individuals and can significantly affect their situation. The UK’s ICO is not an EU law-applying body, but its audits of recruitment tools serve as a practical reference. The regulator highlighted, among other things, the need to perform a DPIA at the procurement stage, clearly divide responsibilities, limit data, test fairness and transparently inform candidates [6].

Retention is also important. Data and scoring results should not be stored indefinitely just because a large database might be useful someday. The organisation should define how long it stores CVs, answers, scores, justifications and logs, and what happens to them after consent is withdrawn or the processing purpose expires.


Legal note: this section is for informational purposes and does not replace a legal analysis of a specific implementation. The scope of obligations depends on how the system operates and is used, the data types, the roles of the parties and national regulations.


27. AI Scoring and the AI Act

The EU AI Act applies a risk-based approach. Annex III lists AI systems intended to be used for recruitment or selection of persons, notably to place job adverts, screen and filter applications, and evaluate candidates [1]. This means AI Scoring used in recruitment is in an area of high regulatory focus. However, we should not simplify this to state that every feature containing AI and CVs automatically has the same legal status.

Classification depends on the intended purpose, design and actual impact on the process. Article 6(3) provides a possibility to deem some systems listed in Annex III as non-high-risk if they do not pose a significant risk of harm to the health, safety or fundamental rights of natural persons, including by not materially influencing the outcome of decision-making, and meet specific conditions. However, a system remains high-risk if it profiles natural persons, and a provider who considers that a system is not high-risk must document its assessment. The Commission published draft detailed guidelines on this classification in 2026 [3].

In practice, a feature that assigns a score to candidates and materially affects review order, advancement to the next stage or rejection requires highly cautious analysis. You cannot resolve classification solely based on names like 'assistant', 'recommendation' or 'scoring'. What matters is the intended use described by the provider and how the organisation actually uses the score.

For systems classified as high-risk, the AI Act outlines extensive requirements for providers. These include a risk management system, quality and data governance, technical documentation, automatic logging of events, transparency for the user, human oversight capabilities, and appropriate levels of accuracy, robustness and cybersecurity, alongside post-market monitoring. Depending on the scenario, a conformity assessment, declaration of conformity, CE marking and registration in the EU database are also required.

Obligations also apply to entities using the system, i.e., deployers. The recruiting organisation should use the solution in accordance with instructions, assign competent persons for oversight, ensure the relevance of input data within its control, monitor operation, and react to issues. In specific situations, individuals subject to a high-risk system must be informed of its use and the type of decisions supported [1]. The AI Act also provides a right to a clear and meaningful explanation of the role of the AI and the main elements of the decision in situations outlined in Article 86.

Human oversight cannot be formal. The person overseeing should understand the system's capabilities and limitations, be aware of automation bias, have access to sufficient information, and possess genuine authority to bypass, reverse or stop operations. A recruiter who automatically accepts every suggestion does not provide real oversight just because they formally clicked a button.

AI literacy is also significant. Organisations should ensure that individuals using AI possess an adequate level of knowledge and competence, taking into account their role, experience, context of use and the persons affected by the system. In practice, recruiter training should cover not only feature usage, but also score interpretation, distinguishing lack of data from lack of competence, discrimination risks, data protection and error reporting procedures [5].

Not every private company using a recruitment system will be required to carry out a fundamental rights impact assessment under Article 27. This obligation applies to specific categories of deployers and uses. Nevertheless, an impact assessment can be a good practice, and a DPIA under GDPR may be required regardless of the AI Act.

The timeline is important. According to current Commission information, following amendments introduced by the AI Omnibus, high-risk system requirements in Annex III, covering employment among others, are to apply from 2 December 2027 [4]. This does not mean organisations can ignore the topic until late 2027. Obligations regarding AI literacy apply from 2 February 2025, and oversight of them starts according to the Commission's timeline in 2026 [5]. GDPR, anti-discrimination laws and employment law apply regardless of this timeline right now.

For organisations, the practical takeaway is simple: do not wait until the last minute. It is worth creating a registry of used AI features, defining their purpose, providers, data sources, impact on decisions, oversight mechanisms and issue reporting procedures. Providers should be asked about classification, documentation, tests, logs, sub-processors, processing locations, model changes and compliance plans. The AI Act is not solely a legal team problem. It impacts product design, procurement, HR, security and the daily work of recruiters.


The safest assumption is: the greater the impact of the score on a candidate's chance of staying in the process, the stronger the documentation, human control and monitoring should be.


28. How to choose an AI Scoring tool?

Comparing tools solely based on demonstration quality is risky. On a demo, you can select perfectly written CVs and roles for which the score looks convincing. Real quality is revealed with incomplete, unusual, multilingual, conflicting documents coming from different industries. Therefore, purchase should involve testing on your own, properly secured data or a prepared set of representative profiles.

The first group of questions relates to features. What is evaluated: the application, the candidate, or both? How are requirements defined? Can the user edit and weight them? How does the system distinguish missing data from unmet criteria? Does it show the justification and source fragments? Does it generate questions? Can the score be recalculated after data changes? Does scoring work in multiple languages?

The second group relates to control. Can automated actions be turned off? Does the system log who triggered the evaluation and which requirement version was used? Can the user report an error? Does the provider inform about model changes? Can the administrator restrict access to the feature and set usage rules?

The third group relates to quality. How does the provider test accuracy, stability and bias? On what roles and languages? Do they share the methodology and known limitations? How do they react to errors? Do they measure differences between groups in a legally compliant way? Does the model generate deterministic or sufficiently repeatable answers?

The fourth group relates to data and security. Where is data processed and stored? Is it sent to third-party model providers? Is it used to train public or proprietary models? How long is it kept in logs? What are the bases for transfers outside the EEA? What do encryption, access control, audit and data deletion look like? Who is the sub-processor?

The fifth group relates to law. How does the provider classify the feature under the AI Act and on what basis? What is the compliance plan? What instructions will the deployer receive? Will documents needed for a DPIA, risk assessment and information obligations be available? Does the contract clearly define roles in data protection?

The sixth group relates to implementation. Does the provider train the team not only on usage, but on responsible interpretation? Do they help calibrate criteria? Do they allow starting with a pilot? Can the client measure effects? The best algorithm won't help if users don't trust the system or trust it too much.

29. How to measure the effectiveness of AI Scoring?

Implementing AI Scoring should have a measurable goal. The simplest indicators relate to efficiency: time from application to first analysis, average profile review time, number of applications analysed in a given timeframe and time to first contact. This data shows whether the tool actually relieves the team.

The second group of indicators relates to utility. You can measure how often the recruiter agrees with the score, how often they correct it, whether they use suggested questions, and whether summaries reduce the need for re-reading. However, agreement with a human is not a measure of truth on its own. It is worth analysing justifications and cases of discrepancy, not just the agreement percentage.

The third group relates to process quality. Do candidates with high scores more often pass screening after competency verification? How many valuable candidates were found in lower brackets? Does the score differ significantly between recruiters using the same criteria? Are fewer candidates waiting long for a response?

The fourth group relates to risk. Complaints, data errors, inaccurate explanations, unequal score distributions and automated actions with negative outcomes should be monitored. Where possible and in compliance with the law, fairness tests should be conducted. It is not enough to check the system once before implementation. Roles, data, models and how users work change.

The fifth group relates to human behaviour. Are recruiters reviewing lower-scored profiles? Are they rejecting candidates faster but without justification? Does the Hiring Manager look at the score first, and only then the CV? You can analyse logs and conduct short qualitative audits. The goal is to detect automation bias and situations where the feature takes on a larger role than planned.

Success should not be evaluated solely by the number of hires with high scores. If recruiters primarily contact this group, such a result is partly a self-fulfilling prophecy. Comparisons, control samples and conscious analysis of candidates outside the top positions are needed.

30. How does AI Scoring work in Recruitify?

AI Scoring in Recruitify is used to evaluate an application or candidate in the context of a specific recruitment project. The feature does not assign a permanent value to a person and does not automatically search for individuals in the database. The starting point is a profile already within the project, and the requirements defined for that recruitment.

The process begins with the Requirements field. The recruiter describes what the project actually requires: what experience is essential, what elements increase fit, which gaps are acceptable, and what should be verified during the interview. The more specific and job-related the requirements, the more useful the assessment will be. Generic phrases like 'good communicator', 'dynamic person' or 'cultural fit' do not give the AI or recruiter a sufficiently precise point of reference.

Recruitify analyses available information from the CV, LinkedIn profile and answers provided in the application form. Thanks to this, the score does not have to rely on a single document. The form can provide data on availability, expectations, working model, licences or experience in an area the candidate did not describe in their CV. LinkedIn can complement history, but is not treated as automatically more reliable than the application document. Discrepancies should lead to verification.

The result has both a numerical and contextual dimension. Recruitify awards a score from 0 to 100, but does not limit itself to the number. The user receives a candidate description in the context of the role, strengths, weaknesses and suggestions for questions to ask during the meeting. The goal is to help prioritise work, grasp the profile faster and prepare the screening better.

A score from 0 to 100 is not a probability of hire. A candidate with a score of 90 is not 'ten per cent better' than someone with a score of 80. The number shows the outcome of comparing available data against the criteria of a specific project. The exact same candidate can receive a different score across two recruitments because requirements change. The score can also change after data updates or refining the Requirements.

After generating the assessment, the recruiter should read the justification and check it against the source material. Particular attention should be paid to weaknesses and areas the system could not confirm. A lack of mention in the CV does not have to mean a lack of competence. In many cases, the best next action is to use a suggested question or ask the candidate to supplement information.

Recruitify clearly separates AI Scoring from AI Matching. Scoring evaluates an application or candidate already being analysed in a project. AI Matching, which is under development, is designed to search the database and suggest individuals potentially fitting the project to the recruiter. These are two different tasks: evaluating a specific profile versus finding and recommending profiles from the database.

The best practice is to start with tests on a few projects with clear requirements. The team can compare scores with their own evaluation, check justification quality, review a portion of low-scoring profiles, and refine how Requirements are written. AI Scoring is meant to support the recruiter's decision, not automatically reject candidates.





31. The future of AI Scoring

AI Scoring will shift from a simple percentage towards more complex, but also more controlled, decision support. The most valuable solutions will not try to create a single 'truth about the candidate'. They will combine different sources, show uncertainty, point out the basis of conclusions and tailor the evaluation method to the stage of the process.

We can expect a greater role for structured rubrics. Instead of a general prompt like 'evaluate the candidate', the system will work on criteria with a definition, weight, source of proof and level of certainty. Development may involve comparing candidates, but should avoid situations where a difference of a few points is treated as an objective advantage of one human over another.

The second direction will be integration with subsequent stages. Scoring before the interview can generate questions, and after the interview, the system can help organise notes according to the same scorecard. It will be important to separate declared from confirmed information and ensure that interview data is processed in compliance with the law.

The third direction will be testing and auditability. Regulations, client requirements and market maturity will increase pressure to document models, logic, data, accuracy and bias. Providers will have to show not just an attractive interface, but also a risk and change management process.

The fourth direction will be combining scoring with AI Matching and talent intelligence. The system can first suggest profiles from the database, then evaluate them in the project, and later support communication and process analysis. This combination increases productivity but also the risk of creating a closed loop of automated recommendations. Every stage should have a clear goal, separate metrics and the possibility of human intervention.

The most significant shift, however, will be organisational. Companies will stop asking simply 'do we have an AI feature?', and start asking 'what decision does this feature support, on what data, with what risk, and who is responsible for control?'. It is the answers to these questions that will determine whether AI Scoring improves recruitment or merely gives old problems a modern interface.

32. AI Scoring FAQ - Frequently Asked Questions

What is AI Candidate Scoring?

AI Scoring is an AI-supported evaluation of a specific application or candidate against the requirements of a specific project. It can take the form of a numerical score, a contextual description, or both elements simultaneously. It should not be understood as a universal assessment of human worth.

Does AI Scoring evaluate the candidate or the application?

It can evaluate both. An application is a specific entry to a project, often alongside form responses. A candidate can be added to the project by a recruiter without a new application. In both cases, the result should relate to the fit with the given project's requirements.

Are AI Scoring and AI Matching the same thing?

No. AI Scoring evaluates an application or candidate already being analysed in a project. AI Matching searches the database and suggests candidates potentially fitting the project. The first feature answers 'how does this profile stack up?', and the second 'who is worth finding and considering?'.

Does a score of 85 mean an 85 per cent chance of being hired?

No. A score from 0 to 100 is a fit scale used by a specific system. It is not a probability of hiring or job success. Its meaning depends on criteria, data and calculation methods.

Is a candidate with a score of 90 better than a candidate with a score of 80?

This cannot be determined solely on the basis of scoring. The difference might stem from CV completeness, how experience was described or a single heavily-weighted criterion. The score helps set analysis priority but does not replace evaluating the entire profile.

Can AI Scoring automatically reject candidates?

Technically, the score can be linked with automation, but automatically rejecting solely based on a threshold significantly increases the risk of error, discrimination and entering the territory of solely automated decisions. A safer approach is using scoring for prioritisation and ensuring a real human review.

Does the recruiter have to read the whole CV after receiving a score?

Scoring can shorten the analysis and point out key fragments, but before a significant decision, the recruiter should check the source material. The review scope can depend on the stage and risk. The score and generated summary alone should not be the sole basis for rejection.

What data does AI Scoring analyse in Recruitify?

In Recruitify, scoring can factor in project requirements and information from the CV, LinkedIn profile and application form responses. The scope of information used depends on the data available in the profile and project.

Is LinkedIn more reliable than a CV?

Not automatically. A LinkedIn profile can be newer or broader, but it is also created by the candidate and can contain gaps. It is best to treat sources as complementary and flag inconsistencies for verification.

What happens if important information is missing from the CV?

A good system should indicate a lack of confirmation, rather than automatically assuming a lack of competence. The recruiter can check other sources or ask a question during the interview. Critical data is best collected in the application form.

Does AI Scoring recognise synonyms and context?

Modern language models can connect similar meanings and recognise experience described in different words. However, they are not infallible. Semantic connections should be explained and checkable in the document.

Does AI Scoring detect if a CV was written by AI?

This is not the primary task of scoring, and there is no foolproof method to detect every AI-generated text. It is more important to verify whether the declared competencies are real. Suggested questions can help verify them during the interview.

Can AI Scoring detect lies in a CV?

No. It can spot inconsistencies between sources or vague fragments, but it does not confirm the truth of experience. Verification requires a conversation, test, references or documents.

Does AI Scoring evaluate soft skills?

It can analyse information suggesting certain experiences, but a CV is not a sufficient source for a reliable evaluation of empathy, communication, resilience or leadership style. The system should suggest questions rather than issue categorical evaluations.

Can the score change?

Yes. The score can change after updating the CV, supplementing the form, changing project requirements or updating features. Therefore, it should be treated as the result of a specific analysis at a given time.

Can you compare scores across different projects?

Usually, this makes no sense. Every project has different criteria, weights and score distributions. A score of 70 in one recruitment does not necessarily mean the same as 70 in another.

Can AI Scoring reduce recruiter bias?

It can increase consistency and limit some random evaluations, but it can also introduce or reinforce bias stemming from criteria, data and the model. Testing, monitoring and the ability to correct are needed.

Does removing names and photos solve the discrimination problem?

Not fully. Anonymisation limits some signals, but other information can indirectly indicate gender, age, origin or social status. Bias must be analysed across the entire process.

Do you need to inform candidates about using AI?

In many situations, information obligations arise under GDPR, and for high-risk systems, the AI Act outlines additional rules for informing individuals subject to the system. The scope of information depends on the specific implementation, but transparency should be a standard even when the law only requires a minimum.

Is AI Scoring subject to the AI Act?

Systems intended to be used for screening and filtering applications and evaluating candidates are listed in Annex III of the AI Act as employment-related uses. The final classification of a specific feature depends on its purpose, impact, usage and exceptions in Article 6(3). This requires an individual analysis.

When do AI Act rules for high-risk systems in recruitment start to apply?

According to the current timeline, following the entry into force of the AI Omnibus, rules for high-risk systems in Annex III are to apply from 2 December 2027. Some other provisions, including those on AI literacy, apply earlier.

Is AI Scoring compliant with GDPR?

The feature name alone does not determine compliance. Compliance depends on the legal basis, data scope, transparency, retention, security, division of roles, the ability to realise candidate rights, and how decisions are made. A tool can support a compliant process, but it does not 'solve GDPR' for the organisation.

Is a DPIA required?

In many implementations of candidate scoring, a Data Protection Impact Assessment may be required or at least highly justified due to systematic evaluation of individuals and potentially significant impact. The decision should be made based on the specific process and consultation with a DPO or lawyer.

What does human-in-the-loop mean?

It means real, not formal human involvement. The recruiter should understand limitations, see the score basis, be able to challenge it, and have actual authority to make a different decision. Merely clicking confirmation is not enough.

Will AI Scoring replace recruiters?

It should not. It automates part of the analysis and structuring of information, but it does not replace conversations, understanding context, evaluating motivation, negotiating or responsibility for the decision. It shifts the weight of a recruiter's work from manual screening to interpretation, verification and human connection.

For what type of recruitment is scoring not a good solution?

It can deliver less value in processes with very few candidates, unclear criteria, a key role for confidential market knowledge or a strong emphasis on competencies evaluated only in interaction. It should not be forced just because it is available.

How to start testing AI Scoring?

It is best to select a few roles with clear requirements, prepare criteria, test the feature on representative profiles and compare results with recruiters' independent evaluations. Then, establish human review rules, train users, and measure both time, quality and risk.

33. Summary - AI Scoring does not choose the person, it supports the evaluation

AI Candidate Scoring is one of the most practical applications of artificial intelligence in modern recruitment systems. It can help teams structure incoming applications faster, match profiles to requirements more consistently, prepare interviews better and collaborate more smoothly with Hiring Managers or clients. Its value, however, does not come from the score of 0 to 100 itself. What matters most is the context, the explanation, the quality of criteria, and the ability to verify and challenge the evaluation.

Scoring is not AI Matching. It is not primarily used to find candidates in the database, but to evaluate an application or candidate in a specific project. It is not a CV parser, though it uses extracted data. Nor is it a simple killer question or an automated rejection decision. It can link with these features, but each has a different task and a different risk profile.

The biggest mistake would be to assume the number closes the discussion. The score can help set the order of work, but it should not define a candidate's value or replace a conversation. Documents are incomplete, criteria can be flawed, and models can generate inaccurate or biased conclusions. That is why a good system highlights not only fit, but also information gaps, limitations and questions for verification.

In 2026, a responsible implementation of AI Scoring requires combining three perspectives. The first is productivity: whether the feature genuinely saves time and improves process quality. The second is recruitment practice: whether criteria are job-related and the recruiter retains their own judgment. The third is governance: whether the organisation controls data, bias, security, transparency and compliance with GDPR and the AI Act.


The best AI Scoring does not tell a recruiter whom to hire. It helps them see faster what is already known, what is still unknown and what is worth asking before they make a decision.


Sources and Further Reading

[1] European Parliament and Council of the EU, Regulation (EU) 2024/1689 - Artificial Intelligence Act, notably Annex III point 4 on employment, recruiting and candidate evaluation: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689

[2] European Parliament and Council of the EU, Regulation (EU) 2016/679 - GDPR, notably Article 22 on automated individual decision-making: https://eur-lex.europa.eu/eli/reg/2016/679/oj

[3] European Commission, Draft Commission guidelines on the classification of high-risk AI systems, including materials on Annex III AI Act, 2026: https://digital-strategy.ec.europa.eu/en/library/draft-commission-guidelines-classification-high-risk-ai-systems

[4] European Commission, Guidelines for providers and deployers of AI high-risk systems - current timeline of high-risk rules application, 2026: https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-high-risk-systems

[5] European Commission, AI Literacy - Questions & Answers, current information on Article 4 AI Act: https://digital-strategy.ec.europa.eu/en/faqs/ai-literacy-questions-answers

[6] Information Commissioner’s Office, Thinking of using AI to assist recruitment? Our key data protection considerations, 2024: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/11/thinking-of-using-ai-to-assist-recruitment-our-key-data-protection-considerations/

[7] Wen A. et al., FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations, 2025: https://arxiv.org/abs/2504.01420

[8] Hunkenschroer A. L., Luetge C., Ethics of AI-Enabled Recruiting and Selection: A Review and Research Agenda, Journal of Business Ethics, 2022: https://link.springer.com/article/10.1007/s10551-022-05049-6

[9] Information Commissioner’s Office, AI tools used in recruitment - audit outcomes and recommendations, 2024: https://ico.org.uk/action-weve-taken/audits-and-overview-reports/2024/11/ai-tools-used-in-recruitment/

[10] National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0): https://www.nist.gov/itl/ai-risk-management-framework

[11] Information Commissioner’s Office, Automated decisions can streamline the hiring process with the right safeguards in place, 2026: https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2026/03/automated-decisions-can-streamline-the-hiring-process-with-the-right-safeguards-in-place/

[12] Recruitify, What is an ATS? A complete guide to applicant tracking systems (2026): https://www.recruitify.ai/blog/what-is-an-ats-complete-guide-to-applicant-tracking-systems-(2026)/


Editorial Note: this article describes principles and good practices at a general level. It does not constitute legal advice. Classification and obligations concerning a specific AI system depend on its purpose, design, data, implementation method and actual impact on decisions.

News & Updates

Stay up-to-date with the latest innovations, features, and tips about Recruitify!

First Name
Email

By providing your email address within the newsletter sign-up form, you confirm its processing to send marketing information regarding the Administrator’s products and services. The Administrator of your personal data processed for the abovementioned purposes is Recruitify Spółka z o.o., based in Warsaw, Poland (KRS 0000709889). For more information on the principles of personal data processing and the rights of data subjects, please check the Privacy Policy.

Share

Published

Category

Applicant Tracking System

Author

Iwo Paliszewski