Este contenido está disponible en inglés mientras se revisa su traducción. · Machine translation preview
SCORING & BENCHMARKS

75th percentile.
75% correct?

They answer different questions. Explore the difference before using a score to make a hiring decision.

EXPLORE THE IDEAIllustrative, not live norm data
75th

percentile
Above 75% of this hypothetical reference group.

1st percentile99th percentile

75 out of 100 dots are highlighted to explain relative position. This does not mean 75% of questions were answered correctly.

BEFORE YOU TRUST A PERCENTILE

Ask for its passport.

A benchmark should arrive with enough information to understand where it came from.

Reference population
Who is being compared?
Sample size
How much data supports the norm?
Test & norm version
Which assessment and scoring rules?
Limitations
Where should this comparison not be used?
CURRENT STATUS

Evidence before precision.

HireValid plans to use pilot data and suitable Intellect Council reference data where available. No completed validation study or live norm sample has been supplied for this website. The explorer above is educational.

How we handle evidence

Scoring methodology and responsible interpretation

What does a percentile mean?

A percentile describes relative position within a reference group. For example, a candidate at the 75th percentile scored above 75% of that documented group. It does not mean they answered 75% of questions correctly or have a 75% chance of succeeding in the job.

Compare results only when test versions, scoring rules and reference groups are appropriate. A score without its norm group can suggest more precision than the evidence supports.

Benchmarked by Intellect Council

HireValid plans to use pilot data and suitable Intellect Council reference data where available. No live norm sample size or validation result has been supplied for this website. The endorsement describes the intended relationship, not proof of a completed norming study.

Before releasing percentile reports, the team needs to document sample recruitment, representation, sample size and the limitations of each norm group. Early norms should be clearly labeled and refreshed as appropriate data accumulates.

How are raw results scored?

The planned initial approach uses classical test theory: correct responses, item difficulty and discrimination. A scaled score can summarize test performance. Overall assessment scores may combine tests with explicit weights.

Writing tasks require an explicit rubric and human review of AI-assisted scores. Employers should be able to inspect the rationale rather than receiving an unexplained number.

What makes a responsible benchmark?

Use a reference population that makes sense for the question you are asking. A broad working-population comparison is different from a role-specific applicant comparison. Neither should be treated as a universal pass mark.

Review item performance and group differences over time. Keep norm versions in reports so a later update does not silently change the meaning of an earlier hiring decision.

Start with what the role needs.

Find relevant testsCompare paid plans