Every Xobin assessment starts from a written construct definition and a job-relevance map that ties each competency back to the role it is meant to predict. Items are authored by subject-matter experts against that blueprint, piloted on a representative sample, and only then promoted into the live item bank. This is the step that makes the difference between a test that feels relevant and a test that actually is.
We characterise every assessment on three axes and document each of them: reliability — using Cronbach's α, split-half, and, where repeat-taker data exists, test–retest — construct validity through item analysis and factor structure, and criterion validity against real hiring outcomes contributed by opted-in customers. Items with weak discrimination or unclear difficulty are pruned before the assessment ships, and the full evidence dossier is available to enterprise customers on request.
Validity is a claim about a test used in a specific context, not a permanent property of the test itself. When the target role shifts, the candidate population changes, or the item bank is refreshed, the evidence has to be re-established. Xobin runs validation as a rolling process with scheduled re-checks and release gates — an assessment that falls out of tolerance is paused, not quietly kept in production.
Every validation study described here ships behind the assessments hiring teams run every day. Teams who want the same rigour in their own hiring can use it directly through Xobin.
The methodology described here is what runs inside every Xobin assessment and AI interview — so hiring teams can make talent decisions on evidence, not instinct. See Xobin.