A score is useful when you can explain it.
Start with your real hiring requirements, agree their importance, and test the same scenarios in every shortlisted system. This tool records one system at a time. Export an assessment for each option using the same priorities.
- Set priorities before the demonstration. Use 1–5 for included criteria. Choose 0 only when a criterion does not apply.
- Record what happened. Rate a completed test from 0–5 and describe the evidence or gap. Leave an untested criterion blank.
- Compare equal coverage. A score based on a few easy tests is not comparable with a complete assessment. Finish the important tests and examine mandatory failures separately.
Your priorities. Your evidence.
The initial weights are an editable starting point for an in-house hiring team. They are not independently validated buyer preferences or a vendor ranking.
Work stays in this browser tab and is lost on refresh. Export a copy before leaving. Use synthetic test records and avoid candidate or confidential information in your notes.
Start blank, or load the clearly labelled synthetic example.
No included criteria have been tested yet. A blank assessment has no score. Evidence notes are your record of a test; their presence does not establish independent verification.
Prefer a spreadsheet? Download the blank CSV template. It includes the same criteria and method; enter your own formulas when calculating offline.
How the numbers are calculated.
Score on rated criteria = 100 × the sum of priority × rating, divided by 5 × the total priority weight of rated criteria. A rating of 0 contributes zero points and remains in that denominator. A blank rating is missing and stays outside it.
Priority weight rated = the weight of rated criteria divided by all included priority weight. Priority weight with evidence notes = the weight of rated criteria with a nonblank note, divided by all included priority weight. Notes can describe a failure; evidence coverage is not a measure of product quality.
A priority of 0 excludes a criterion from all calculations. If all priorities are excluded, the tool shows no score or coverage percentage. If nothing has been rated, there is no score and rating coverage is 0%.
Exported CSVs contain the current values and denominators, rather than spreadsheet formulas. User-entered text that could be interpreted as a spreadsheet formula is prefixed as text. If you edit or resave a CSV in another application, review that application’s handling of text before opening it elsewhere.
A high-looking number can still hide a blocker.
The “Replace with example” button loads a fictional assessment, not a rating of TiLab or another vendor. Only four of ten criteria have been tested:
| Criterion | Priority | Rating | Weighted points |
|---|---|---|---|
| Application workflow | 5 | 4 | 20 |
| Job publishing | 4 | 3 | 12 |
| Team evaluation | 4 | 4 | 16 |
| Access and confidentiality | 5 | 0 | 0 |
The result is 48 / 90 = 53.3% on tested criteria, with 18 / 40 = 45% of priority weight rated. Every rated criterion has a note, so evidence coverage is also 45%. The six untested criteria are not treated as successes or failures.
The confidential-job test failed. If that access requirement is mandatory, the team should resolve it before selecting the system even if later tests raise its overall score. An arithmetic average cannot overrule a requirement you agreed was essential.
Turn the assessment into a decision record.
Export the same scorecard for each shortlisted system and review them together. Keep the package name, test context, unresolved questions, and the people who participated with your decision notes. Ask a future operator to reproduce the most important tests.
For a small peer review, ask three recruiting colleagues to use the scorecard on a synthetic scenario. Record which criteria they interpreted differently, which evidence they could not obtain, and whether the weights reflected the intended workload. These are suggested next steps; this tool does not claim that peer testing has already taken place.