Methodology
Understand how UXit calculates evaluation scores, category results, grades, and trend comparisons.
Overview
UXit scores an interface against the criteria in a selected Guideline Set. Each included criterion is answered as Pass, Fail, or NA (not applicable), producing a score that shows how closely the reviewed experience matched the criteria used for that evaluation.
This is an operational score: it measures whether defined criteria were satisfied. It is not a general-purpose guess at how people felt, how fast they moved, or whether every possible user would experience the interface the same way. That boundary is intentional. It keeps the result auditable, repeatable, and easier to defend in review.
Because each result is tied to a specific guideline ID, Analytics can show both the summary and the source of the summary: the category, criterion, result, annotation, and image evidence behind the number.
Validation Model
UXit's scoring model is built around traceable evidence instead of broad sentiment. The score is valid when the criteria are clear, the evaluator applies them consistently, and the same Benchmark and Guideline Set are used for comparison over time.
The model depends on four ideas:
- Defined criteria: The
Guideline Setestablishes what is being measured before the evaluation starts. - Binary scoring: Scored criteria resolve to
PassorFail, which keeps the math direct and repeatable. - Excluded non-applicable criteria:
NAand excluded criteria are removed from scoring so they do not inflate or punish a result unfairly. - Traceability: Each score can be traced back to the category, guideline ID, condition text, notes, and images that produced it.
Methodology note
The score is only as strong as the criteria behind it. Clear, objective criteria make results easier to compare. Vague or heavily changed criteria make trend data less reliable.
Scoring Model
Each included criterion is converted into a scoring state:
Pass= 1Fail= 0NA(not applicable) = excluded from scoring
Only Pass and Fail are counted in the score calculation. Criteria excluded from the Guideline Set do not appear in the evaluation and are not counted. Included criteria marked NA are also excluded from the calculation for that evaluation only, because they did not apply to the reviewed version.
This keeps the denominator honest. A criterion that does not apply should not help the score, hurt the score, or create noise in the category result.
Category Scoring
For each category, let:
- Pass = number of criteria marked
Pass - Fail = number of criteria marked
Fail
Then:
Example:
If , the category is excluded from aggregation because it has no scored criteria in that evaluation.
Overall Score
If there are valid category scores, the overall score is:
The result is expressed as a percentage and then mapped to a grade.
Flat Score Variant
Without categories, let:
- TotalPass = total number of
Passresponses across all included criteria - TotalFail = total number of
Failresponses across all included criteria
If categories are ignored:
Worked Example
| Category | Pass | Fail | NA |
|---|---|---|---|
| A | 4 | 1 | 0 |
| B | 2 | 2 | 1 |
| C | 3 | 0 | 2 |
Using the table above:
Grade Thresholds
| Grade | Interval |
|---|---|
| A | 90 to 100% |
| B | 80 to 89.9% |
| C | 70 to 79.9% |
| D | 60 to 69.9% |
| F | 0 to 59.9% |
Evaluation Method
A reviewer goes through each included criterion and marks whether the current interface satisfies it. Scored criteria are treated as binary checks:
The final score is calculated from the binary Pass and Fail outcomes. NA and excluded criteria are left out of the scoring math.
This keeps the model consistent, easy to audit, and traceable to individual failed criteria. It is intended to measure criterion satisfaction and track change over time, not infer quality from ambiguous user-dependent signals that may vary between users or sessions.
In practice, that means a score should be read as evidence of how well the interface matched the selected Guideline Set, not as a universal claim that the design is good or bad in every possible context.
What the Score Represents
- Percentage of scored criteria that the interface satisfies
- Category-level performance against the selected
Guideline Set - Stable metric for comparing the same
Benchmarkover time - Directional trend to track improvement, regression, or tradeoffs
- Evidence that can be reviewed at the condition level
What the Score Does Not Represent
- User satisfaction or emotional response
- Perceived ease of use or aesthetic appeal
- Efficiency, speed, task completion time, or click accuracy
- Cognitive demand or user behavior patterns
- Quality outside the criteria included in the selected
Guideline Set
Using Results
- Focus on changes over time rather than single scores
- Review failed items to understand specific gaps
- Keep older evaluations to track trends and regressions
- Use the score to guide decisions, not to define success or failure by itself
- Pair scores with
Condition Detailswhen a decision needs proof, context, or handoff evidence