Justice Fairness Metrics
Understanding Fairness Metrics
Fairness metrics are the scientific backbone of the Fairness Diagnostic Kit (FDK™).
They transform fairness and ethics — traditionally qualitative ideas — into quantitative, measurable indicators.
Each metric evaluates how fairly a dataset or decision-making system treats different social groups such as gender, race, age, or region.
By comparing outcomes mathematically, these metrics reveal whether any group receives disproportionate results in areas like sentencing, bail, or predictive risk assessments.
Together, they provide an objective framework for diagnosing and improving fairness in justice-related algorithms.
Categories of Fairness Metrics in the Justice Pipeline
A) Core Group Fairness
These metrics assess how equally different demographic groups are represented in outcomes — the foundation of fairness evaluation.
1. Statistical Parity Difference
2. Disparate Impact
3. Selection Rates by Group
4. Predicted Positives per Group
5. Predicted Negatives per Group
B) Error and Performance Fairness
These metrics examine whether all groups experience similar error rates in AI predictions — ensuring fairness in both accuracy and mistakes.
6. False Positive Rate (FPR) Difference
7. False Positive Rate (FPR) Ratio
8. False Negative Rate (FNR) Difference
9. False Negative Rate (FNR) Ratio
10. True Positive Rate (TPR) Difference
11. True Positive Rate (TPR) Ratio
12. True Negative Rate (TNR) Difference
13. True Negative Rate (TNR) Ratio
14. Error Rate Difference
15. Error Rate Ratio
16. Predictive Equality
17. Disparate Mistreatment Index
C) Equality of Opportunity and Treatment
These metrics evaluate whether every group has an equal chance of receiving a positive outcome under similar conditions.
18. Equalized Odds Difference
19. Equal Opportunity Difference
20. Average Odds Difference
21. Average Absolute Odds Difference
D) Error Distribution and Subgroup Fairness
These detect bias concentrated in smaller subgroups that general averages might miss, offering deeper diagnostic precision.
22. False Discovery Rate (FDR) Difference
23. False Discovery Rate (FDR) Ratio
24. False Omission Rate (FOR) Difference
25. False Omission Rate (FOR) Ratio
26. Error Disparity Subgroup
27. MDSS Subgroup Discovery Score
E) Robustness and Worst-Case Fairness
This category identifies the least-fair performance scenario across all groups — ensuring fairness even under adverse conditions.
28. Worst Group Accuracy
29. Worst Group Loss
30. Composite Bias Score
31. Validation Robustness Score
F) Calibration and Predictive Reliability
Calibration metrics ensure predictions have consistent meaning and reliability for every group.
32. Slice AUC Difference
G) Causal and Counterfactual Fairness
These metrics assess fairness through what-if scenarios, testing whether outcomes would change if a person’s group attribute were different — holding all else equal.
33. Counterfactual Fairness Score
34. Causal Effect Difference
H) Explainability and Temporal Fairness
The final category addresses how transparent model behaviour is over time and whether fairness remains stable as data evolves.
35. Feature Attribution Bias
36. Temporal Fairness Score
Conclusion
These 36 metrics collectively convert ethical judgement into empirical measurement — turning fairness from a philosophical discussion into a reproducible, data-driven science.
They make it possible for both professionals and citizens to see fairness numerically, fostering transparency and accountability in justice systems.
For detailed mathematical definitions, equations, and validation studies, please refer to Appendix C of the Fairness Diagnostic Kit (FDK™) Book, where each metric is formally described and referenced.