Evidence-Based Detection
BiasClean v2.0 uses a transparent, UK-aligned detection framework that identifies demographic disparities in datasets before any mitigation is applied. The detection process is built on statistical evidence, regulatory principles and demographic structures documented in official UK data sources.
This ensures that every detected disparity is meaningful, measurable and grounded in real-world inequality patterns.
1. What “Evidence-Based Detection” Means
Evidence-based detection means BiasClean does not guess, assume or apply generic global rules.
Instead, it follows three principles:
1. The detection uses real UK demographic evidence
Demographic groups are defined using Census-aligned structures for Ethnicity, Region, Age, Disability and Migration.
2. The system measures disparities using statistical ratios
BiasClean computes representation ratios and proportional imbalance scores for each demographic feature.
3. Findings are compared against known UK inequality patterns
Detected disparities are interpreted within the context of documented national-level inequality.
This creates a detection method that is interpretable, policy-aligned and sector-specific.
2. How BiasClean Measures Disparity
BiasClean analyses each demographic feature separately and reports where the dataset shows meaningful imbalance.
Step A — Group Definition
Groups are built using UK-standard categories from ONS and UK regulators.
Examples:
Ethnicity: White, Mixed, Asian, Black, Other
Age Bands: 18–21, 22–29, 30–39, 40–49, 50–59, 60+
Regions: Nine English regions + Wales
Disability, Gender and Migration categories follow EHRC guidance.
Step B — Representation Ratios
For each group, the system computes:
Group Representation Ratio (GRR)
GRR = (group proportion in dataset) ÷ (expected proportion)
A ratio below 1.00 indicates under-representation.
A ratio above 1.00 indicates over-representation.
Step C — Disparity Scoring
BiasClean converts each GRR into a disparity level:
Critical Disparity — < 0.60 or > 1.40
Major Disparity — < 0.75 or > 1.25
Moderate Disparity — < 0.90 or > 1.10
Aligned — within ±10% of expected structure
These thresholds are used across all domains for consistent measurement.
Step D — Weighted Detection
After raw disparities are computed, the system applies domain-specific weights so findings match real UK inequality patterns.
For example:
Ethnicity has the highest influence in Justice and Hiring.
SES and Region dominate in Finance.
Disability and Gender carry stronger impact in Health.
The final “Detection Score” therefore represents:
Disparity × Domain Weight × Evidence Strength
3. Why Detection Needs UK-Specific Evidence
Countries differ in demographic structure, regional inequality, protected characteristic definitions and census frameworks.
A US-style or EU-generic detection method would misdiagnose gaps in a UK dataset.
BiasClean solves this by using:
UK Census 2021 population baselines
ONS regional distribution patterns
MoJ, NHS, FCA and DfE disparity evidence
EHRC definitions of protected characteristics
This grounds the detection system in reliable national data.
4. What Users See in the Detection Report
When a user uploads a dataset and selects a domain, BiasClean returns a clear, domain-aware detection summary.
The report includes:
• Feature Disparity Table
Listing GRR values for each demographic group.
• Four-Tier Disparity Colours
Critical (red)
Major (amber)
Moderate (yellow)
Aligned (green)
• Weighted Detection Score
Showing domain-specific impact level.
• Narrative Interpretation
Plain-language explanation of which groups are most affected and why.
• Evidence Reference
Quotes the corresponding UK evidence source (e.g., MoJ, ONS, NHS England) that validates the detected pattern.
This ensures detection results are understandable for both technical and non-technical users.
5. Example of Evidence-Aligned Interpretation
Here are examples of how BiasClean interprets detected disparities using UK evidence:
If Black groups are under-represented in Finance datasets:
This is flagged as structurally consistent with documented UK credit access disparities (FCA, Bank of England).If disability groups are under-represented in Education outcomes:
BiasClean links this to known SEND and SEN support gaps (DfE, EHRC).If 18–21 year-olds are over-represented in Justice outcomes:
This aligns with MoJ evidence of age-linked disproportionality.
This evidence-aligned interpretation is what makes BiasClean suitable for compliance, auditing, and ethical AI deployment.
6. Citations and Official Data Sources
ONS – Census 2021 (Population, Region, Age, Disability, Migration)
ONS – Regional Inequality & IMD
https://www.gov.uk/government/statistics/english-indices-of-deprivation-2019
Ministry of Justice – Race and the Criminal Justice System
https://www.gov.uk/government/collections/race-and-the-criminal-justice-system
NHS England – Health Inequalities Hub
https://www.england.nhs.uk/about/equality/equality-hub/health-inequalities/
Financial Conduct Authority – Consumer Fairness & Inclusion
https://www.fca.org.uk/firms/consumer-duty
Bank of England – Credit & Household Finance Data
https://www.bankofengland.co.uk/statistics
Department for Education – Attainment & Exclusion Statistics
https://explore-education-statistics.service.gov.uk/find-statistics
Sutton Trust – Social Mobility Reports
Education Policy Institute – Inequality Research
British Business Bank – Entrepreneurship & Diversity
https://www.british-business-bank.co.uk
Electoral Commission – Participation & Representation Data
https://www.electoralcommission.org.uk