Evidence-Based Detection

BiasClean v2.0 uses a transparent, UK-aligned detection framework that identifies demographic disparities in datasets before any mitigation is applied. The detection process is built on statistical evidence, regulatory principles and demographic structures documented in official UK data sources.

This ensures that every detected disparity is meaningful, measurable and grounded in real-world inequality patterns.

1. What “Evidence-Based Detection” Means

Evidence-based detection means BiasClean does not guess, assume or apply generic global rules.
Instead, it follows three principles:

1. The detection uses real UK demographic evidence
Demographic groups are defined using Census-aligned structures for Ethnicity, Region, Age, Disability and Migration.

2. The system measures disparities using statistical ratios
BiasClean computes representation ratios and proportional imbalance scores for each demographic feature.

3. Findings are compared against known UK inequality patterns
Detected disparities are interpreted within the context of documented national-level inequality.

This creates a detection method that is interpretable, policy-aligned and sector-specific.

2. How BiasClean Measures Disparity

BiasClean analyses each demographic feature separately and reports where the dataset shows meaningful imbalance.

Step A — Group Definition

Groups are built using UK-standard categories from ONS and UK regulators.
Examples:

  • Ethnicity: White, Mixed, Asian, Black, Other

  • Age Bands: 18–21, 22–29, 30–39, 40–49, 50–59, 60+

  • Regions: Nine English regions + Wales

  • Disability, Gender and Migration categories follow EHRC guidance.

Step B — Representation Ratios

For each group, the system computes:

Group Representation Ratio (GRR)
GRR = (group proportion in dataset) ÷ (expected proportion)

A ratio below 1.00 indicates under-representation.
A ratio above 1.00 indicates over-representation.

Step C — Disparity Scoring

BiasClean converts each GRR into a disparity level:

  • Critical Disparity — < 0.60 or > 1.40

  • Major Disparity — < 0.75 or > 1.25

  • Moderate Disparity — < 0.90 or > 1.10

  • Aligned — within ±10% of expected structure

These thresholds are used across all domains for consistent measurement.

Step D — Weighted Detection

After raw disparities are computed, the system applies domain-specific weights so findings match real UK inequality patterns.

For example:

  • Ethnicity has the highest influence in Justice and Hiring.

  • SES and Region dominate in Finance.

  • Disability and Gender carry stronger impact in Health.

The final “Detection Score” therefore represents:
Disparity × Domain Weight × Evidence Strength

3. Why Detection Needs UK-Specific Evidence

Countries differ in demographic structure, regional inequality, protected characteristic definitions and census frameworks.
A US-style or EU-generic detection method would misdiagnose gaps in a UK dataset.

BiasClean solves this by using:

  • UK Census 2021 population baselines

  • ONS regional distribution patterns

  • MoJ, NHS, FCA and DfE disparity evidence

  • EHRC definitions of protected characteristics

This grounds the detection system in reliable national data.

4. What Users See in the Detection Report

When a user uploads a dataset and selects a domain, BiasClean returns a clear, domain-aware detection summary.

The report includes:

• Feature Disparity Table
Listing GRR values for each demographic group.

• Four-Tier Disparity Colours

  • Critical (red)

  • Major (amber)

  • Moderate (yellow)

  • Aligned (green)

• Weighted Detection Score
Showing domain-specific impact level.

• Narrative Interpretation
Plain-language explanation of which groups are most affected and why.

• Evidence Reference
Quotes the corresponding UK evidence source (e.g., MoJ, ONS, NHS England) that validates the detected pattern.

This ensures detection results are understandable for both technical and non-technical users.

5. Example of Evidence-Aligned Interpretation

Here are examples of how BiasClean interprets detected disparities using UK evidence:

  • If Black groups are under-represented in Finance datasets:
    This is flagged as structurally consistent with documented UK credit access disparities (FCA, Bank of England).

  • If disability groups are under-represented in Education outcomes:
    BiasClean links this to known SEND and SEN support gaps (DfE, EHRC).

  • If 18–21 year-olds are over-represented in Justice outcomes:
    This aligns with MoJ evidence of age-linked disproportionality.

This evidence-aligned interpretation is what makes BiasClean suitable for compliance, auditing, and ethical AI deployment.

6. Citations and Official Data Sources

ONS – Census 2021 (Population, Region, Age, Disability, Migration)

https://www.ons.gov.uk/census

ONS – Regional Inequality & IMD

https://www.gov.uk/government/statistics/english-indices-of-deprivation-2019

Ministry of Justice – Race and the Criminal Justice System

https://www.gov.uk/government/collections/race-and-the-criminal-justice-system

NHS England – Health Inequalities Hub

https://www.england.nhs.uk/about/equality/equality-hub/health-inequalities/

Financial Conduct Authority – Consumer Fairness & Inclusion

https://www.fca.org.uk/firms/consumer-duty

Bank of England – Credit & Household Finance Data

https://www.bankofengland.co.uk/statistics

Department for Education – Attainment & Exclusion Statistics

https://explore-education-statistics.service.gov.uk/find-statistics

Sutton Trust – Social Mobility Reports

https://www.suttontrust.com

Education Policy Institute – Inequality Research

https://epi.org.uk

British Business Bank – Entrepreneurship & Diversity

https://www.british-business-bank.co.uk

Electoral Commission – Participation & Representation Data

https://www.electoralcommission.org.uk

Equality and Human Rights Commission – Disparity Reviews

https://www.equalityhumanrights.com