Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification

1. Introduction

  • Bad facial recognition can lead to false accusations
  • algorithms trained with biased data have resulted in algorithmic discrimination
  • With Word2Vec, the analogy man is to computer programmer as woman is to “X” was completed with “homemaker”
  • computer vision systems with inferior performance across demographics can have serious implications
  • Some face recognition systems have been shown to misidentify people of color, women, and young people at high rates
  • our work introduces a new face dataset composed of 1270 unique individuals that is more phenotypically balanced on the basis of skin type than existing benchmarks
  • this work introduces the first intersectional demographic and phenotypic evaluation of face-based gender classification accuracy.
  • 2. Related Work

    Automated Facial Analysis

  • Face recognition software is now built into most smart phones
  • companies such as Affectiva (Affectiva) and researchers in academia attempt to identify emotions from images of people’s faces
  • Faception has developed software that purports to determine an individual’s characteristics (e.g. propensity towards crime, IQ, terrorism) solely from their faces.
  • “The Perpetual Lineup” provides an in-depth analysis of the unregulated police use of face recognition
  • accuracies of face recognition systems used by US-based law are lower for people labeled female, Black, or between the ages of 18—30
  • The lack of datasets that are labeled by ethnicity limits the generalizability of research exploring the impact of ethnicity on gender classification accuracy.
  • our work explores gender classification on African faces
  • Benchmarks

  • Megaface (largest publicly available set of facial images) was composed utilizing Head Hunter (Mathias to select one million images from Flicker
  • LFW, gold standard benchmark for face recognition, was estimated to be 77.5% male and 83.5% White
  • 3. Intersectional Benchmark

  • We use the labels “male” and “female” to define genders
  • Due to phenotypic imbalances in existing benchmarks, we created a new dataset with more balanced skin type and gender representations.
  • 3.1. Rationale for Phenotypic Labeling

  • demographic labels for protected classes like race and ethnicity have been used for performing algorithmic audits
  • phenotypic (observable characteristics) labels are seldom used for these purposes
  • benchmarks consisting of lighter-skinned Blacks would not adequately represent darker-skinned ones
  • racial and ethnic categories are not consistent across geographies
  • Since race and ethnic labels are unstable, we decided to use skin type as a more visually precise label to measure dataset diversity
  • skin type was chosen as a phenotypic factor of interest because default camera settings are calibrated to expose lighter-skinned individuals
  • 3.2. Existing Benchmark Selection Rationale

  • IJB-A is a US government benchmark released by NIST
  • the dataset consisted of 500 unique subjects who are public figures
  • each subject was manually labeled with one of six Fitzpatrick skin types
  • The Adience benchmark contains 2, 284 unique individual subjects.
  • only one image of each subject was labeled for skin type.
  • 3.3. Creation of Pilot Parliaments Benchmark

  • Preliminary analysis of the IJB-A and Adience benchmarks revealed overrepresentation of lighter males, underrepresentation of darker females, and underrepresentation of darker individuals in general
  • We developed the Pilot Parliaments Benchmark (PPB) to achieve better intersectional representation on the basis of gender and skin type.
  • PPB consists of 1270 individuals from three sub-Saharan countries and three Nordic countries
  • 3.4. Intersectional Labeling Methodology

    Skin Type Labels – We chose the Fitzpatrick six-point labeling system to determine skin type labels

    Gender Labels – male and female

    Labeling Process
  • one author labeled each image with one of six Fitzpatrick skin types and provided gender annotations for the IJB-A dataset
  • Gender labels were determined based on the name of the parliamentarian, gendered title, prefixes such as Mr or Ms, and the appearance of the photo.
  • 3.5. Fitzpatrick Skin Type Comparison

    PPB provides substantially more darker-skinned unique subjects than IJB-A and Adience.

    4. Commercial Gender Classification Audit

    We evaluated 3 commercial gender classifiers.

    4.1. Key Findings on Evaluated Classifiers

  • All classifiers perform better on male faces
  • All classifiers perform better on lighter faces
  • All classifiers perform worst on darker female faces
  • Microsoft and IBM classifiers perform best on lighter male faces
  • Face++ classifiers perform best on darker male faces
  • 4.2. Commercial Gender Classifier Selection: Microsoft, IBM, Face++

    Only IBM provided confidence scores for face-based gender classification labels.

    4.3. Evaluation Methodology

  • TPR- true positive rate
  • Dark females had the highest gender classification error rate
  • 4.4. Audit Results

    Male and Female Error Rates

  • NIST- gender classification on female faces was 1.8% to 12.5% lower than male faces; PPB- 8.1% to 20.6%
  • The relatively high positive predictive value for females indicate that when a face is predicted to be female the estimation is more likely to be correct than when a face is predicted to be male.
  • false positive rates (FPR) for males are triple or more than for females
  • Darker and Lighter Error Rates

    All classifiers perform better on lighter subjects than darker subjects in PPB

    Intersectional Error Rates

  • darker females account for the largest proportion of misclassified subjects
  • though darker females make up 21.3% of the PPB benchmark, they constitute between
  • 61.0% to 72.4% of the classification errors
  • The Microsoft gender classifier performs the best, with zero errors on classifying all males and lighter females.
  • On the South African subset of the PPB benchmark, all the error for Microsoft arises from misclassifying images of darker females.
  • Face++ is flawless on lighter males.
  • the presence of more darker individuals is a better explanation for error rates than a deviation in how images of parliamentarians are composed and produced.
  • 4.5. Analysis of Results

  • Classification is 8.1% − 20.6% worse on female than male subjects and 11.8% − 19.2% worse on darker than lighter subjects
  • Darker females have the highest error rates for all gender classifiers ranging from 20.8% − 34.7%.
  • lighter males are the best classified group with 0.0% and 0.3% error rates respectively.
  • 4.6. Accuracy Metrics

    the IBM API is most confident in classifying lighter males and least confident in classifying darker females.

    4.7. Data Quality and Sensors

  • pose, illumination, and expression (PIE) can impact the accuracy of automated facial analysis.
  • Default camera settings are often optimized to expose lighter skin better than darker skin
  • Images from Rwanda and Senegal had more pose and illumination variation than images from other countries
  • classification accuracy does not appear to be confounded by the quality of sensor readings.
  • error rates would be higher on more challenging unconstrained datasets.