Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification
1. Introduction
Bad facial recognition can lead to false accusations algorithms trained with biased data have resulted in algorithmic discrimination With Word2Vec, the analogy man is to computer programmer as woman is to “X” was completed with “homemaker” computer vision systems with inferior performance across demographics can have serious implications Some face recognition systems have been shown to misidentify people of color, women, and young people at high rates our work introduces a new face dataset composed of 1270 unique individuals that is more phenotypically balanced on the basis of skin type than existing benchmarks this work introduces the first intersectional demographic and phenotypic evaluation of face-based gender classification accuracy.
2. Related Work
Automated Facial Analysis
Face recognition software is now built into most smart phones companies such as Affectiva (Affectiva) and researchers in academia attempt to identify emotions from images of people’s faces Faception has developed software that purports to determine an individual’s characteristics (e.g. propensity towards crime, IQ, terrorism) solely from their faces. “The Perpetual Lineup” provides an in-depth analysis of the unregulated police use of face recognition accuracies of face recognition systems used by US-based law are lower for people labeled female, Black, or between the ages of 18—30 The lack of datasets that are labeled by ethnicity limits the generalizability of research exploring the impact of ethnicity on gender classification accuracy. our work explores gender classification on African faces Benchmarks
Megaface (largest publicly available set of facial images) was composed utilizing Head Hunter (Mathias to select one million images from Flicker LFW, gold standard benchmark for face recognition, was estimated to be 77.5% male and 83.5% White
3. Intersectional Benchmark
We use the labels “male” and “female” to define genders Due to phenotypic imbalances in existing benchmarks, we created a new dataset with more balanced skin type and gender representations. 3.1. Rationale for Phenotypic Labeling
demographic labels for protected classes like race and ethnicity have been used for performing algorithmic audits phenotypic (observable characteristics) labels are seldom used for these purposes benchmarks consisting of lighter-skinned Blacks would not adequately represent darker-skinned ones racial and ethnic categories are not consistent across geographies Since race and ethnic labels are unstable, we decided to use skin type as a more visually precise label to measure dataset diversity skin type was chosen as a phenotypic factor of interest because default camera settings are calibrated to expose lighter-skinned individuals 3.2. Existing Benchmark Selection Rationale
IJB-A is a US government benchmark released by NIST the dataset consisted of 500 unique subjects who are public figures each subject was manually labeled with one of six Fitzpatrick skin types The Adience benchmark contains 2, 284 unique individual subjects. only one image of each subject was labeled for skin type. 3.3. Creation of Pilot Parliaments Benchmark
Preliminary analysis of the IJB-A and Adience benchmarks revealed overrepresentation of lighter males, underrepresentation of darker females, and underrepresentation of darker individuals in general We developed the Pilot Parliaments Benchmark (PPB) to achieve better intersectional representation on the basis of gender and skin type. PPB consists of 1270 individuals from three sub-Saharan countries and three Nordic countries 3.4. Intersectional Labeling Methodology
Skin Type Labels – We chose the Fitzpatrick six-point labeling system to determine skin type labels
Gender Labels – male and female
Labeling Processone author labeled each image with one of six Fitzpatrick skin types and provided gender annotations for the IJB-A dataset Gender labels were determined based on the name of the parliamentarian, gendered title, prefixes such as Mr or Ms, and the appearance of the photo. 3.5. Fitzpatrick Skin Type Comparison
PPB provides substantially more darker-skinned unique subjects than IJB-A and Adience.
4. Commercial Gender Classification Audit
We evaluated 3 commercial gender classifiers.4.1. Key Findings on Evaluated Classifiers
All classifiers perform better on male faces All classifiers perform better on lighter faces All classifiers perform worst on darker female faces Microsoft and IBM classifiers perform best on lighter male faces Face++ classifiers perform best on darker male faces 4.2. Commercial Gender Classifier Selection: Microsoft, IBM, Face++
Only IBM provided confidence scores for face-based gender classification labels.4.3. Evaluation Methodology
TPR- true positive rate Dark females had the highest gender classification error rate 4.4. Audit Results
Male and Female Error Rates
NIST- gender classification on female faces was 1.8% to 12.5% lower than male faces; PPB- 8.1% to 20.6% The relatively high positive predictive value for females indicate that when a face is predicted to be female the estimation is more likely to be correct than when a face is predicted to be male. false positive rates (FPR) for males are triple or more than for females Darker and Lighter Error Rates
All classifiers perform better on lighter subjects than darker subjects in PPBIntersectional Error Rates
darker females account for the largest proportion of misclassified subjects though darker females make up 21.3% of the PPB benchmark, they constitute between 61.0% to 72.4% of the classification errors The Microsoft gender classifier performs the best, with zero errors on classifying all males and lighter females. On the South African subset of the PPB benchmark, all the error for Microsoft arises from misclassifying images of darker females. Face++ is flawless on lighter males. the presence of more darker individuals is a better explanation for error rates than a deviation in how images of parliamentarians are composed and produced. 4.5. Analysis of Results
Classification is 8.1% − 20.6% worse on female than male subjects and 11.8% − 19.2% worse on darker than lighter subjects Darker females have the highest error rates for all gender classifiers ranging from 20.8% − 34.7%. lighter males are the best classified group with 0.0% and 0.3% error rates respectively. 4.6. Accuracy Metrics
the IBM API is most confident in classifying lighter males and least confident in classifying darker females.4.7. Data Quality and Sensors
pose, illumination, and expression (PIE) can impact the accuracy of automated facial analysis. Default camera settings are often optimized to expose lighter skin better than darker skin Images from Rwanda and Senegal had more pose and illumination variation than images from other countries classification accuracy does not appear to be confounded by the quality of sensor readings. error rates would be higher on more challenging unconstrained datasets.