On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
ABSTRACT
How big is too big? What are the possible risks associated with this technology and what paths are available for mitigating those risks?
1. INTRODUCTION
LMs got bigger. environmental impact scales with model size In collecting ever larger datasets we risk incurring documentation debt(?) LMs are not performing natural language understanding (NLU), Focusing on results without understanding of the mechanism can cause misleading results human interlocutors to impute meaning where there is none
4. UNFATHOMABLE TRAINING DATA
the training data has been shown to have problematic characteristics resulting in models that encode stereotypical and derogatory associations along gender, race, ethnicity, and disability status
4.1 Size Doesn’t Guarantee Diversity
the voices of people most likely to hew to a hegemonic viewpoint are also more likely to be retained. Internet data overrepresents younger users and those from developed countries multiple cases where people on the receiving end of death threats on Twitter have had their accounts suspended while the accounts issuing the death threats persist. a limited set of subpopulations can continue to easily add data, sharing their thoughts and developing platforms that are inclusive of their worldviews; if populations who feel unwelcome on mainstream sites set up different fora for communication, these may be less likely to be included in training data Finally, the current practice of filtering datasets can further attenuate the voices of people from marginalized identities. 4.2 Static Data/Changing Social Views
Social movements produce new norms, language, and ways of communicating. LMs can't keep up social movements which are poorly documented and which do not receive significant media attention will not be captured at all. media outlets that tend to ignore peaceful protest activity and instead focus on dramatic or violent events Due to the cost, LLMs won't be kept up-to-date 4.3 Encoding Bias
LLMs exhibit various kinds of bias BERT associates phrases referencing persons with disabilities with more negative sentiment words, gun violence, homelessness, and drug addiction are overrepresented in texts discussing mental illness 4.4 Curation, Documentation & Accountability
LMs trained on large, uncurated, static datasets from the Web encode hegemonic views that are harmful to marginalized populations. Feeding AI systems on the world’s beauty, ugliness, and cruelty, but expecting it to reflect only the beauty is a fantasy. undocumented training data perpetuates harm without recourse.
5. DOWN THE GARDEN PATH
we explore some of the risks and harms that can follow from deploying technology that has learned those biases we focus on the risk that misdirected research, specifically around the application of LMs to tasks intended to test for NLU This research brings with it an opportunity cost, on the one hand in terms of time not spent applying meaning capturing approaches to meaning sensitive tasks, and on the other hand in terms of time not spent exploring more effective ways of building technology with datasets of a size that can be carefully curated and available for a broader set of languages (?) no actual language understanding is taking place in LM-driven approaches to these tasks training data for LMs is only form; they do not have access to meaning.
6. STOCHASTIC PARROTS
LLMs present real-world risks of harm\6.1 Coherence in the Eye of the Beholder
human language use takes place between individuals who share common ground and are mutually aware of that sharing, who have intention, and who model each others’ mental states. Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. How do LLMs understand sarcasm? an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.6.2 Risks and Harms
As people in positions of privilege with respect to a society’s racism, misogyny, ableism, etc., tend to be overrepresented in training data for LMs While some of the most overtly derogatory words could be filtered out, not all forms of online abuse are easily detectable using such taboo words, propagating or proliferating overtly abusive views and associations, GPT-3 could be used to generate text in the persona of a conspiracy theorist, which in turn could be used to populate extremist recruitment message boards. A case in point is the story of a Palestinian man, arrested by Israeli police, after MT translated his Facebook post which said “good morning” (in Arabic) to “hurt them” (in English) and “attack them” (in Hebrew). 6.3 Summary
7. PATHS FORWARD
we urge researchers to shift to a mindset of careful planning, along many dimensions, before starting to build either datasets or systems trained on datasets. Value sensitive design provides a range of methodologies for identifying stakeholders working with them to identify their values, and designing systems that support those values we would like to consider use cases of large LMs that have specifically served marginalized populations