How AI Systems Work
- Explain how neural networks learn from data
- Describe what LLMs are doing when they generate text
- Evaluate the "Stochastic Parrots" argument with the understanding needed to assess it fairly
- Identify how technical choices in AI development embed philosophical assumptions
1. From Rules to Learning
AI moved from symbolic programs to neural nets
2. Neural Networks and the Basic Unit: Artificial Neurons
An artificial neuron receives numeric inputs, multiplies each input by a weight, sums the results, and passes the result through a simple function.
A network neurons, organized in layers, can compute complex functions. The inputs enter the first layer, get transformed, pass to the next layer, get transformed again, and so on until the output layer produces a result
Training: Adjusting the Weights
The procedure works like this: show the network an input, observe its output, compare the output to the correct answer, measure the error, and then adjust the weights in the direction that would have reduced the error. (Backpropagation)
Model Scale
emergent abilities: capabilities that appear weak or absent in smaller models but become visible in larger ones.
Some researchers argue that these abilities represent genuine qualitative changes with scale; others argue that some apparent “emergence” comes from the way we measure performance.
3. Large Language Models and Next Token Prediction
An LLM is a neural network trained on a specific task: predict the next token.
When you prompt an LLM:
- Your text is converted into a sequence of tokens.
- The model computes a probability distribution over all possible next tokens.
- A token is sampled or selected from that distribution.
- The sampled token is added to the sequence, and step 2 repeats.
- This continues until the model generates a stop token or reaches a length limit.
Plausibility and coherence are properties of the output.
Understanding would be a property of the process that produces the output.
4. The Stochastic Parrots Argument
Syntax is not semantics.
Risks-
- Mistaking fluency for competence.
- Encoding and amplifying bias.
- Research misdirection. too much emphasis on LLMs
5. Philosophical Assumptions in Technical Choices
A model trained predominantly on English-language internet text will have a picture of the world shaped by the perspectives overrepresented there: primarily Western, educated, and English-speaking. A system optimized to perform well on benchmarks will optimize for what benchmarks measure The Alignment Problem
The process of fine-tuning large language models, or training them on human feedback to be helpful, harmless, and honest, embeds assumptions about what "helpful," "harmless," and "honest" mean.Reading: Gebru et al., "On the Dangers of Stochastic Parrots" (2021,select sections)
Does understanding how LLMs generate text change your view of the Searle/Dennett debate? Does it change your view of the Turing Test, or your view on whether an AI has sentience/consciousness?