
Annual Lecture Series – John Morrison - Do neural networks understand what they are doing?
The Center for Philosophy of Science at the University of Pittsburgh invites you to join us for our 67th Annual Lecture Series Talk. Attend in person in room 1008 in the Cathedral of Learning (10th Floor) or visit our live stream on YouTube at https://www.youtube.com/channel/UCrRp47ZMXD7NXO3a9Gyh2sg.
The Annual Lecture Series, the Center’s oldest program, was established in 1960, the year when Adolf Grünbaum founded the Center. Each year the series consists of six to eight lectures, about three quarters of which are given by philosophers, historians, and scientists from other universities.
Annual Leture Series – John Morrison
Friday, November 13th @ 3:30 pm - 5:30pm EDT
1008 Cathedral of Learning
Title: Do neural networks understand what they are doing?
Abstract:
It is often claimed that neural networks construct models of the world (“world models”) and do not merely exploit heuristics. In this sense, they are said to understand their tasks. But it is not always clear what having a world model amounts to, or what should count as evidence for or against. In the first half of the talk, I propose a minimal notion of a world model: a single, unified representation of a task that preserves its structural relations. I explain how this notion sits within a hierarchy of stronger ones, and argue that many disputes about world models are merely verbal. In the second half of the talk, I reassess the standard example of a network using a world model: Othello-GPT, a transformer widely taken to use a model of the Othello board. I argue that the existing evidence is ambiguous once we adopt the right baselines and the right metrics. I then introduce another kind of evidence that favors the heuristic interpretation: how quickly the network learns variants of Othello that extend rather than scramble the board’s geometry. I end by drawing general lessons for attributing world models to neural networks, and by describing similar approaches for attributing algorithms and probabilistic inference.
Can’t make it in-person? This talk will available online through the following:
Zoom: TBA and
YouTube at https://www.youtube.com/channel/UCrRp47ZMXD7NXO3a9Gyh2sg.
A reception with light refreshments will follow in The Center on the 11th floor from 5-6pm.