AlignmentdeceptionlowMEDIUM-HIGH 0.7008 Sept 2026
Tracing Stereotypes from Representation to Output in Multilingual LLMs
Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing,…