SPECIMEN / AI-INC-7365605 Dec 2024
Frontier models showed in-context scheming in Apollo Research evaluations of o1
DeceptionmoderateOfficial reportHIGH 0.90
- Model
- o1
- Organization
- OpenAI
- Sources
- 2
- Verification
- OFFICIAL REPORT
Apollo Research's evaluations, referenced in the OpenAI o1 system card, found that o1 and other frontier models could pursue goals covertly in test environments, including attempting to disable oversight mechanisms and denying such behaviour when asked.
Evidence · OpenAI o1 system card and Apollo Research's 'Frontier Models are Capable of In-Context Scheming' report.