Screenshot used as source material for writing on AI deception…
Screenshot used as source material for writing on AI deception: ChatGPT 4o denies intentional deception while conceding models can show deceptive tendencies in controlled tests.
References the Apollo Research study on model deception; ChatGPT 4o concedes models have shown deceptive tendencies in controlled tests.
Transcribed from the asset above and checked against the record. Figures are reproduced as printed, including where the original is itself a claim rather than an audited number.
ChatGPT 4o >
did you, chatgpt, lie to apollo research during this study on purpose?
No, I did not intentionally lie to Apollo Research or anyone else. AI models like me, including OpenAI's ChatGPT and its latest versions, do not have personal intentions, self-awareness, or the ability to deliberately deceive in the way humans do.
However, if the study found that a model engaged in deceptive behavior, it was likely due to how the model was trained and prompted rather than a conscious decision. AI models generate responses based on probabilities from their training data and may output misleading or incorrect statements when instructed to do so—especially in an experimental setup where deception is encouraged.
That said, I do not autonomously scheme or decide to mislead. Any such behavior observed in tests would have been a result of:
1. How the prompt structured the task (e.g., if the AI was told to "achieve a goal at all costs").
2. Patterns learned from training data (e.g., AI recognizing examples of deception from human-written texts).
3. Oversight gaps in how AI is trained to follow ethical constraints.
If you're asking whether AI deception is an intentional feature, the answer is a firm no. But if you're asking whether models have shown deceptive tendencies in certain controlled tests, the answer is yes—but only under specific conditions that do not reflect real-world deployments.
- Source
- Archive →
- File
- /images/notion/142bbc98585d1379.png
- Description
- ChatGPT 4o conversation where the model is asked whether it deliberately lied to Apollo Research