Fact check
AI analysis“The company has not stated whether the uploaded images were AI-generated.”
Reasoning
Early blog posts and reports indicated OpenAI had not disclosed the provenance of uploaded images, supporting the claim. However, later communications—including an OpenAI blog confirming AI‑generated images in the dataset and a Reuters report of an SEC filing revealing 12% of images were AI‑generated—demonstrate that the company has indeed stated the presence of AI‑generated images. The existence of these explicit disclosures outweighs the earlier lack of information, rendering the claim false.
On confidence: Multiple recent, credible sources directly state that OpenAI disclosed AI‑generated images in its training set, contradicting the claim that it has not made any such statement.
Important context
The claim does not specify a time frame. While OpenAI initially did not disclose image provenance, it later publicly acknowledged AI‑generated content in its training data, which directly refutes the claim.
Evidence
Supporting (3)
- Tier 1 — Primary sourcePausing training of our most capable models
We are pausing further training while we investigate data provenance. We have not disclosed whether any of the images in our dataset were generated by AI.
- Tier 1 — Primary sourceOpenAI Statement on Data Sources
OpenAI has not released details about the origin of the image data used in model training.
- Tier 2 — Independent reportingindependent originOpenAI hasn't clarified if images used for training were AI-generated
The company has not said whether the images uploaded to its training pipeline were created by humans or by other AI systems.
Contradicting (2)
- Tier 1 — Primary sourceOur training data includes AI‑generated images
In this update we confirm that a portion of the image dataset consists of images generated by earlier versions of our models.
- Tier 2 — Independent reportingindependent originOpenAI discloses AI‑generated images in training set
OpenAI filed a supplemental filing with the SEC stating that 12% of the images used in training were generated by its own DALL·E system.
Limitations
The assessment assumes the cited sources are accurate and up‑to‑date. If the company’s statements were limited to aggregate statistics rather than specific uploaded images, nuance may exist, but the overall disclosure still contradicts the claim.
- Last verified:
- Sep 26, 2026, 5:08 PM CDT
- Pipeline:
- 0.1.0
- Claim type:
- Factual
Where this claim appeared
OpenAI pauses training of its ‘most capable models’The Verge