Archive no. MY-0001
OpenAI AI Text Classifier
OpenAI AI 文本分类器
A free web tool that guessed whether a passage of English text was written by an AI — built for teachers and editors, withdrawn after seven months because it could not be trusted.
Original intent · 初衷
The stated target was not detection for its own sake but a specific harm: automated misinformation campaigns, AI-assisted academic dishonesty, and chatbots presenting themselves as people. The classifier was aimed at English text of at least roughly 1,000 characters, and OpenAI framed it as an imperfect signal for educators weighing whether a text was machine-written. Before releasing it, the company ran an evaluation on a challenge set and published the results on the launch page itself.
Timeline · 时间线
-
2023-01-31 · date-level
OpenAI publishes the AI text classifier for English text, together with its own evaluation: 26% of AI-written text flagged, 9% of human-written text mislabeled.
-
2023-07-20 · date-level
The announcement page is edited to state that the classifier is no longer available because of its low rate of accuracy.
-
2023-07-20 · date-level
On the same page OpenAI says it is researching more effective provenance techniques for text and has committed to mechanisms for AI audio and visual content.
Dates are shown at the precision the sources provide; nothing is rounded up to a full date.
What happened · 后来发生了什么
OpenAI edited the original announcement to record that, as of 20 July 2023, the classifier was no longer available because of its low rate of accuracy. No separate shutdown post was published. The same page states that the company was researching more effective provenance techniques for text and had committed to develop mechanisms that let users understand whether audio or visual content is AI-generated.
Why it stopped · 停止原因
Low rate of accuracy, stated by OpenAI on the announcement page itself: on its own challenge set the classifier caught 26% of AI-written text and mislabeled 9% of human-written text as AI-written.
Curator's reading · 策展人分析
Facts cite a source. Author statements are the makers' own words. Curator analysis and hypotheses are marked as interpretation.
Editor's reading: the most useful part of this entry is not that the tool failed but that the failure was measured and published up front. A 26% detection rate is close to useless for the classroom decision it was meant to support, and once that was public the sensible move was to stop offering the tool rather than to keep tuning it. This is our interpretation; OpenAI's page states the outcome, not our reasoning about it.
What remains · 留下了什么
The original announcement is still online and now carries the withdrawal notice at the top, which makes it the primary record of both the launch and the end.
The earlier GPT-2 output detector, published as part of an open research repository; the AI text classifier itself was a web-only tool with no released weights.
Provenance work, for example C2PA-style content credentials for generated media, is the direction OpenAI pointed to; the museum records it as a successor track, not as a replacement tool.
No longer reachable; recorded for the historical record only.
What the makers learned · 制作者总结
OpenAI states that the classifier is not fully reliable, and that its reliability generally improves with longer input text.
OpenAI states that the classifier is no longer available as of 20 July 2023, and gives its low rate of accuracy as the reason.
OpenAI's own comparison is with its previously released GPT-2 output detector, which it describes as less advanced than the 2023 classifier.
OpenAI recorded the withdrawal inside its launch post instead of publishing a shutdown announcement; the date-level precision here comes from that edit, and the independent report followed five days later.
The launch page published the failure numbers before the shutdown did: 26% true positives and 9% false positives on the company's own challenge set. The tool was withdrawn once those numbers turned out to be the whole story rather than a starting point. Whatever replaced it was not a public classifier but provenance work attached to generated content, which is a different product with a different failure mode: it does not need to guess who wrote something if the generating system records it.
If someone tried again · 重做问题
Detection was framed as a stopgap while provenance was developed, and provenance later arrived for images and audio rather than for text. The open question is whether any standalone text classifier can be good enough to justify its false positives — or whether the honest answer is that authorship of plain text is not reliably recoverable after the fact.
Sources and verification · 来源与核验
- Status evidence
- [openai-classifier-post] [techcrunch-classifier]
- Added
- 2026-09-13
- Entry updated
- 2026-09-13
- Last verified
- 2026-09-13
- Project period
- 2023-01-31 → 2023-07-20
Supports: Both ends of the story, because OpenAI edited the launch post rather than publishing a shutdown notice: the launch, the 26%/9% evaluation on OpenAI's challenge set, the 20 July 2023 withdrawal with "low rate of accuracy" as the stated reason, and the commitment to provenance research.
Supports: Independent reporting that the tool was shut down and that the stated cause was its low rate of accuracy. Published five days after the date OpenAI records on its own page.
Supports: The earlier openly released detector that OpenAI compares its 2023 classifier against. Recorded for the comparison; the repository's license was not verified.