Another researcher is challenging OpenAI about the data driving its increasingly impressive array of mathematical discoveries. Just days after a bitter row erupted over whether the company’s models benefited from unpublished work, a second mathematician has come forward accusing the AI giant of unethical and “dishonest” behavior and a lack of transparency about the origins of its training data.
In a series of posts on Mastodon, mathematician Andreas Thom raised concerns that interactions he and his colleagues had had with the ChatGPT chatbot before OpenAI’s triumphant announcement may have contributed to its success in the field. One of the 10 results OpenAI announced with great fanfare last month involved Thom’s area of expertise, so-called non-sofic groups, and OpenAI acknowledged that their result built heavily on previous work by Thom and fellow mathematician Gábor Kun.
Thom said he began reflecting on his own interactions with OpenAI after Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether the company’s AI models had benefited from his use of OpenAI’s Codex. After OpenAI announced its non-sofic groups result, it was widely criticized in mathematical circles for failing to acknowledge recent contributions from Thom and Kun and the company quietly amended its writeup. Non-sofic groups are, roughly speaking, infinite mathematical structures that cannot be approximated by finite ones.
Thom said he was also struck by “OpenAI’s detailed command of our techniques,” which he said were neither the most obvious nor the most promising routes to a solution at the time. He said he wrote emails to OpenAI researchers Sébastien Bubeck and Mark Sellke, also a statistician at Harvard, to ask whether his interactions with ChatGPT were “part of the training data or accessible to the reasoning process” and could therefore have contributed to the result.
But the answer did not satisfy Thom, who said it only addressed whether his conversations with the chatbot could be accessed directly, not whether they had entered into the vast pools of training data the company uses to improve its models. “No such qualification, explanation, or evidence was given,” he wrote. “I take this as dishonesty to say the least.”
Thom said researchers aren’t equipped to reverse-engineer OpenAI’s training pipeline to figure out whether their work has been used or not. “Only OpenAI has the relevant data for that.” If the company is going to deny doing this, he said the responsibility is on them to prove that by disclosing all necessary datasets and clarifying various settings and terms setting out how it uses data.
OpenAI’s reluctance to conclusively rule out any use of user data echoes the way it defended its recent Millennium Prize breakthrough, both in its public messaging and its communications with Buckmaster — who was working on the problems with Anthropic researcher Levent Alpöge in a personal capacity. In the blog post announcing the Navier-Stokes solution, which concerns the movement of fluids, OpenAI flatly denied using any specific user data: “We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.”
But it would not conclusively rule out an indirect influence: “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Thom said it is the same obfuscatory distinction the company drew in its communications with him. “De-identification may remove a name; it does not remove the intellectual content of a mathematical idea,” he said.
In light of recent events, Thom said “Sellke’s categorical answer was, at minimum, unjustifiably broad and materially misleading; looking back it was plainly dishonest.”
Thom said it “would be ethically indefensible” if nonpublic research supplied by users helped to improve models that the company then used to race those very same users to publication, without consent, proper disclosure, or credit.
OpenAI did not immediately respond to The Verge’s request for comment.
His comments add to mounting unease over OpenAI in mathematical circles at what should be a moment of triumph for the company. Its announced solution to one of mathematics’ legendary Millennium Prize problems is an extraordinary achievement that, should it be verified, few would deny. But this was complicated by the unusual circumstances OpenAI said led it to pursue the problem in the first place: It heard rumors online that other researchers had made major progress and thought it would try too.
The ongoing incident has left a sour taste in mathematicians’ mouths. Numerous researchers told The Verge they worry behavior like this will push the field into a more secretive state if mathematicians know that even rumors they are close to a big breakthrough could ignite a race with a well-resourced tech giant eager for glory.
