GDPR gives you the right to ask whether your data trained a model. It gives you no way to check the answer, and no way back.
Two mathematicians have just demonstrated this in public. On 8 September 2026, hours after OpenAI announced a result on the Navier-Stokes equations, Tristan Buckmaster (NYU) said that he and his collaborator Levent Alpöge had put all of their research drafts into Codex, and asked publicly whether the model might have benefited from them. OpenAI denies accessing their user data, while stating it cannot rule out having benefited from de-identified data derived from the use of its products.
For our purposes here, it does not matter who is right. What matters is that the question went unanswered in any verifiable way. The rest of this page explains why that was predictable, article by article.
What happened, in the conditional tense it deserves
OpenAI announced on 8 September 2026 that an internal model, run by a swarm of roughly ten thousand agents, had produced a proof relating to the Navier-Stokes equations, verified in Lean. Tristan Buckmaster questioned the coincidence: the route the model took was, he says, precisely the uncommon approach he had been pursuing with Levent Alpöge for close to a year, and "not the direction one arrives at in a few days by giving a model the problem statement". He reports being told that the model does not look up user session data, with no direct answer on the training question.
A second mathematician, Andreas Thom (TU Dresden), raised a related concern about an earlier result on non-sofic groups: he says he opted out of training on 29 June, which leaves everything before that date outside the opt-out. A separate strand of the dispute concerns authorship, which OpenAI also contests.
We report these as what they are: public allegations, met with denials, which no outsider can settle. That is exactly the problem.
A draft proof is not personal data. The session holding it is.
First distinction, and it catches people out: GDPR does not protect your ideas. It protects information relating to an identified or identifiable natural person (Article 4(1)). A functional inequality is not personal data. The fact that you wrote it, on that date, from that account, in that session, is.
The consequence is practical. For the idea itself your remedies lie elsewhere: copyright in the expression, trade secrets, and, in academia, the conventions of priority and citation, which are not law but often carry more weight. For the link between the idea and you, GDPR applies. The two regimes do not overlap, and invoking the wrong one wastes time.
Four rights, and exactly where each one stops
| What GDPR gives you | Where it stops |
|---|---|
| Purpose limitation (Art. 5(1)(b)): data collected to deliver a service cannot be reused for an incompatible purpose. | "Incompatible" is assessed case by case. A provider whose terms announce training up front will argue the purpose was disclosed from the start. |
| Right of access (Art. 15): you can demand to know what data is processed, and for what purposes. | You will get a description of the processing. You will not get an audit of the training set: nobody outside the company is in a position to look. |
| Right to object (Art. 21): you can object to processing based on legitimate interest. | An objection operates going forward. It has no effect on what has already been ingested. |
| Right to erasure (Art. 17): your data must be deletable. | You delete a record from a database. You do not remove a contribution from the weights of a trained model. At best you retrain, which nobody does for one user. |
All four rights exist, are enforceable, and are exercised successfully every day. None of the four gives back what was learned.
Opting out is not retroactive
This is the costliest point, and the least understood. On most consumer services, using your content to improve models is on by default and turned off in a setting. The day you turn it off, you protect what follows. You recover nothing from before.
In the case above, that is precisely the fault line Andreas Thom points to: an objection raised in late June says nothing about earlier exchanges. Take the general rule away with you: on these services the date that counts is not the day you find the setting, it is the day of your first session.
What the European regulator has already settled
The European Data Protection Board adopted Opinion 28/2024 on AI models on 17 December 2024. Three points frame this dispute exactly:
- A model is not anonymous by construction. It is anonymous only where the likelihood of extracting personal data from it, directly or through queries, is negligible, assessed case by case, against a high bar.
- Legitimate interest remains available as a basis for developing a model, at the price of a documented three-step test: a real and non-speculative interest, necessity, and balancing against the rights of data subjects.
- Unlawfulness upstream contaminates downstream. A model developed on unlawfully processed data may have its subsequent operation called into question, absent demonstrated anonymisation.
So the law sides with the user in principle. It stays silent on proof: the opinion requires the controller to demonstrate compliance, but creates no mechanism letting a third party establish what sits in a training corpus. Hence disputes that end in duelling statements rather than findings.
The only question worth asking a provider
"Are you GDPR compliant?" has never taught anyone anything: everyone says yes, and most are telling the truth. The three questions that actually discriminate:
- Does my content enter a training set: by default, as an option, or never? And if it is an option, is that option on or off when the account is created?
- From what date does my choice apply? This is the question that reveals whether you are being told about the past or only about the future.
- What exactly does a withdrawal trigger: the end of collection, or deletion of what was already collected? Either answer is acceptable. Conflating them is not.
Add the hosting question, which determines the applicable regime and the transfer mechanisms under Chapter V of GDPR: we cover it in our Notus / Plaud Note comparison.
Takeaways
- GDPR protects the link between data and you, not the idea the data contains.
- Access, objection and erasure operate on databases. None of them undoes training.
- An opt-out only works forward: the date that counts is that of your first session.
- EDPB Opinion 28/2024 sets real requirements without creating any external means of verification.
- Faced with an unverifiable question, the only durable answer is architectural: do not hand the content to a system that could learn it.
Sources: Regulation (EU) 2016/679 (GDPR), Articles 4, 5, 15, 17, 21 and Chapter V; EDPB Opinion 28/2024 of 17 December 2024 on AI models; public statements by Tristan Buckmaster (7-8 September 2026) and OpenAI's responses as reported by TechCrunch, VentureBeat, Axios and The Verge (8-10 September 2026). The allegations reported are contested by those named.
Notus makes the opposite commitment: voice processed in France, European models, never listened to by our teams and never used to train our models. The device ships in H1 2027; that commitment goes into the privacy policy at commercial launch, with the retention periods.
Discover Notus →