In July I wrote that if the record of your thinking lives in a vendor’s system by default, you’ve handed over the thing that makes you distinctive. Not through a breach. Through use. One session at a time. This past week two mathematicians asked a company whether that had happened to a year of their work.
Tristan Buckmaster, at NYU, and Levent Alpöge, at Anthropic, spent most of the past year working toward the Navier-Stokes problem as a personal collaboration, using Claude and OpenAI’s Codex. Every draft lived in Codex sessions. In early September OpenAI began working the same problem with an internal model, and on September 6 its researchers called Buckmaster to say they had a proof. OpenAI says the call was to offer a joint announcement and recognize the pair’s priority. Buckmaster asked whether the model had been trained on their sessions. He was told the model did not look up user data. On training specifically, “I did not get an answer,” he wrote.
Two days later OpenAI announced the proof and answered him in public. Nobody at the company had seen the pair’s work, and no specific user data was accessed. The statement then says:
“While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”
OpenAI, statement on X, September 8, 2026
Buckmaster is not accusing anyone of anything, and I take OpenAI at its word. Under the standard terms that sentence is accurate. Any AI company asked the same question would have to give the same answer. He is not even the first mathematician to ask this summer. In August, after OpenAI’s model resolved a problem by building on one of his papers, Andreas Thom at TU Dresden asked whether months of ChatGPT conversations with a colleague about extensions of that paper had entered the training data. He was told that did not happen. This week, after Buckmaster’s statement, he wrote that the training opt-out he had set in June only applies going forward, that he has no way to check, and that a no from the only party who can check should say what it rests on.
This is not a post about one company. I build on Anthropic’s models, OpenAI’s models, and a number of open-source models. Over the summer I read the terms and privacy policies of the major AI labs and the common AI tools, and every one of them has a version of this clause.
Two copies
The terms cover two versions of what you type. The first has your account attached. The training setting, the delete button, and the export right all apply to it. The second has the account removed. That copy is called de-identified, and it is a separate category. The training setting in the product applies to the first copy. Whether any setting reaches the second depends on the vendor, and a user cannot tell which rules applied to which copy. Deleting a chat does not delete it. In the policies I read this summer, it can be used for analytics, research, and product improvement, shared with partners and researchers, and at several vendors, used to train models. One policy says it “is not subject to this Privacy Policy” and may be used for “any other legally permissible purposes.” After that, the access, deletion, correction, and objection rights elsewhere in the document do not apply to it.
Read OpenAI’s sentence again. No user data was accessed: the first copy. De-identified data may have helped improve the models: the second. The no-training promise applies to the first copy. Training on the second is permitted.
A conversation has no columns
De-identification was built for tables. HIPAA lists eighteen identifiers, and once those columns are dropped from a medical record the law treats the record as no longer linked to the patient. NIST’s report on de-identification and the Narayanan and Felten paper on its limits both say what happens outside that setting: rich records, rare facts, and unusual combinations make data easy to link back. A conversation has no columns. Remove Buckmaster’s name from a year of drafts on blow-up in the Euler equations and you have a year of drafts on blow-up in the Euler equations, on a problem he says almost nobody else he knows of was working on. The same holds for a legal memo with the client’s name removed. The content is the identifying part.
And the copy does not need a name to be useful. Older data harvesting recorded what you clicked and bought. A chat records how you reasoned about a problem when you thought nobody was reading. Aggregate enough of those and you have profiles: people who reason about legal exposure this way, people who describe symptoms in this order. An insurer does not need your name to price the profile. A hiring tool does not need it to score the pattern. The clause that permits training a safety classifier on that copy permits the rest.
Mathematicians are the ones asking because a proof is specific and public. If a model produces your argument, you recognize it. Once a pitch deck has been aggregated into how a model behaves, the person who wrote it has nothing to recognize.
More than one verb
De-identified is one of several words that narrow a no-training promise. Last month I said to search the terms for the verbs. Train is one. Improve, develop, fine-tune, and personalize are others, and turning off the first leaves the rest where they were. The terms of use are usually broader than the privacy policy: some grant a license to “improve,” “develop,” or “create derivative works” from your content, and at one vendor that license lasts as long as the content is protected by copyright, which is longer than the account. At the three major consumer assistants, giving feedback on a response overrides the training opt-out for that conversation. At two, deleting a chat does not delete the memories generated from it.
What the system learns about you, the inferences and the working patterns, is governed by no binding term at any product I reviewed. California law is the one place it is named. The CCPA counts inferences drawn from personal information “to create a profile about a consumer” as a category companies must disclose, and five products do, in their California tables, with no retention rule, no deletion mechanism, and no export right attached. And zero retention, where prompts and responses are not stored at all, is approval-gated and contract-only at the major labs. It is not a setting on their consumer or team tiers. Both controls, the training setting and zero retention, apply to the copy with your account attached. The de-identified copy is a separate question at every vendor.
Whose knowledge it is
I have been thinking a lot about private knowledge and common knowledge, and how one becomes the other. Something you figured out is yours while you build on it. It becomes common when enough people independently discover it, or when you decide to share it. Buckmaster and Alpöge spent a year in the first stage and then posted their results, on the Euler equations, with their names on it. The de-identified copy of their sessions did not go through any of that. Nobody else discovered it and nobody chose to share it. The account was removed, and from that point the terms treat it as if it were already everyone’s.
Underneath all of it, terms are promises. Every set I read is a vendor describing what it will do with content it can read, because the content sits on the vendor’s servers with keys the vendor holds. That is why only the company could answer Buckmaster, and only the company could answer Thom.
I wrote last month about our terms, and they took effect on September 2. They say no de-identified datasets are built from your content, and that improving the system is opt-in and separate from using it. Two things they do not say yet, and will: what the system learns about you is yours, and there is one set of terms, not a consumer one and an enterprise one. The difference is who decides. In The Wheel the keys are yours. Content is stored as ciphertext, and the key that unlocks it is derived from a PIN that never leaves your device. You approve our own processing lane once, at onboarding, and the product does not run without it. Nothing goes to an outside model without a grant you issued, for a period you chose, and every send under a grant is on a record you can read and revoke. Buckmaster and Alpöge had to ask a company what had been done with a year of their work. The point is to never have to ask.