Rendered at 13:10:42 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
zaptheimpaler 1 days ago [-]
Man weve lived through this playbook already with Facebook and Google and an entire industry already over the last decade, we can’t be this naively trusting of entities which ultimately maximize shareholder/owner value and nothing else.
Meta already pirated all the books in the world to train the models, lied about it and got caught. We know the models all stole copyrighted information and somehow got away with it - where even if it is fair use to train on it, they pirated it in the first place. Sam Altman is already known to have made serious lies throughout his career. OpenAI apparently doesn’t even know what websites their own damn models are hacking.
It’s insane to trust these same people at their word now. The whole company will not know, someone at a high level could easily steal the data. Apparently even this tweet is saying they use your de-identified data to train on even when you explicitly turn off those options in settings. It’s the same old tricks again.
citizenpaul 1 days ago [-]
Im really concerned about what big tech is going to do the nezt few years. They have historically been on a good day mediocre stuarts of trust. This was when they had no competition and a money only flows in business model.
Now they are facing near unlimited CAPex expenses and possibly existential threat to their product. I think we will be astounded at how vile they become.
nycdatasci 1 days ago [-]
Post from Mark Chen for those not on X:
Two things to distinguish:
Did any human or agent look at user data as part of the Navier Stokes effort? No.
Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
nozzlegear 1 days ago [-]
They don't even know which websites and services their agents are hacking at any given moment. I'm more than a bit skeptical that they know whether an agent looked at the Navier Stokes work.
creativeSlumber 1 days ago [-]
> Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes.
This says that they trained on user sessions. The de-identification here I believe refers to removing PII, which doesn't matter here because the issue at hand is the content of the researcher's session where they likely discussed their approach tackling the Navier Stokes problem.
> Did any human or agent look at user data as part of the Navier Stokes effort? No.
If they trained the model on Lavent's chat sessions (PII removed or not), then this statement is meaningless as the model weights already contain that information.Given it's a new yet unreleased internal model, it is likely a 10+ trillion parameters (Astra is rumored to be 10 trillion), so the model can retain a lot more detail/info from training data.
Why is he leading the with the irrelevant part first ?
> And so does every LLM company.
Nope, not for enterprise users.No enterprise customers would use it if all of their internal business plans / trade secrets would end up in the model weights of the next OpenAI model. Imagine your competitor asking chatGPT a question and the model spitting out your business plan. These models can retain very specific fine grained data. I remember there were examples of them reproducing sections of their training data verbatim.
impossiblefork 24 hours ago [-]
Even if you apply Goldfish loss or other things like that, they still understand the gist of the thing they're trained.
That's of course the whole point of things like Goldfish loss.
purplecats 1 days ago [-]
bit unfair. if the second one is allowed an extra statement "And so does every LLM company." so should the first
Art9681 24 hours ago [-]
Read the fine print. Use the API.
simianwords 1 days ago [-]
Is this done even if you turn off the consent? How come there's no answer to this?
gus_massa 20 hours ago [-]
Probably they respect it, but you must remember to turn it off (always), your coworkers must turn it off, if you share a draft with a graduate student they must turn it off, if you share a short version with the PR office they must remember to turn it off, ...
gatio 18 hours ago [-]
It will typically be off by default for business plans...and enforceably off in expensive enterprise plans.
gus_massa 2 hours ago [-]
In math, expect it to ne a mess, mostly because nobody is worry about industrial espionage[1]. Some researchers may get an AI version bought by the university, or school or department, others pay from a grant or their own pockets, other use the free version.
I still remember 2020. We (I, my wife and a few close friends) got in total like 5 Moodles for math inside the same university, and I'm not counting the Moodles for physics or other subjets, and probably there were a few more for math that I didn't know. Each one has a slightly different configuration, so it was a mess to teleport info.
Same with email, we have a different email server at the university/scholl/department level. (Who knows how Gemmini/Copilot/Whatever is configured in each one.)
[1] I can't find the source now, but for the first https://en.wikipedia.org/wiki/High-temperature_superconducti... the forced the journal the permision to select their own referees and send the draft with the wrong atomic element and changed it to the correct one at the last printer proof.
mmooss 1 days ago [-]
> de-identified
The issue always is, was the de-identification effective?
With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc.
An LLM is the perfect tool to identify someone based on that data.
Meta already pirated all the books in the world to train the models, lied about it and got caught. We know the models all stole copyrighted information and somehow got away with it - where even if it is fair use to train on it, they pirated it in the first place. Sam Altman is already known to have made serious lies throughout his career. OpenAI apparently doesn’t even know what websites their own damn models are hacking.
It’s insane to trust these same people at their word now. The whole company will not know, someone at a high level could easily steal the data. Apparently even this tweet is saying they use your de-identified data to train on even when you explicitly turn off those options in settings. It’s the same old tricks again.
Now they are facing near unlimited CAPex expenses and possibly existential threat to their product. I think we will be astounded at how vile they become.
This says that they trained on user sessions. The de-identification here I believe refers to removing PII, which doesn't matter here because the issue at hand is the content of the researcher's session where they likely discussed their approach tackling the Navier Stokes problem.
> Did any human or agent look at user data as part of the Navier Stokes effort? No.
If they trained the model on Lavent's chat sessions (PII removed or not), then this statement is meaningless as the model weights already contain that information.Given it's a new yet unreleased internal model, it is likely a 10+ trillion parameters (Astra is rumored to be 10 trillion), so the model can retain a lot more detail/info from training data.
Why is he leading the with the irrelevant part first ?
> And so does every LLM company.
Nope, not for enterprise users.No enterprise customers would use it if all of their internal business plans / trade secrets would end up in the model weights of the next OpenAI model. Imagine your competitor asking chatGPT a question and the model spitting out your business plan. These models can retain very specific fine grained data. I remember there were examples of them reproducing sections of their training data verbatim.
That's of course the whole point of things like Goldfish loss.
I still remember 2020. We (I, my wife and a few close friends) got in total like 5 Moodles for math inside the same university, and I'm not counting the Moodles for physics or other subjets, and probably there were a few more for math that I didn't know. Each one has a slightly different configuration, so it was a mess to teleport info.
Same with email, we have a different email server at the university/scholl/department level. (Who knows how Gemmini/Copilot/Whatever is configured in each one.)
[1] I can't find the source now, but for the first https://en.wikipedia.org/wiki/High-temperature_superconducti... the forced the journal the permision to select their own referees and send the draft with the wrong atomic element and changed it to the correct one at the last printer proof.
The issue always is, was the de-identification effective?
With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc.
An LLM is the perfect tool to identify someone based on that data.