Rendered at 22:48:50 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Aurornis 38 minutes ago [-]
The article is vague about the companies in the paper, for some reason.
In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one.
Companies like Google aren't included because they aren't unicorn startups.
dominotw 34 minutes ago [-]
databricks is a top ai startup?. i thought they did spark hosting or something.
what makes them a top ai startup
tfrancisl 20 minutes ago [-]
$160 billion valuation and market capture among some top (non-AI, traditional enterprise) companies, I think.
8 minutes ago [-]
egonschiele 48 minutes ago [-]
As far as I can tell, the paper never actually mentions the companies who aren't publishing papers. Open AI, Anthropic, and hugging face are all specifically mentioned as companies that do publish papers. Just FYI for anyone else who reads "AI's top startups" and immediately assumes OpenAI and Anthropic.
TimCTRL 1 hours ago [-]
Yet none of them would have been here if Google hadn't published "Attention is all you need", the irony.
arjie 46 minutes ago [-]
The authors of that paper are all at AI companies not named Google.
gowld 35 minutes ago [-]
They wouldn't be if Google hadn't published "Attention is all you need".
usef- 6 minutes ago [-]
Lab employees do seem to be "knowledge sharing" as they move employers as far as I can tell, so I don't think that's true,
47 minutes ago [-]
randomImmigrant 40 minutes ago [-]
What the blogificafion of AI research has done is allowed all kinds of claims and terminology related to AI to be introduced and taken up in a manner replicating social media dynamics. And that is simply not healthy. We’re fast reaching a place where any claim can be backed up with a set of numbers from a number of experiments run in some gamified environment or the other, with little concern for if it all adds up to anything.
It’s a vicious loop, because this same junk then goes in to train the next models which help spit out the next set of models AND blogs/papers.
The net effect is not dissimilar to setting termites loose in a library.
gowld 39 minutes ago [-]
How is that different from traditional research publication? It's faster?
randomImmigrant 4 minutes ago [-]
Top line difference is speed. The impact of speed is deeper and gets felt over time.
Traditional research publications are no angels. They gatekeep research, and also allow financial incentives to drive them to publish junk with their stamp on it.
But a flood of papers doesn’t actually mean more knowledge. In bypassing this route entirely, AI has swiftly lost the ability to engage with itself as a field. And the cost of that is only beginning to be felt.
shimman 23 minutes ago [-]
No traditional research is mostly done by post-docs and have a phd level of education rather than a tech bro that passed leetcode. That seems like a good bar to have; not too mention the whole peer review thing, hard to really understand anything if you purposely withhold it and tell people to kick rocks.
janalsncm 3 minutes ago [-]
Two separate axes here: whether the research is in a blog, and whether it’s peer reviewed.
On the first, I mostly don’t care.
On the second, that’s mostly unavoidable if they want to keep their IP.
greazy 8 minutes ago [-]
Plenty of research is performed by non PhDs.
The main difference is reproducibility and peer review process.
sndgndgndgndy 5 minutes ago [-]
Science is a method of inquiry, not an institution or group of people with titles.
KoolKat23 60 minutes ago [-]
Perhaps I'm imagining it but the entire industry was build on published research, this "AI wave" is at odds with that and seems to be driven by greed (although they'll claim some arms race or something to help themselves sleep at night).
There should be a new ESG (Environmental, Social, and Governance) policy being pushed recognizing the important role this plays. Although ESG and all norms have been set aside in this grim new world it seems.
fultonn 13 minutes ago [-]
> Perhaps I'm imagining it
You are not. A disproportionate amount of value in the computing industry was created by geniuses who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.
This observation pre-dates the current wave of AI hype by a half century or so.
> driven by greed
I can only speak for myself.
For me it's exactly the opposite. If I want to explain how something works, I can just... do that. If I want to share an artifact demonstrating how to solve a particular type of problem, I can just... do that. If I want to mentor/teach, I can just... do that.
Doing those things within the confines of Academia Approved Institutions is exhausting and distracting.
To wit, and the actual point of this post: the term "Publishing Research" in this article doesn't mean "post it on a blog and share the source code". It means engaging in a very specific and peculiar and extremely political modality of communication.
And it really only makes sense to do that specific and peculiar and political thing you're at a stage in your professional/personal development where you need to play that particular prestige game. (Which there's nothing wrong with, but it is a deeply cargo culted version of the actual scientific process.)
Aurornis 30 minutes ago [-]
> Perhaps I'm imagining it but the entire industry was build on public research
There is some excellent publicly-funded research in there, but pivotal papers like Attention Is All You Need and the numerous pivotal OpenAI publications were privately funded.
OpenAI is at the top of the chart in the study.
I think you're bringing some assumptions into this conversation that aren't supported by the evidence.
anon373839 15 minutes ago [-]
> OpenAI is at the top of the chart in the study.
A distinction should be made between the old nonprofit OpenAI and the current organization. They don't publish technical research anymore.
KoolKat23 27 minutes ago [-]
Sorry I mean published i.e. available to the public.
Aurornis 25 minutes ago [-]
I was also referring to published papers available to the public.
Like I said, I think you're bringing some assumptions to this conversation that aren't based on the how the industry came about.
23 minutes ago [-]
cute_boi 54 minutes ago [-]
Forget publishing, companies like misanthropic are buying rare books and destroying it.
When they do this, are they buying a single copy of say Book A and destroying a single Book A copy? Or are they buying every copy they can find of Book A and destroying all of them?
ComputerPerson 47 minutes ago [-]
Elon recently (2d) tweeted about making sure they preserve rare books and "scan them the hard way". That tweet launched a cultural wave of opposition regarding the debinding of rare books.
The discussion has been around, but it's flared significantly recently. Not sure if that's what the person you're responding to is specifically inflamed about.
Regardless, old/rare books are certainly being acquired and destroyed.
gowld 36 minutes ago [-]
No, he didn't tweeted about "making sure" of anything.
Elon tweeted "I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
which is no sense a credible source for what he actually asked SpaceXAI team to do, or if they are doing it, or what they were doing before yesterday. Elon has an extremely long and thorough track record of lying.
The problem remains either way. Small operations around the globe are debinding rare/unique books. I'm involved with an organization considering (not strongly) that exactly. It's bigger than a single man's tweets.
Internet Archive scans their books one page at a time keeping the book intact and stored in a warehouse.
usef- 1 minutes ago [-]
They do in this case: the judge said that destroying them means that the text is being "transferred" rather than "copied". So the law wants them to destroy.
paxys 43 minutes ago [-]
I’m not sure why this is so surprising? AI company does not automatically mean research company. The vast majority of new startups popping up over the last few years have commercial motivations, and use models built by someone else. Why are you expecting them to publish scientific papers?
50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
odyssey7 46 minutes ago [-]
Once something goes commercial, academics have to consider what progress would be research-worthy, rather than a half-baked prototype for a product or feature.
Companies can’t be expected to publish their confidential and proprietary information about their feature development, and academics should consider projects that would have higher impact.
If you’re taking up a seat in a PhD program tinkering with would-be feature ideas for an existing major tech company, you should really just get hired by that tech company, where the resources are abundant and the degree is not required.
TrackerFF 28 minutes ago [-]
I don't expect non-foundational or non-frontier startups to do, or publish much heavy hitting research. "AI" has become such a ubiquitous description, it is probably easier to find non-AI startups.
raincole 46 minutes ago [-]
Do startups in other fields constantly publish their research?
aeternum 38 minutes ago [-]
The age of the dark forest is upon us
deadbabe 21 minutes ago [-]
There is no reason to publish anything in the AI era until you have reaped as much benefit from it as you are satisfied with. "Building in public" and "Researching in public" now means someone can swoop in with an AI and copy you instantly and become your competition overnight. Keep secrets.
Even at work, I have come to realize if I simply horde my accumulated custom AI built tools and productivity boosting tricks for myself, I can make myself more competitive as an employee.
I think this is how you get hired now, not by having a good resume, but by making claims of having special processes and personal tooling design that gets massive productivity ROI.
annoyingnoob 28 minutes ago [-]
Startups exist to profit, not necessarily publish.
amazingamazing 41 minutes ago [-]
Trade secrets are back, baby!
More to the point, without determining how much work is “worthy” of a paper it is unclear how much this matters.
Most AI companies are either a product and marketing layer over a model or not meaningfully moving any dimension to be worthy of a paper.
Also, formal papers and blogs and “cards” are all being intertwined.
ltbarcly3 44 minutes ago [-]
All startups barely publish their research, because doing that would be incredibly stupid for the most part, because they want to sell the stuff they invent not give it away for free.
gowld 34 minutes ago [-]
The stuff they invent isn't what's in the papers. The progress in AI is in the engineering hurdles.
ltbarcly3 7 minutes ago [-]
What you just said is not in any relationship to reality.
sublinear 59 minutes ago [-]
> “If we were racing forward on cancer-curing AI, I would be like, ’Fantastic, full steam ahead,’” she says. “But that’s not what we’re racing toward, right?”
We're not racing towards anything. We've been going in circles for years.
datakan 54 minutes ago [-]
Feels more like 3 steps forward and 2 steps back.
warkdarrior 55 minutes ago [-]
> We're not racing towards anything. We've been going in circles for years.
We reached the NASCAR-racing equivalent of scientific research.
xyst 46 minutes ago [-]
That’s because most of its slop, and not reproducible
Revanche1367 39 minutes ago [-]
Live by the non-determinism, die by the non-determinism.
kekku 45 minutes ago [-]
[flagged]
EGreg 41 minutes ago [-]
Funny, as ONE PERSON with Claude, I've been on a tear when it comes to publishing my own research:
In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one.
Companies like Google aren't included because they aren't unicorn startups.
what makes them a top ai startup
Traditional research publications are no angels. They gatekeep research, and also allow financial incentives to drive them to publish junk with their stamp on it.
But a flood of papers doesn’t actually mean more knowledge. In bypassing this route entirely, AI has swiftly lost the ability to engage with itself as a field. And the cost of that is only beginning to be felt.
On the first, I mostly don’t care.
On the second, that’s mostly unavoidable if they want to keep their IP.
The main difference is reproducibility and peer review process.
There should be a new ESG (Environmental, Social, and Governance) policy being pushed recognizing the important role this plays. Although ESG and all norms have been set aside in this grim new world it seems.
You are not. A disproportionate amount of value in the computing industry was created by geniuses who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.
This observation pre-dates the current wave of AI hype by a half century or so.
> driven by greed
I can only speak for myself.
For me it's exactly the opposite. If I want to explain how something works, I can just... do that. If I want to share an artifact demonstrating how to solve a particular type of problem, I can just... do that. If I want to mentor/teach, I can just... do that.
Doing those things within the confines of Academia Approved Institutions is exhausting and distracting.
To wit, and the actual point of this post: the term "Publishing Research" in this article doesn't mean "post it on a blog and share the source code". It means engaging in a very specific and peculiar and extremely political modality of communication.
And it really only makes sense to do that specific and peculiar and political thing you're at a stage in your professional/personal development where you need to play that particular prestige game. (Which there's nothing wrong with, but it is a deeply cargo culted version of the actual scientific process.)
There is some excellent publicly-funded research in there, but pivotal papers like Attention Is All You Need and the numerous pivotal OpenAI publications were privately funded.
OpenAI is at the top of the chart in the study.
I think you're bringing some assumptions into this conversation that aren't supported by the evidence.
A distinction should be made between the old nonprofit OpenAI and the current organization. They don't publish technical research anymore.
Like I said, I think you're bringing some assumptions to this conversation that aren't based on the how the industry came about.
[1] https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
The discussion has been around, but it's flared significantly recently. Not sure if that's what the person you're responding to is specifically inflamed about.
Regardless, old/rare books are certainly being acquired and destroyed.
Elon tweeted "I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
which is no sense a credible source for what he actually asked SpaceXAI team to do, or if they are doing it, or what they were doing before yesterday. Elon has an extremely long and thorough track record of lying.
https://xcancel.com/elonmusk/status/2081844165881594362
The problem remains either way. Small operations around the globe are debinding rare/unique books. I'm involved with an organization considering (not strongly) that exactly. It's bigger than a single man's tweets.
The scanning process destroys the book.
Internet Archive scans their books one page at a time keeping the book intact and stored in a warehouse.
50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
Companies can’t be expected to publish their confidential and proprietary information about their feature development, and academics should consider projects that would have higher impact.
If you’re taking up a seat in a PhD program tinkering with would-be feature ideas for an existing major tech company, you should really just get hired by that tech company, where the resources are abundant and the degree is not required.
Even at work, I have come to realize if I simply horde my accumulated custom AI built tools and productivity boosting tricks for myself, I can make myself more competitive as an employee.
I think this is how you get hired now, not by having a good resume, but by making claims of having special processes and personal tooling design that gets massive productivity ROI.
More to the point, without determining how much work is “worthy” of a paper it is unclear how much this matters.
Most AI companies are either a product and marketing layer over a model or not meaningfully moving any dimension to be worthy of a paper.
Also, formal papers and blogs and “cards” are all being intertwined.
We're not racing towards anything. We've been going in circles for years.
We reached the NASCAR-racing equivalent of scientific research.
https://arxiv.org/search/?searchtype=author&query=Magarshak%...
So I know that smart people in those companies can definitely publish. In fact, a whole team should probably be publishing like no tomorrow!
To be fair — from about half of the papers.