Rendered at 13:10:41 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
kennywinker 1 days ago [-]
Sounds like they did. IMO, building your business on anything but open-weight models is a bad idea. Unless you're on the s&p 500 you are an insect to google, anthropic, grok (ew), and openai - and they could crush you at any time without even noticing.
torvin92 1 days ago [-]
Sunsetting a model with a two-month notice is exactly why the open-weight argument keeps winning. The API is a dependency you don't control.
grahamnorton39 1 days ago [-]
Might be missing something —- are there any issues with Gemini 3.1 Pro that aren’t there in 2.5?
I agree, though. 2.5 Pro is a great model. Very competent, knows a lot, and can process tons of text (and videos, and images, and audio too iirc?). Basically unlimited access to it too via AI Studio. I used it for processing and transforming bucketloads of data, ingesting masses of transcripts and converting them to flashcards, etc. I’ll be sad to see it go. None of the newer, cheaper, but obviously less intelligent benchmaxxed smaller models really seem to hold a candle to it for lots of things.
waldrews 1 days ago [-]
It's still 'preview' and not generally available, so can't run it for US restricted workloads.
dzonga 1 days ago [-]
2.5 flash was also good to use with as the llm layer for voice products.
but google gonna google.
davedx 1 days ago [-]
Nope. The projects I'm on where we use it, we're carefully migrating to the newer models. Where we can we test with evals to try and get an understanding of how the models have changed.
It's not all roses -- I've seen some regressions -- but generally the 3.x Flash models are pretty great for our use cases.
The great thing about LLMs though is it's incredibly easy to diversify and have fallbacks. But of course that means additional costs, mostly centered around engineering efforts to test and integrate them.
ernsheong 1 days ago [-]
Flash is the new Pro, try it first
waldrews 19 hours ago [-]
We sure did. It's a great writer, better in a harness, will process lots large context, but complex reasoning with convoluted rules and low hallucination tolerance? That's still larger model territory.
ernsheong 15 hours ago [-]
There's still 3.1 Pro though, as ancient as it sounds now
waldrews 14 hours ago [-]
Yup. The problem is that it's bizarrely still not in General Availability status.
PaulShin 23 hours ago [-]
Google is falling behind in this competition.
trio8453 5 hours ago [-]
And it's mostly from lacking vision and a clear direction, not from being behind on research or engineering. For example, Gemini 3.8 Flash is really impressive.
But their whole model product line is confusing, even the version numbering barely makes sense. They're barely selling agentic coding or the Gemini chat, the fumbled the coding harness race with the weird Gemini CLI / Antigravity rebrand - the whole thing is a mess.
yieldcrv 1 days ago [-]
Check model garden on vertex ai for other models that you can access
Models you can download and use elsewhere if Google nixes access
OutOfHere 23 hours ago [-]
As an alternative, it's not a bad idea to first convert each document to markdown via a thinking or agentic LLM. Embedded figures can even be embedded as readable tables or Latex or Mermaid. Do record the name and parameters of the model that performs the conversion. You can then query the markdown using any model with an input token cost that is exactly equal to the encoding of the markdown. For multiple queries you can also use input caching.
waldrews 19 hours ago [-]
Yup, tried all variations of that. There's the advantage that you can use a lower hallucination OCR specific model for the pre-processing, at least for clean text. But for something hard like handwritten forms, applying VLM with context is less error prone than preprocessing to text.
Also - and this is bizarre - the token cost of doing that is higher, not lower, at least in Gemini world, and by a large margin. That's very counterintuitive, but a page encoded as image tokens can be smaller than same page as text, and is not meaningfully lossy on documents that are just typed text because the models are well trained on those.
I agree, though. 2.5 Pro is a great model. Very competent, knows a lot, and can process tons of text (and videos, and images, and audio too iirc?). Basically unlimited access to it too via AI Studio. I used it for processing and transforming bucketloads of data, ingesting masses of transcripts and converting them to flashcards, etc. I’ll be sad to see it go. None of the newer, cheaper, but obviously less intelligent benchmaxxed smaller models really seem to hold a candle to it for lots of things.
but google gonna google.
It's not all roses -- I've seen some regressions -- but generally the 3.x Flash models are pretty great for our use cases.
The great thing about LLMs though is it's incredibly easy to diversify and have fallbacks. But of course that means additional costs, mostly centered around engineering efforts to test and integrate them.
But their whole model product line is confusing, even the version numbering barely makes sense. They're barely selling agentic coding or the Gemini chat, the fumbled the coding harness race with the weird Gemini CLI / Antigravity rebrand - the whole thing is a mess.
Models you can download and use elsewhere if Google nixes access
Also - and this is bizarre - the token cost of doing that is higher, not lower, at least in Gemini world, and by a large margin. That's very counterintuitive, but a page encoded as image tokens can be smaller than same page as text, and is not meaningfully lossy on documents that are just typed text because the models are well trained on those.