Rendered at 19:09:12 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Sol- 1 days ago [-]
Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
miki123211 22 hours ago [-]
I find that "vibe coders" (that is, people who do not know anything about programming, but nevertheless produce useful tools for themselves and others) are using a lot more tokens than we do as programmers.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
inopinatus 15 hours ago [-]
It's because they don't know data structures.
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and essentially the same sentiment, three decades later:
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.
If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.
Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.
avmich 14 hours ago [-]
Watch LLM start paying attention to data structures.
jappgar 8 hours ago [-]
If you frame the conversation in those terms, they will.
One of the problems is that by default, they'll avoid changing data structures or architecture that is already written down.
Like a junior dev, they're correctly cautious about breaking things, so they prefer to write more code instead.
sdeframond 3 hours ago [-]
> If you frame the conversation in those terms, they will.
Indeed I realized recently that, when we complain about LLMs producing slop, that's in part because we dont ask them to refactor.
Coding agents won't, on their own, make a big change the user did not ask for. And this is fine.
stymaar 13 hours ago [-]
It's not going to happen naturally, the labs first need to implement a reinforcement learning pipeline that promotes it.
potbelly83 6 hours ago [-]
Falling back on a RL pipeline to cover gaps always strikes me as a more sophisticated version of the mechanical turk. If what we had was truly AGI wouldn't they be able to derive this from the data they already have.
sdeframond 3 hours ago [-]
Why would we care wether something truly is AGI or not?
It is useful. It may be dangerous. It has an impact. I care about that.
7 hours ago [-]
inopinatus 13 hours ago [-]
A coding model that can make genuinely well-considered data structure choices won't be a LLM, it'll be something more general.
hathawsh 12 hours ago [-]
While I agree that a coding model (such as Opus) by itself tends to act very shallowly, when it's driven by a harness like Claude Code, the combination seems to be a far more general thing than a LLM. It's capable of consistently making excellent data structure and architectural choices over large code bases. It imitates thinking about anything and it can drive itself for hours.
Honestly, if I simply fed it a sense of presence (I would repeatedly tell it what's going on right now and ask it to react if it thinks it should), it would feel eerily like AGI.
logicchains 12 hours ago [-]
Coding models can already make well-considered data structure choices if given all the relevant context, but a non-programmer doesn't know the context to give it or even to tell it to optimize the data structure choice.
gchamonlive 18 hours ago [-]
I think it's not only a matter of token efficiency. If you don't know what you are doing development will eventually crawl to a halt invariably.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
baq 13 hours ago [-]
We’re at 99% for a lot of stuff today, you get a third nine from council reviews and labs have another one or two nines in the pipeline. At five nines your task length horizon extends far beyond the current frontier model release cadence. More out of distribution tasks lose a nine or two, still revolutionary. You can get an extra nine from a good set of skills around slicing and distributing work according to model capabilities.
xbmcuser 17 hours ago [-]
The llm are improving though maybe a year from now they can use it to fix the code
gchamonlive 17 hours ago [-]
Maybe, it'll be exciting to see, and I'm all about accessibility, but in this case I also don't think it's about model capability or intelligence, it's the low information to noise ratio in the codebase. There just won't be enough information in the code itself to know what to fix. Fix how? What should it do? I'm really not sure you can reconstruct intention from a codebase created unsupervised.
alexytsu 13 hours ago [-]
Commit the intention as specs. If tokens/intelligence become that much cheaper over time, then "throwing away the code" to start again becomes feasible.
gchamonlive 8 hours ago [-]
Yeah, but not quite. I've done this with a project of mine, wrote it in python so it's easy and quick to read and debug, saved every work item and decision, done everything through agents so there was logs to go back to, and when I tried reimplementing in elixir everything fell apart.
Even after documenting decision and specs you need somehow to replay these in the correct sequence after you've validated these specs are still updated. Imagine you spec a feature, it works well, but in an edge case while doing a separate work you see something wrong, will you stop, fix and update the related spec? You will trust the AI will do this, and you guessed right, the counter-probability of success also has effect here, so eventually you will do undocumented changed in the codebase that won't reflect in the spec.
Now imagine all this but in the hands of someone that is an expert in their field but has zero notion what we are talking about here. Just look at the state of packages in R, the programming language, you'll see that technical competency and intelligence don't translate immediately to efficiency in a programming role.
I haven't had time to dive into it yet, but I think it might structure things in the way you want.
gchamonlive 7 hours ago [-]
Thanks! Will look into it.
This is something I didn't think about. Have a tech illiterate friendly harness that will keep asking technical questions the mainstream user won't be aware they needed be addressed, until there is enough evidence to either start an implementation or outright reject the project with suggestions where the user might look into to better prepare for another session.
noisy_boy 15 hours ago [-]
1. Implement feature and write tests for the code
2. Make sure tests pass
3. <Every now and then> Review code for quality and fix - make sure tests pass.
4. Go to #1
Overly simplistic? Yes. But I would wager that this can go a long way, even for vibe coders.
gbalduzzi 13 hours ago [-]
I keep seeing this but I'm not sure it is effective in the long run *if unsupervised*.
"Review quality and fix" doesn't mean a lot without context.
Does it mean to remove unused features and simplify the underlaying code? Does it mean changing the data structures to better support future development? Does it mean improving performance because of bottlenecks?
You are supposed to tell an LLM what your codebase needs, but if you just vibe code without knowing the code, "review quality and fix" will have unexpected results
gchamonlive 8 hours ago [-]
It's glaring to see who's still buying the positivists AI hype and who's done actual work with it.
AI is amazing, but people need to realise meaning and intention can't exist in a vacuum.
neya 16 hours ago [-]
Fixing code is not the same as fixing a fundamentally broken architecture. The latter requires understanding that the architecture is broken in the first place and that understanding comes from experience.
gchamonlive 8 hours ago [-]
And most of the time the architecture is not just broken, it's just meaningless because of so much accumulated error. If it was just broken then other comments would make sense, just throw a more capable model at it, but LLMs can't really reproduce shakespeare out of white noise just yet.
internet2000 16 hours ago [-]
Can confirm, I'm currently using Opus 5.5 to fix some Opus 4.5 slop from around this time last year.
gobdovan 21 hours ago [-]
I think this is valid now, but not guaranteed to be valid forever. For engineers, there was a period where more checks, more tests, more auto code reviews improved results quite a bit. People were consuming tokens like crazy (including me). Then things improved via better effort/thinking levels, where you could see repeated code reviews plateaued, so now people don't really do that quite as much.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
djmips 17 hours ago [-]
> you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about.
You sure that wasn't just working at Microsoft?
senderista 17 hours ago [-]
Astra is less nitpicky IME than Fable.
wahnfrieden 17 hours ago [-]
I’m seeing hundreds of review loops before settling, with Astra/Sol
gchamonlive 8 hours ago [-]
Could come down to the different nature of the work you guys do.
furyofantares 19 hours ago [-]
For my normal work I take ownership of the code, and end up with the exact code I want. I still have it go off and do a good amount of work a lot of the time, still queue up multiple tasks at the same time a lot of the time. Sometimes I throw it all away and re-prompt once it's time to commit to it, sometimes edit what it made, sometimes have it edit what it made etc.
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
arceister 8 hours ago [-]
Because that "vibe coders" didn't know and go through the fundamentals, thus they're wasting tokens with probably continuing the AI hallucination suggestions.
I've seen bunch of persons like this and that's kinda stupid because they're just blindly following AI's "suggestions" while they actually don't know what they're doing, then results on terrible code and architecture with "if it works, it works" mentality.
alkonaut 10 hours ago [-]
I (a programmer) just did a pretty large task with Opus 5.5 that was a perfect fit for a big token eating task. It ate into my weekly budget in a way that made me have to use anthropics one-off "reset" they offer now.
Long story: we have a big legacy desktop app. It uses a big legacy UI component (a grid control), which we had a license for in an old version. Fast forward 20 years, and to be able to move to a new runtime for our app, we need to update the component. Someone had bought the company making the component and now charges north of $1k per developer per year. So instead of doing this, we had just lived with the very old version.
We had long thought of writing our own control to replace the proprietary, but it was always going to be a man-year of work we thought. But I thought I'd give it a try with AI now. I told Opus: look at our uses of that control (tens of thousands of lines of code, it has over 100 instances across our User Interface). Write a new control that would compile with the exact same app syntax. First just make a dummy implementation that throws on every call. Then start implementing. Make a test suite that can run both with our new control and the proprietary control, and test everything, every function that can be called in its public interface and every state that can be inspected from the public API. Verify that everything behaves exactly the same, and lock it in with thousands of tests. Finally, check that the control _looks_ exactly the same as the proprietary one. Render to bitmaps, figure out the rendering logic from observation, such as arithmetic for padding, font sizes, and so on. Compare pixels until it's exactly the same.
Basically: it was a mammoth coding task, but it was so extremely well specified that an LLM could easily just do it. It's a clean-room implementation of something with no tests, but we had a test double that could provide 100% of the expected behavior. The description was extremely short. "Make a new thing that works like the old thing, and prove that it does". Opus 5.5 finished this in a number of hours. 500 source files, several thousand unit tests, and html reports with image diffs from the reimplementation and the original control. It did not use any disassembly or such "cheating". Only observation of the public API and the behavior.
Do we need to deeply understand the implementation? Does the architecture matter? Not much in this case I'd argue. It was a black box to begin with and it remains a black box. If we notice a bug, we can always point it to the original proprietary control and say "there's a behavioral difference when doing X" and it will fix it, and lock it down with tests.
As a programmer it's kind of chilling. I had recreated for a few tens of dollars something that would cost $1000 per year to buy. Obviously it's not a complete implementation only the parts of the API we use. It likely still has some bugs. We don't get support, we get to maintain it ourselves. But the rate of reverse engineering this thing "black box" was frightening. It hasn't created anything novel. But we must realize that as programmers some times we have man-years of work that just isn't novel. And in the past, we didn't do this work at all.
I wonder if those who write and sell libraries like this will start having explicit no-reverse-engineering EULAs soon? Perhaps even explicitly mentioning AI/LLM use in analysis and reimplementation?_ Obviously the library we reimplemented was from 2005 so didn't mention AI... (It doesn't mention reverse-engineering either, luckily).
Pannoniae 8 hours ago [-]
Most products do in fact have an anti-reverse engineering clause in their EULA, to be fair, it's been a standard EULA term for a long while. It's just that no one cares anymore...
alkonaut 7 hours ago [-]
Yes, and usually in the form "You may not reverse engineer, decompile, or disassemble the SOFTWARE or any of its constituents, except and only to the extent that applicable law expressly permits". (This is the concrete example from this software). And as far as I understand, this means that so long as you stay short of decompilation - you can reimplement as much as you want.
The law that covers this (in the EU) is EU Directive 2009/24/EC, where Article 5 is the reverse-engineering-without-decompilation.
> The person having a right to use a copy of a computer program shall be entitled, without the authorisation of the rightholder, to observe, study or test the functioning of the program in order to determine the ideas and principles which underlie any element of the program if he does so while performing any of the acts of loading, displaying, running, transmitting or storing the program which he is entitled to do.
This is pretty difficult to parse, but luckily there is a ruling from the European Court of Justice on this: SAS Institute Inc. v World Programming Ltd (Case C-406/10), delivered on May 2, 2012.
SAS Institute claimed that World Programming Ltd (WPL) infringed its copyright by studying the behavior of the SAS software system and writing a competing program (the World Programming System) that emulated its exact functionality and used the same data file formats. WPL did not have access to SAS's source code and did not copy any of its literal text or internal structural design.
CJEU:
> "It must therefore be held that the copyright in a computer program cannot be infringed where, as in the present case, the lawful acquirer of the license did not have access to the source code of the computer program to which that license relates, but merely studied, observed and tested that program in order to reproduce its functionality in a second program".
Which is a good find. But this is where I wonder if LLM-based reverse engineering is going to creep into either law (via lobbying) and/or EULA's, because this "observe every single state of the program for every single mutation" was simply not a viable mode of reverse engineering in the past. Or, it was at least always cheaper than just buying the software! Not so any more.
Or alternatively, that programs stop having so many observable states, making more things public. But for libraries as in this case, the whole product IS the public API. Without a rich public API, the library can't be sold. And with it, I can observe it and copy it - because it's internal workings are "too simple" not to be deduced from the public API. In short: a UI control is a ton of hard-to-write but easy to copy boilerplate code. And selling this has been an industry, but I wonder if it will be for very long.
Pannoniae 5 hours ago [-]
"And as far as I understand, this means that so long as you stay short of decompilation - you can reimplement as much as you want."
Yes, but my point is that.... go on github, you'll find tons of decomps. And many more done just privately too. One of the No Man's Sky devtalks start with "yeah we decompiled the terrain generation from this other game, implemented it in our prototype, it didn't work okay, here's how we've learnt from it to make something better". This was in 2016. More recently, this has been going on way more openly, even full AI-assisted decomps thrown up onto GitHub casually. It might be the letter of law or included in Terms of Service but no one cares really.
drbojingle 16 hours ago [-]
That and some have bridged the gap with tooling. starter projects+ Strick typing + dead code detection, linting rules and today's models can get you pretty far, especially if you plan out some basic architectural patterns with your starter kit.
It's not perfect but any means but it helps manage ones sanity.
spicyusername 17 hours ago [-]
Plenty of vibe coders who know a lot about programming producing useful tools and using tokens too.
drusepth 14 hours ago [-]
Indeed, I've been coding for ~25 years (competitively and in open source for many of them) and I max out at least two subscriptions' worth of tokens every week across a half dozen active projects.
I still code "by hand" sometimes (mostly Ruby/Rails, C#, and random languages for code golf) but just for fun at this point. Serious projects started being 95-100% AI over a year ago.
taliesinb 14 hours ago [-]
What are your half dozen active projects? And did AI use make you more ambitious about what those projects could be?
seanhunter 10 hours ago [-]
I’ve noticed some people re-prompt big tasks from scratch rather than iterating. This burns tokens very fast.
onevsall 9 hours ago [-]
Architecture is essential for anything complex. With bad architecture, complexity quickly outruns any model.
bensyverson 19 hours ago [-]
If you have well-specified tasks, you can easily reach 10 simultaneous agents working on disjoint parts of the code in worktrees. That consumes tokens pretty quickly!
daemonologist 18 hours ago [-]
Personally, it takes me longer to write the specifications than it takes the model to implement them (and it takes me much longer to review the resulting code, although maybe that makes me old-fashioned). Consequently I do not have enough tasks to run more than one agent at a time.
gbalduzzi 13 hours ago [-]
Exactly. I don't understand how so many developers seem to have a long tail of well written task specifications ready to submit to the LLM.
Who produces them?
bensyverson 6 hours ago [-]
I have conversations with a smart model, and then the model writes the spec. I review and approve the spec, and it dispatches.
For a concrete example, check out this random plan [0]. A detailed spec followed by the exact implementation tasks that will be executed by the subagents.
You don't need to have well written specifications. Just need to point it in the right direction and to know what you are doing so you can ask it to generate a plan, review and then ask it to implement.
8n4vidtmkvmk 12 hours ago [-]
LLMs are good at cleaning up tech debt. Give them lots of small refactorings or dig out those crusty old P4 tickets.
There's a lot of easy stuff for them that requires very little specification and very high probability they'll get it right the first time, especially if you can point them at an example done right.
stymaar 13 hours ago [-]
I'm exactly in this situation, and at the same time I get so many jumpscares when reviewing the code that I'm not going to stop anytime soon.
bensyverson 6 hours ago [-]
Not sure why I'm being downvoted for stating the obvious. To the parallel agent skeptics: I was also a skeptic until a month or two ago. I would run one agent, watch it carefully, and check its work. However the models got good enough that it was more efficient to do more work in parallel, then have a single agent integrate the changes with a critical eye, then run a code quality pass, and then I would take a look a it and kick the tires behaviorally.
It does involve letting go and not micromanaging every code convention and implementation detail, but that is the same skill you need when leading engineering teams.
thunky 19 hours ago [-]
Are they also telling each other what to do? Because I for one can't assign and keep track of 10 things at once.
bensyverson 19 hours ago [-]
I create large hierarchical plans, and have a coordinator agent divvy up the work. It's extremely effective.
nsonha 17 hours ago [-]
I created an orchestration skill for myself (using herdr but any persistent mechanism works). So I then only interact with a front session and it will triage and dispatch each request to the relevant spaces (each of them can have multiple worktrees of the same project), summarize movements and pending decisions for me all at once. I do not directly interact with a multiplexer or any dashboard.
thunky 4 hours ago [-]
This seems complex and expensive, and I suppose the only reason to do this is because you want to generate code faster? Do you really have so much code to write that a single LLM is too slow?
And then on top of that you need another agent to manage merging all of the subagent code together?
nsonha 3 hours ago [-]
I don't do this at my day job (coworkers would be pretty mad). This is for software ideas that come up weekly that I need to execute to at least MVP before I ever need branching/merging
Not complex at all, only one extra session other than the ones doing work and it's on a dumb model and can be thrown away & restarted because it only dispatches work, not doing anything.
I do everything in there, collecting requirements, kick off research, branching, merging, not one other agent on top. I considered making that orchestration command llm-powered but it's not justified at my current use.
It's not more expensive, in fact I could have just chugged along with the slow and manual session by session work but I have a claude subscription and another GLM one (the most low cost basic tier, not even much), that just sit there collecting dust if I don't put them to use in a more efficient way.
And doing session by session would face your problem when context switching too much become unscalable.
meowface 19 hours ago [-]
I am actually going to go out on a limb and guess the opposite of this is true, and that on average vibecoders burn tokens less readily than veteran software engineers. I could list several reasons why I think this would be likely. No idea which of us is empirically right, though.
(With exceptions for what I can only call the "manic vibecoders" with like 10 simultaneous weird slopprojects they're spewing out at once. Generally with each project itself being something related to vibecoding. Steve Yegge being an example of a "manic vibecoder-actual programmer" hybrid.)
vineyardmike 18 hours ago [-]
I’d think this matches my hypothesis. I’d say that I spend more tokens rewriting and fixing things, so that contributes more.
Also, I’d imagine the token-maxed user is a programmer that lives in chat. I’ll admit to having asked the LLM to move a method up/down in a file, and watched it burn tokens for a minute thinking and executing a menial task.
meowface 18 hours ago [-]
Same. I spend tokens on so much more than just the initial implementation of a feature I have an idea for.
gbalduzzi 13 hours ago [-]
This I don't understand. I'm faster at moving the method then at prompting the LLM to do so
VMG 9 hours ago [-]
LLMs are sometimes stupid and get confused when you change the files without them knowing. So asking them to do even simple things keeps the context in sync.
Plus you get a bonus random line "methods are all on the top" in the commit message that makes no sense to anybody.
hasbot 11 hours ago [-]
Sure, if it's right there in front of you in the editor. But if have to locate the file and the location within the file, it's easier to just type out what I want and have the LLM do it.
avadodin 10 hours ago [-]
I would be faster than Claude or Gemma if I did it.
That's a big if though and the blank page syndrome was already getting worse long before AI.
With age, it becomes easier and easier to get angry at someone or something until they work as expected than it is to actually do it.
I think this is why we've been seeing the genius coders from two generations ago embracing vibe coding even before it was cool or any good.
vineyardmike 12 hours ago [-]
Sometimes I don’t have the text editor/IDE open, and I’m just looking at the code in a PR or similar UI.
pmg101 9 hours ago [-]
"Sometimes"? I and I think many other people have been doing exactly this for most of 2026, assuming they look at the diff/PR at all and aren't all-in on dark factories.
FpUser 18 hours ago [-]
>"but maybe less important than they once were"
On browser based front ends it seems to be the case for me even though I still impose certain guidelines. On my C++ backends, no fucking way. Even the best models produce working but absolutely disastrous non scalable (performance wise and design wise) code unless watched over like a hen. Having said that - the value I get in either case is enormous.
tikhonj 18 hours ago [-]
I mean, they're trading off the time to learn to program for LLM time. Which might make sense! Locally.
But, in my experience, the projects where I have a constant pulse on the core design and abstractions in the code end up moving much faster than the ones where I don't. And I've been working on one of each at work recently, so I have a decent point of comparison.
erida_counter2 7 hours ago [-]
[flagged]
satvikpendem 19 hours ago [-]
LLMs these days write better architected and produced code than most programmers, so the fallacy that LLMs produce slop code is increasingly false.
drewnick 19 hours ago [-]
Every new generation of model goes and cleans up the slop of its predecessor in my code bases, and it has turned out to be quite effective.
Last year I held off on implementing a few features knowing that a model like Opus 5.5 was around the corner. I'm now implementing them in a much more efficient and quality manner than I could have fall of 2025.
satvikpendem 13 hours ago [-]
Indeed. Maybe people down voting me don't like to admit it but when models train on the entirety of human input you can assume they'd be better than the average human.
Tanjreeve 10 hours ago [-]
This is probably why web development and scripts are much more effective domains while anyone working on anything even slightly off the track is either tearing their hair out or writing a new layer of software to write the software.
maherbeg 1 days ago [-]
There's lots more you can do! Use the model to monitor your deployments after they get deployed. Have them fix and watch CI issues for you. Run adverserial review. Automatically watch metrics every day and highlight performance regressions. Start reviewing your previous sessions to find ways to statically reject different failure modes and have the agent have more success earlier on etc.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
mattm 14 hours ago [-]
> what would it take for you to care less about the understanding
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
sdeframond 3 hours ago [-]
I find that we dont need to go all in. I can use LLMs to make tooling custom to my project: linters, skills, rules, some doc etc. Then iteratively improve on that.
For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.
Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.
It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.
Continue this improvement process long enough and you may find yourself with an AI Software Factory.
maherbeg 5 hours ago [-]
Yeah, what things can you write to statically eliminate the things that are not acceptable. What would you have to change in your prompting process to get that outcome? What bad patterns is it copying from the code base that maybe you should spend tokens fixing?
I do agree that they're not great at program design by default and that's where we as engineers should spend our time. Data structures and data flow are king. But once you suss that out, they're pretty good at writing the resulting code.
This is also where I disagree with dhh about just using lower level languages. Good abstractions make for excellent program understanding and we should continue to build extremely good building blocks that make program design naturally solid.
ncruces 10 hours ago [-]
Not true. Claude is laser focused at finding the next load bearing thing. /s
klardotsh 23 hours ago [-]
The thing with watching CI in an agent loop is that it burns tons of tokens. At work I ended up writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code, and then updated my `/glab-ci-feedback` skill to use that. Saved a ton of token churn, and now I have a runbook a human could just as easily use if they don’t want to (or can’t) use an agent loop.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
ipsi 22 hours ago [-]
FWIW, Claude Channels[1][2] are probably going to be the solution for that, eventually. While I'm not sure how the WebHook receiver example will work with, say, GitHub and a local Claude, the Chat side of things _would_. So you'd have GH send its web hook to Telegram (for example), and then the Telegram Channel MCP would inject that into Claude, and Claude would start working on the problem. Still experimental, but functional enough to play with.
This sounds... horrible? I mean, it's certainly a solution to the "wake up when this thing happens" problem, but... $SERVICE -> webhook -> $CHAT_APP -> MCP -> remote wakeup sounds both brittle and - as you said - the local code harness route is entirely unserved by something like this.
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
wren6991 12 hours ago [-]
Hey now, if programming is solved they still need to achieve vendor lock-in somehow. Won't somebody please think of the vendors?
solatic 15 hours ago [-]
> The thing with watching CI in an agent loop is that it burns tons of tokens.
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
klardotsh 14 hours ago [-]
I guess folks' experiences with this stuff will vary wildly by what environment they work in. I use LLMs mostly at work, where I don't have any subscription plans, everything is billed per-token, and there's multiple coding harnesses with different token quotas available (and vastly different functionality). So the sharable CLI that works whether I'm in Claude Code (where tokens cost some outrageous amount) or Devin CLI (a horrible harness that also lacks any sort of scheduling system as far as I've ever figured out, but hey, there's GPT Luna and GLM available, at least) is a huge win.
unddoch 22 hours ago [-]
I think they are trying now to to bake CI awareness into Claude Desktop, didn't use it yet.
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
maherbeg 22 hours ago [-]
Yeah the codex app can deterministically poll and watch for you too. Consider it like an event based trigger, where the event can be anything you can dream of (like webhooks!)
miki123211 22 hours ago [-]
I think that's what a future dev team is going to look like.
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
emkoemko 18 hours ago [-]
why would there be customers? only customers will be of AI companies, you ask the AI and it will build the tool for you
bdamm 17 hours ago [-]
You're thinking that the entire economy will collapse down to just 3-4 vendors?
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
xoac 8 hours ago [-]
robot make sandwich, robot drive bus, robot cook soup
bdamm 59 minutes ago [-]
Who's making the robots? Who's owning the bus? Who's growing the ingredients in the soup? Who's deciding when the bus needs to stop due to a security or safety concern? Who's flirting with the customers and recognizing the power brokers? Not the AI companies.
Just because robots can do stuff doesn't mean the human power structures or service preferences evaporate.
Hauthorn 23 hours ago [-]
> Another thing to think about is, what would it take for you to care less about the understanding.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
miki123211 22 hours ago [-]
Because humans are currently the bottleneck.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
tshaddox 21 hours ago [-]
Humans could already produce more code than a human can understand. Even a single human in the pre-agentic era could produce more code than they could understand, certainly over a career and often even in the short term given the resources many companies give to maintenance.
A lot of old-school software engineering is about how to deal with this reality.
satvikpendem 20 hours ago [-]
No they couldn't. You can't create software you don't understand because you wouldn't even know what to type into the IDE in the first place. I don't understand claims like these, how exactly are people especially individuals producing more code than they could understand? Even at a huge corporation one might not understand all the code but surely they understand the part they're modifying because otherwise they wouldnt know how to modify it.
massysett 19 hours ago [-]
Yes, it’s possible for a person to create software he doesn’t understand himself. In the old days this was pasting from Stack Overflow and changing things until it worked.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
tshaddox 17 hours ago [-]
I’m referring to competent engineers maintaining understanding over time of all the code they’ve produced. Long before agentic coding, codebases routinely grew beyond the comprehensive understanding of their own authors.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
satvikpendem 13 hours ago [-]
As I said, you understand the part you're modifying because otherwise you wouldn't know how to modify it.
> literally hand-write code they don’t understand even as they write it
I find this literally impossible. How can you even start typing anything without knowing what to type?
tshaddox 4 hours ago [-]
I suppose I am confused about why you’re confused. There’s a long history in computing of describing pieces of programming languages syntax syntax as “incantations” and similar. I suspect it has been very common, especially in the early part of developers’ careers, to know what you’re trying to do and to know that this code accomplishes it, but to not understand how the code works, to not be able to use the technique more generally, and to not understand all the effects your change has on the rest of the system.
klausa 11 hours ago [-]
There's understanding and there's understanding.
Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?
People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.
It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.
gr_norm 20 hours ago [-]
I want to make better software, not more software. Making software development faster isn't necessarily the goal. Making it better in the many, many ways that matter (of which speed is just one part) is.
satvikpendem 20 hours ago [-]
There is more software to be written than there were programmers so lots of people do indeed want more software, for example small tools and one off projects that aren't worthwhile to make pre LLM.
jimbokun 16 hours ago [-]
Many people want less software to deal with now, not more. Being forced to download and update apps on your phone that previously could be done without an app, for example.
satvikpendem 13 hours ago [-]
Which people? They want the right kind of software, that which works for them, not garbage.
jimbokun 18 minutes ago [-]
A lot of people would find it easier to pull out a few quarters and put it in the parking meter than download and update an app and give it your financial information. Or hand over a paper ticket to get into an event that’s easy to transfer instead of yet another app with a barcode.
Etc.
smeej 19 hours ago [-]
I'm looking forward to the making of more software. I think there are probably people who have had really useful software ideas for a long time that they'd never be able to raise money for, but now for $20 a month, they can get up and running, serving their local and/or niche communities, without having to hire a team of engineers.
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
ThrowawayR2 18 hours ago [-]
The Android/iOS app store is already flooded with low quality apps. Almost 20,000 games have been released on Steam so far in 2026, that averages to about 70 games per day, every single day. Getting any kind of traction in such a environment is hopeless; every new app might as well be a scream into the void.
andrewaylett 21 hours ago [-]
A human can produce far more code than a human can understand, too, but pre-LLM we always viewed someone overwhelming their colleagues like that as being bad at their job.
jimbokun 16 hours ago [-]
Right.
So accelerate the vibe coding of shit nobody wants or asked for, just to see some metric go up somewhere.
Are we still getting bonuses for the number of tokens we can burn?
maherbeg 5 hours ago [-]
True, but if you can think about tests as "how do we validate that the outcome I want has been solved", then you can orient your tests around that. I do wonder if BDD style integration tests will end up being the path we go down.
Frontier models today don't really write incorrect code at the micro level. They do miss edge cases at the high level though, and that's what we want to test, is the scenarios.
maherbeg 22 hours ago [-]
There's different layers of understanding the system. I generally care about high level data flow, concurrency and performance (batching, holding transactions too long, back pressure etc.) rather than the mechanics of how the code actually does a thing. I still look to see what the final output looks like and ask my agent questions on how it fits in the larger system and evolve things if necessary, but agents are pretty good at writing code if the rest of the code base looks pretty decent.
23 hours ago [-]
tshaddox 21 hours ago [-]
More tests that aren’t written by you don’t help you understand the system, and I would argue the there’s no confidence without understanding. That was true in the pre-agentic era and is perhaps even more true now.
jimbokun 16 hours ago [-]
Absolutely fucking nothing.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
2 hours ago [-]
datadrivenangel 24 hours ago [-]
Opus 5.5 on Low seems smarter, cheaper, and faster than sonnet on medium, so what's the point of sonnet?
xgb84j 23 hours ago [-]
Claude Code has the issue that sub agents inherit the thinking level. This means that to use a smarter or dumber sub agent you need a different model. That's not a particularly good reason, but that's my one use case for Sonnet.
mnicky 21 hours ago [-]
You can also create custom agents with defined effort levels and use those.
copperx 23 hours ago [-]
Or just use a better harness.
jpease 22 hours ago [-]
Being that my first prompt can be something like: for task x/issue y, which model would strike the best balance between cost and capability…
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
willtemperley 15 hours ago [-]
> Another thing to think about is, what would it take for you to care less about the understanding
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
2 hours ago [-]
maherbeg 5 hours ago [-]
The cat's out of the bag already. We can't undo the idea of LLMs or coding agents. If training progress stopped today, we have years and years of harness improvements to extract more performance out of today's models.
We also have open weight models too, and ways to host those at home.
Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.
crooked-v 23 hours ago [-]
> Run adversarial review.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
maherbeg 22 hours ago [-]
lol yeah, our review bot does a cost based analysis and pauses itself until you re-resume if it goes over a threshold.
mattm 14 hours ago [-]
Yeah, I made this point above but LLMs just don't have a good sense of importance. They treat everything at the same level of importance and can spend considerable effort on things that just don't really matter.
phainopepla2 1 days ago [-]
It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
Imustaskforhelp 1 days ago [-]
> It's the "semblance of understanding" you're holding on to that is keeping your demand limited. I'm holding onto it as well, but I think these companies are assuming that human understanding will no longer be relevant for most codebases going forward.
In short, seems to describe vibe-coding to me?
What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
> Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
ihumanable 23 hours ago [-]
It sorta feels to me like extrapolating from "the internet has all the knowledge for free" to "we won't need tradespeople anymore"
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
theturtletalks 20 hours ago [-]
The plumbing analogy doesn’t match software. You pay for a plumber once and you can always choose a new plumber. With software, it’s a monthly fee that will continue to increase over time. More features behind higher tiers. And probably taking and selling your data. So why wouldn’t a person try to build something custom for their needs? They have the ultimate feedback loop of actually using the product and telling AI the issue and having AI fix it. Most software people create are probably something they never used it their lives. But the best software comes from people building it that also use it. Even Shopify started when Tobi started a snowboarding store online and couldn’t suitable e-commerce software.
You’re right about people not watching plumbing videos and doing it themselves. But the equivalent example would be open-source software in the tech example. But instead of reading open-source code to see how different features were implemented, AI can go and dig into the code and figure it out.
zer00eyz 18 hours ago [-]
> With software, it’s a monthly fee that will continue to increase over time.
The fact that you think this is the way software is priced is telling.
The model being destroyed here is that every piece of software is something that needs to generate recurring revenue.
theturtletalks 17 hours ago [-]
You’re right, the monetization of software is fading and building actual moats is becoming difficult. If you have distribution, regardless of your app, you still have a long run way.
I was more pointing out that software gets enshittified. A plumber necessarily doesn’t and if the plumber does get worse, you can call a different one next time. Software, especially B2B, has switching costs and lock in. So you just have to put up with it.
Another point is that most software started with a few features and to get more market share and support more use cases, it became worse for the users using the early features. That’s why they try to build their own so it’s not bloated with features you will never use.
customguy 7 hours ago [-]
> A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
You couldn't move the goal post further from "Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons".
So this one falls under denial for me.
munificent 23 hours ago [-]
> I wonder how others are coping with it aside from denial and cynicism.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
busssard 6 hours ago [-]
i am in a similar state, i joined a bigger family owned company as a source of safety.
And AI is not the only "threat" that we are facing.
I recently built this tool to come to terms with where to live, because i want at least some certainty on geographical factors:
https://om-intelligence.ch/projects/polycrisis.html
but the AI thing is on one side using lots of energy to keep up with it, and on the other side gives you an edge, because most people are not aware what even is already possible. so suddenly you are the "AI-Expert" just because you try to keep up to date. so if it all goes to shit, at least we have a chance to sniff it in the wind a couple moments beforehand.
Or make memes from it.. that helped me cope with it: https://t.me/RobotComrades
I hope you find a peer group to talk with and exchange and build community. it is so rewarding to talk to likeminded people that have a similar knowledge base and soothe some fears that someone might have, and have them help you with the ones i have... (i recently did a deep dive in custom DNA synthesis, and how connected those services are already to API https://www.twistbioscience.com/tapi )
as always accepting what is seems to be a healthy strategy
RGS1811 22 hours ago [-]
This made me a bit emotional. Thank you for the beautiful response.
munificent 21 hours ago [-]
You're welcome, and I'm glad it helped. Now more than ever, we have to try to connect with actual humans and take care of each other.
Imustaskforhelp 15 hours ago [-]
> So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
One of the quotes which might help as well (I think I have this even in my HN profile): The only thing we know about the future is that it will surprise us.
Not even experts are much more likely to predict for what its worth than a coin toss in many cases (especially if they believe that only one theory/idea will mostly predict the future)
It's a blend of things and ideas and the sheer interconnectedness of them where a small pocket can grow large and then also shrink and taking into account all variables and factors is just simply impossible for a mind. I think that although we feel we are being more informed about the world, that in it of itself doesn't prevent things in the future from happening. It just makes us alert and sad and anxious about it.
Yet this life is one which shouldn't be lived with sorrow and anxiety. It is one of beauty and greatness. In many ways, we humanity have come so far from the past (Our medicine is something that not even the mightiest of kings could get) and yes, there are many problems in the world and some things feel as if they are staying just the same or getting worse real-time.
But even then, worrying about it could lead to nowhere other than a path of misery. Also these problems are complicated enough that its extremely hard for a single person to bring change (not that I wish to demotivate that person but rather seeing the system as a complex nature)
So to me, its also a form of inward action. I can work on myself to be better prepared for the world that comes next. In the same time, I think that the present for me as well is good. I have great family and friends and have many qualities that I am proud of and I wish to share that gratitude to the people who have helped me along the way (my family/friends/ Hackernews!.)
Within the hustle culture, there is no time to relax but it is within the time of relax that I believe some of the most fruitful actions can come. I believe it just makes my mind more productive being in a calmer state.
here's a quote from how to measure your life that I hope can help some people:
I genuinely believe that relationships with family and friends are one of the greatest sources of happiness in life. It sounds simple but like any important investment, it needs constant attention and care(...)
You'll be tempted to invest your resources elsewhere but if you don't nurture these relationships, they won't be there to support you in hardships or as one of the most important sources of happiness in your life.
So thank you hackernews and have a nice day and please, please try to say gratitude towards someone close to you (within these tough times) and try to keep a balance towards inward focus, sharing time with friends/family and also writing on hackernews (as is my past time nowadays), balance is necessary :-D
So once again, I hope that its a call to action to say gratitude towards anyone. Just send them a big message thanking them and make their day as well as yours memorable, have a nice day!
Also you are allowed to make mistakes (everyone makes them!) and even though I am saying (preaching?) these things, I have found myself sometimes failing to act on these things as well but I just think that these help in being more mindful about them hopefully and can help provide a perspective. I wish to adopt more of these things in my life myself as well hopefully :-D
[Pardon me for the long post]
jimbokun 15 hours ago [-]
Being part of an organized faith community really helps with this sort of thing.
Not sure how atheists are coping with the existential risks we are facing.
boromisp 12 hours ago [-]
How does it help?
You either keep in mind all the horrifying little possibilities the future could hold for us and try to prepare, accept and cope, or you push it out of your mind using whatever techniques available to you to not drive yourself into nervous spirals.
Ultimately it comes down to what your brain chemistry allows in combination with ways you practiced dealing with stress, existential dread, cognitive dissonance, etc.
I guess submission to a higher power is one way to deal with it? That way it's no longer your problem (alone).
jimbokun 20 minutes ago [-]
It creates a bigger context, and a longer time horizon.
Even some very basic questions, like “are humans inherently valuable?” have been thought through and discussed thoroughly in many faith traditions. For many people not part of such a community it’s a question that’s suddenly very important and they lack the tools to address it.
andrepd 23 hours ago [-]
Damn, yet they still hire programmers, marketers, researchers like there's no tomorrow. I thought everything would be vibe coded and we wouldn't need to even understand code anymore. Which one is it?
The proof of the pudding.
egeozcan 24 hours ago [-]
I created a team of agents using Opus 5.5 to review and address findings on a job system I have in a side project with medium reasoning, and I burned through the 20x plan weekly limit in 2.5 days. They were using GPT-6-Sol for reviews, and it also used 85% of my OpenAI x5 weekly limit. Three hundred something commits in total.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
czhu12 16 hours ago [-]
The other thing too is that I’m having a hard time reviewing these massive PRs that are being generated. So much so that I’m having it write me a book (also vibe coded) that I can read through to learn about all the stuff it did as part of the pr review
https://ai-lessons.oncanine.run/
16 hours ago [-]
gregwebs 1 days ago [-]
> I want to retain some semblance of understanding
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase.
Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
jwpapi 22 hours ago [-]
I think going for more understanding is the way you need less understanding. The more solid your core understanding of your codebase is the less you need to know the details, the less missunderstandings the less iterations needed, the less mental capacity consumed
raincole 17 hours ago [-]
If I follow the principle "every single line of AI-generated code has to be reviewed and understood by me," I can't even use up the $20 subscription.
selcuka 15 hours ago [-]
I do that and can use up the $20 subscription, but I haven't been able to hit the limits with the 5x subscription so far (even though I use Fable for making plans).
mrbonner 18 hours ago [-]
I have been using the “free” Ling Flash model on openrouter for a bit over a month now. My work is an aside project at home building a Rust binding for an open source Zig code base library. The result is nothing impressive but also not a total failure: I have a feature parity binding to use in Rust vs. Python/java/typescript.
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.
chrismustcode 1 days ago [-]
Cache read is the same as Opus as well where most agentic workflow cost comes from.
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
pjjpo 6 hours ago [-]
Smart people delegate to dumb people. Of course ideally dumb people get slightly smarter. From my experience sonnet is mostly for opus or fable to delegate to. So delegates getting more useful is always a good thing.
Seeing 2000 years of history being replayed by the AI startups is pretty weird right?
busssard 6 hours ago [-]
that.
also i find it interesting how the capabilities are growing.
first speech, then code, then simple tools, then 3d objects, then desktop use
afro88 1 days ago [-]
I've been vibe coding a game and running multiple Opus 5.5 in parallel on Claude Code Cloud, 5x Max plan, and I'm yet to hit a session limit too. Not sure when I'd use Sonnet. Though it would be nice to switch back to Pro I guess
alansaber 1 days ago [-]
When they inevitably drop allocation after post-launch hype dies down.
randerson 9 hours ago [-]
While I mostly use Opus for coding, I use Sonnet in my actual application, because it is far cheaper at scale. I have Sonnet parse user free-text input, understand it and return the meaning as structured json. It doesn't need advanced reasoning, just a good enough understanding.
huntertwo 23 hours ago [-]
Plan longer chains of work / higher level goals that can be broken down into multiple chains of work. This will allow you to automate more work units to be worked on.
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
busssard 6 hours ago [-]
for very extensive agent swarms, you could tell the orchestrator to use sonnet agents for the eval of papers.
this way you can digest much more papers (content of whatever form)
so whenever you are dealing with volume rather than independence. opus5.5 can define the goals of a sonnet well enough, that i would trust it with a group of 100s of agents
y42 12 hours ago [-]
When working with a whole bunch of sub-agents, like you probably would do when coding, 2-3 sessions at a time are probably not enough. Like I build very specific sub-agents for code review, ui review and so on. Within one project I work with at least 3 - 4 agents then. Using Opus on every task would not leave room for other everyday-tasks.
dionian 1 hours ago [-]
Subagents
konsnos 12 hours ago [-]
Sometimes it's not about token usage but speed of the task.
I am using Sonnet 5.0 in browser (btw Claude in Chrome extension works in Edge) to download Datadog logs with multiple filters. It's running for about an hour, doesn't run out of tokens and does a splendid job.
swalsh 8 hours ago [-]
If you ever had claude code fan out a task to sub agents it will often use sonnet or haiku for those tasks. Web scraping is an example.
hamburglar 13 hours ago [-]
This is boat I’m in too. I get a ton out if my pro subscription, and I don’t hit the limits, but my Microsoft buddy was just griping about how the company recently imposed $10000/month token budgets on his team and he blew through his quota on under a day. The mind boggles.
losvedir 1 days ago [-]
Useful for API requests, when using AI in the product rather than to build the product.
neuronexmachina 1 days ago [-]
Most business/enterprise accounts also have to pay API rates.
losvedir 24 hours ago [-]
Exactly. I'm saying that Sonnet 5.5 might not be useful or necessary in a Claude Code session but it could be good value in the API when you pay per token.
doctoboggan 1 days ago [-]
I am mostly at the same point right now you are, but I think in the future with those "gas town" ideas we might be managing even more agents each.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
jcrben 19 hours ago [-]
Does that mean you're doing work week by week that is well-scoped and planned for that week and you can't pick up the work for next week until that week happens? It just seems like mostly programmers work on things that extend out for a long time and you can kind of just throw more at the problem and pull it forward earlier
schwarzrules 1 days ago [-]
The only advantage I could anticipate is I still hit session limits with Opus 5.5. My usage shows I'm on-track reach my weekly reset with room to spare, but yesterday I ran into a session limit. I switched down to Sonnet 5 for the next session, but performance benefit of Sonnet 5.5 is a compelling alternative for managing session limits.
eonmatrix 10 hours ago [-]
I'd use it for some of the lesser roles in oh-my-pi (omp) like commits.
ohyes 7 hours ago [-]
I honestly rarely use the “more powerful” models, I find they don’t really follow my instructions very well. So it’s medium effort sonnet for most coding tasks for me, escalating to opus for code review.
I’m just hoping they didn’t “improve” sonnet too much or it will become annoying to wrestle into doing what I ask it to do.
richardw 10 hours ago [-]
API. Get significantly smarter than Haiku at lower cost than Opus.
JMKH42 1 days ago [-]
One reason might be that Sonnet tends to be a lot faster, so since its almost as smart as opus maybe you use it to get work done quicker. In latency terms not throughput.
mgaunard 21 hours ago [-]
I find that I can do 4 to 7 sessions in parallel, and still review everything in depth and co-design.
I mostly use Fable though, Opus only via sub-agents.
smb06 20 hours ago [-]
I expect they’ll change usage limits in some of the plans or introduce new plans with new limits
Imustaskforhelp 1 days ago [-]
I understand the point that you are making but why do we have to fulfill the supply just as much as demand. There is a demand frenzy going on right now with still being substantially subsidized.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
eptcyka 20 hours ago [-]
Try telling an agent to go through your backlog.
jtrn 21 hours ago [-]
Subagents.
cft 10 hours ago [-]
I ran out of Opus 5.5 20x plan in 4 days out of 7, so my experience is different from yours.
bossyTeacher 13 hours ago [-]
> Probably a first world problem, but with Opus 5.5's efficiency, the limits on the 5x plan are simply sufficient for my everyday work, even when running 2-3 sessions at a time. So I wonder when I would use Sonnet 5.5.
Not every user of Claude is a programmer. Or even exclusively a worker. Claude has uses beyond work. Something that many in HN struggle to understand.
bbor 1 days ago [-]
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding.
Yes. Our career is over, as is our economy. Soooo... FYI :(
jobs_throwaway 1 days ago [-]
> we now have programmatic intelligence powerful enough to do most white-collar work
> the economy is over
Hackernews' neuroticism remains undefeated
Atreiden 7 hours ago [-]
What does the existence of this do to wages and compensation for the workers in the industries it upsets?
It depresses them. Significantly. White collar jobs constitute the bulk of global purchasing power. What happens to the economy when aggregate purchasing power drops? The naive response is "prices fall until equilibrium is reached again"
But what if the needle continues moving so quickly that equilibrium is never reached?
This is the K-shaped-economy concern. The ultra wealthy and those who own the "AI means of production" will become unfathomably wealthy at the expense of everyone else.
Why should this not be a concern? Historically, this trend has always precipitated bloody conflict.
bbor 20 hours ago [-]
Just listening to the science, my friend.
jobs_throwaway 24 minutes ago [-]
What 'science' says that the economy is over, friend?
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.
thefourthchime 24 hours ago [-]
Yes! The point being that up until yesterday, every model struggled with this, and now they don't.
copperx 23 hours ago [-]
"this" being recreating Pacman specifically, or games?
londons_explore 23 hours ago [-]
I made ~10 games with opus 5.5 (all multiplayer web games over web sockets).
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
igleria 9 hours ago [-]
> "the blaster weapon is way too powerful, divide it's hit points by 10"
this is literally faster to do it yourself
> "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS"
these would not.
weird-eye-issue 7 hours ago [-]
> this is literally faster to do it yourself
It's literally not unless you already know exactly where it is in the code
Also even if it is I find that the extra mental switching is not worth it, that's why I even have it do basic things like updating the text in buttons these days. There is no point in using my mouse and keyboard to track down a file and then make the change when I can just use my voice to tell it what to change and then wait a few seconds.
copperx 22 hours ago [-]
Other models fail at oneshot creation of similar games?
thefourthchime 23 hours ago [-]
Your welcome!
sixtyj 23 hours ago [-]
I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
lukan 11 hours ago [-]
" If we compare it with pelicans that are still not-perfect…"
I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
sally_glance 23 hours ago [-]
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
thefourthchime 20 hours ago [-]
Just added! I had Opus 5.5 look at them, not a perfect way to score them but it's close-ish -- Best would be a ELO, where people play both and rank a winner, but I don't know if people want to bother doing that.
ilamont 23 hours ago [-]
Thank you for doing this. It is very helpful not just for capabilities but also for costs.
pyaamb 22 hours ago [-]
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
sunaookami 21 hours ago [-]
GPT models really have no taste huh.
coopykins 11 hours ago [-]
Surprising how cheap Sol 6 was.
russellbeattie 24 hours ago [-]
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
thefourthchime 23 hours ago [-]
Thanks!
fakedang 22 hours ago [-]
Interesting. Sonnet 5 was horrible, and Opus 5 was unplayable, but both Sonnet 5.5 and Opus 5.5 were about as close to the real thing.
formvoltron 22 hours ago [-]
oh! How about pengo, dig dug, & defender?
thefourthchime 20 hours ago [-]
I'm afraid those will be too easy. I'm not sure what the next game should be...
abejora 1 days ago [-]
Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
eli 1 days ago [-]
Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails
abejora 1 days ago [-]
You're right about its real world performance, and I worded my original comment wrongly.
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
joeyhage 1 days ago [-]
Claude, is that you?
bb-connor 22 hours ago [-]
you're absolutely right to push back
ramon156 23 hours ago [-]
Your clarification makes sense. The distinction between overall benchmark performance and why Terminal-Bench is an outlier is important
swiftcoder 23 hours ago [-]
> You're right about its real world performance, and I worded my original comment wrongly.
Damn, HN commenters starting to talk in claudisms now
verdverm 23 hours ago [-]
this is human writing...
this is claude writing...
corporate needs you to find the difference
spider-mario 10 hours ago [-]
You can read “a little bit” (i.e. “not too much”) into it (it does indeed tell you about the out-of-the-box experience), but e.g. being able to know when a fallback model has been used means that in terms of pure accuracy, you might still be better off defaulting to Opus 5.5 and re-routing to Sonnet 5.5 yourself when you get the fallback.
chis 23 hours ago [-]
Well presumably now it’ll fall back to Sonnet 5.5 lol
subscribed 24 hours ago [-]
I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.
Anthropic made it that way, and I'd say the lower score is accurate.
bitexploder 19 hours ago [-]
If you care about the things terminal bench cares about, yes. Sonnet was probably trained aggressively on agentic coding and things that align well with deepswe and terminal bench and or tuned heavily for those tasks. Sonnet is an agent likely to do more of those tasks and be given the more grunt work tasks. Whilst Opus' wider knowledge pool means it can deal with a much higher variety of real world situations successfully. And, those benches are often timed or limited. Opus may have been running out of time. Looots of factors.
subscribed 10 hours ago [-]
I see you.
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
cromka 19 hours ago [-]
It matters if it's not Sonnet performing the task, doesn't it?
subscribed 10 hours ago [-]
*Opus
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
shmel 2 hours ago [-]
It'd be relevant if I could disable safeguards. As long as I can't, this is Opus 5.5 performance I have to deal with.
Leary 1 days ago [-]
And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!
radlad 1 days ago [-]
I believe you meant to cite the Opus 5.5 System Card which states:
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max
AA intelegence index (agent harness doesn't have sonnet data yet) on max:
Astra 27k
Fable 5.1 78k
(Sonnet 5) 118k
Opus 5.5 119k
Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
MadameMinty 1 days ago [-]
That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
manojlds 1 days ago [-]
Fallback was usually Opus 4.8
Aissen 12 hours ago [-]
Note that it seems that it no longer falls back automatically. So the actual score of Opus 5.5 will be even lower (fail vs fallback than can succeed)
manojlds 1 days ago [-]
Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?
verdverm 22 hours ago [-]
this "feature" is one of the primary that caused me to cancel and move to exclusively open weight based systems
falcor84 23 hours ago [-]
So I suppose the easy fix for Anthropic would be to have Opus 5.5 now fall back to Sonnet 5.5, right?
azuanrb 19 hours ago [-]
Unless you’re using frontier models like Astra, Sol, Fable, or Opus, I think you’re often better off using Chinese models for a fraction of the price. I’m not sure people realises just how competitive they’ve become.
GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.
I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.
The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
AbstractH24 17 minutes ago [-]
It's still unclear to me how much I'd have to use a Chinese model to equal my Claude Max subscription price.
There's something to say for flat pricing rather than per token. Even if its not a better deal.
persedes 17 hours ago [-]
Those models are cheaper per token, but depending on your use case you might still end up paying more with the cheaper model. DeepSeek likes to burn through a lot of tokens for example, which can quickly ameliorate those savings. They are great models and have most likely helped anthropic and openai drop their prices lately, but I still don't see the monetary benefit of using them atm.
noisy_boy 15 hours ago [-]
That has been my experience with DeepSeek after the price increase. I was so used to it being so frugal, it took couple of top-ups for me to realize that it wasn't the budget-king anymore.
dools 13 hours ago [-]
Deepseek 4.1 flash is back to peanuts pricing and more capable than 4 pro
dkarvik 12 hours ago [-]
The main difference I notice is that 4 pro has more "world knowledge" which makes sense given its lineage. For most tasks flash is great though and the model can just fetch the code or library information it needs.
grahamnorton39 15 hours ago [-]
You might mean “eliminate” or similar - to ameliorate something is to improve it or better it :)
benced 19 hours ago [-]
Luna is the main exception to this. It's such a cheap model and the tokens come so fast.
tripleee 17 hours ago [-]
Luna is genuinely my favorite model. I like developing in small chunks instead of huge sweeping changes, and Luna is so good for that.
CharlieDigital 8 hours ago [-]
Try using something like Herdr and have an Astra or Opus drive/orchestrate multiple Lunas. Best of both worlds.
pkulak 14 hours ago [-]
I cancelled my codex subscription 2 days ago, but only because I knew I could wire into Luna API pricing. Luna xhigh writes very good code, basically for free.
shepherdjerred 15 hours ago [-]
Yup. Luna is fast, cheap, and pretty intelligent. IMO the Codex harness is a bigger limiter than the model itself
user43928 8 hours ago [-]
It's pretty dumb below xhigh reasoning as far as I know.
It seems weird to me that just using a ton of output tokens manages to produce a decent result in the end.
It seems to work well though. Sometimes I fear that it might be more likely to eg. run an incorrect, destructive command, but maybe that concern is not justified.
solarkraft 10 hours ago [-]
What bothers you the most about Codex? Looking from the outside, it looks like a pretty capable harness with a pretty good client.
broodbucket 11 hours ago [-]
Give omp a try.
scuppernong 18 hours ago [-]
if you're paying API prices, you probably already know this. everyone else is using a subscription which is massively subsidized rel API prices. or am I missing a third case?
phoghed 17 hours ago [-]
Does anyone beat Luna on price? It’s surprisingly capable for a lot of things.
Most enterprise customers are paying per token at this point afaik, whether that’s to gh copilot, Anthropic, or running models on Vertex/Azure/Whatever
dools 13 hours ago [-]
Deepseek v4.1 flash is way cheaper and comparable intelligence
toasty228 13 hours ago [-]
Is it? The artificial analysis "price to run intelligence index" is almost 4x cheaper with Luna than deepseek
In practice for whatever reason I find Luna surprisingly less capable, though. I keep trying, and it keeps failing in seriously odd ways for the coding tasks I'm using it for, in a way that GLM 5.3 Flash doesn't.
dools 5 hours ago [-]
I go for weeks without topping up my Deepseek balance, and have to top up my OpenAI balance at least every week.
GPT 6 Luna closed the gap significantly for sure (it seems to be about twice as expensive as DS v4.1 Flash), but Deepseek v4.1 Flash is still the best value model and capable enough for almost everything I need to do. Sometimes if it's babbling or can't nail down a solution I switch to Sol for one prompt, get the solution, then switch back to DS. I use 4.1 Flash almost exclusively though for both planning and implementation these days.
I was a heavy Kimi k2.5/2.6 user but since 2.7 Kimi has gone way downhill -- even the previous models. I think they got under heavy load and had to quantise their models to avoid going broke.
fallingbananna 10 hours ago [-]
I know it's not a perfect benchmark, but in the "Create a Pacman" prompt it did cost 4 times as much, while producing way worse result than Sol-6: https://jonclegg.github.io/pacman-bakeoff/
CharlieDigital 8 hours ago [-]
Quick scan of the repo and I didn't find the prompt or the setup of the harness, but I'd say that most people using DeepSeek are probably also not use it raw and without any guidance.
Is a zero-shot, zero-context prompt a useful benchmark? Yes, in the absolute sense. Does it reflect how teams would use it in the real world? I think in real-world use cases (IME), DeepSeek gets the job done.
dools 4 hours ago [-]
Did you actually play the games? The Luna version is super buggy. Like when you eat a cherry you don't properly eat the ghosts, and you don't see your points when you eat a ghost. The DS looks worse but is more solid gameplay.
The DS one used Claude Code via OpenRouter, the Luna one used Codex. I'd say a big difference in cost is coming from the harness, and quite possibly the different meanings of "one shot" in each of those harnesses. The Luna one probably spent less money, but might have also done far less real browser testing and testing/verification is the expensive part.
The better graphics is probably just Codex system prompt.
In terms of real world coding usage, I am finding that DS v4.1 Flash is about half the price of Luna for comparable workloads. G6L is super cheap for sure, and super capable. It's by far the best coding model from a frontier lab for everyday coding work, but DS v4.1 is even cheaper and no less capable in my experience.
madeofpalk 17 hours ago [-]
Third case is you’re just an employee at a company who pays for an enterprise codex/claude/whatever for you, and you just use whatever the best available is. Maybe they’ve locked away Astra or Fable, but if cost literally isn’t an issue (at this moment) is there still a benefit to Deepseek?
raincole 4 hours ago [-]
Chinese models are not for a fraction of the price though. If you only have a budget of $20 a month, luna with chatgpt subscription actually gives you much more room than deepseek. The idea that Chinese models = cost efficiency is rather outdated.
BoorishBears 16 hours ago [-]
I've seen the opposite: if you're doing work that doesn't need the frontier, it's really hard to beat lower-tier frontier models.
Luna is obviously very competitively priced, and I'm expecting Haiku 5.5 will be strong based on this release and get back to more competitive pricing since the model family shrank this generation (though Anthropic has proven my expectations wrong on the latter before)
Where GLM, Kimi and co shine for me is when you need to offer near-frontier capabilities in your product and straight up can't afford frontier models: if you're offering Opus in a product with API pricing, $20 a month Claude Pro is offering about $500 of comparable usage in a harness that flexes to a lot of tasks.
Offering GLM/Kimi increases the max complexity of problems you can solve successfully compared to stepping down to Luna/Haiku, while letting you offer a reasonable amount of usage. Once $20 a month Pro is comparable to "just" $100 of usage in your product, it's much easier to close the gap with UX, a better constrained harness, etc.
vincengomes 13 hours ago [-]
For the folks asking what is the point of Sonnet 5.5 when Opus 5.5 is better in every way, Sonnet is now the new default in the free tier and now people using Claude Web in the free tier get access to an almost close to the frontier model.
ddxv 12 hours ago [-]
I've been using the Claude free tier (browser chat) for awhile and never seem to hit restrictions anymore. There used to be many more restrictions a year ago. At the same time, if they enforce a minimum paid tier I'd probably just switch to Gemini/ChatGPT or whatever other free model is around at the time.
etatester 10 hours ago [-]
Hello fellow copy-paster. I hit them regularly when asking for medium refactors with higher effort. I have ChatGPT in another tab for when that happens, but to me it's awful for anything non-trivial. Gemini I don't even consider unless I'm just asking "what would Gemini do"
geokon 10 hours ago [-]
I've never hit anything with Qwen. Very rarely it will silently switches from 3.8 to 3.7 (which makes it notably dumber) but you just click the drop down at the top to switch back. It's rare though. I have >week convos with no switching
wongarsu 1 days ago [-]
"Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Aissen 10 hours ago [-]
The fun part is that the cash grab the frontier labs are running on cyber tasks might motivate enough people to pay for third parties; i.e it might bring enough cash to sustain Chinese competitors (and their open weights marketing strategy, which we all benefit from).
ttul 1 days ago [-]
Daybreak Blue is not bad and the bar to get into OpenAI's program is reasonable.
watusername 24 hours ago [-]
Hold on, is there any bar to begin with? For OpenAI's Daybreak Blue, I only had to go through the Persona KYC to gain access. With Anthropic's I had to submit links to my profile and briefly describe my use cases, which I doubt were read by any human being but at least there's some semblance of barrier.
ttul 23 hours ago [-]
Yeah, I did say the bar is low :)
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
nicce 22 hours ago [-]
How well it is documented or known that how they use the passport information and so on. Current blocker for EU citizen is to share that data for AI company…
stavros 18 hours ago [-]
Do you have a link for getting into the Anthropic one? I couldn't find anything.
> The Cyber Verification Program (CVP) is a free, application-based program that is designed to enable professionals to continue working on legitimate dual use tasks safely while minimizing interruption. If your use case has a legitimate defensive purpose and is being affected by these safeguards, we encourage you to apply for the CVP. See our Help Center article[1] on the CVP for more information.
what is the easyest way to use the chinese models and which harness does work with them well?
Iolaum 23 hours ago [-]
OpenCode harness with their subscription would be my recommendation.
malshe 23 hours ago [-]
Between OpenCode and Openrouter which one would you suggest? Sometimes I have pure grunt work to be done on non-sensitive data for which I want to use Chinese models. For example, tasks like extracting something from publicly available large pdf files.
Flere-Imsaho 22 hours ago [-]
Opencode Go, with the Deepseek 4.1 Flash model feels like a bottomless pit, which is great for grunt work.
malshe 54 minutes ago [-]
OK, I will try it out. Thanks
asp_hornet 18 hours ago [-]
Opencode for harness.
I use GLM directly from z.ai, they do not retain or train on your data accordingly to their TOS.
beveradb 23 hours ago [-]
opencode with model inference on cheaperinference.com has been working well for me - glm-5.3-flash is shockingly cheap (i've spent a total of a few dollars over several weeks of heavy usage), fast and capable for cyber tasks
AbstractH24 19 minutes ago [-]
I had to do a double take to realize why this wasn't week old news.
Getting hard to keep track of opus vs sonnet vs....
Assume 6 will be announced on around an IPO?
MisterMunchkin 24 hours ago [-]
It costs 20x more than the Chinese models I use. I just don’t need them anymore. Sure I’d use them if forced to for a job, but I don’t pay them outside of that anymore.
And my job won’t even pay for Claude now because it’s so ruinously expensive.
yipinwong 23 hours ago [-]
Say that to Luna's face. Ya all bringing up this not-so-cheap-nowadays chinese models and not that more intelligent than luna and bringing "cost" as the only factor.
cromka 19 hours ago [-]
Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs. Don't get me started on Luna.
From my today's session with Sol:
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/amend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
- hallucinated several facts despite me asking beforehand to check online.
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to what Sol 5.6 was a month ago.
I can't deal with this sh*t anymore, I have no trust in the tools I use and both OpenAI and Anthropic do the same thing.
gruez 18 hours ago [-]
>Not even GPT6 Sol cannot match DeepSeek 4.1 in my work, with outrageous bugs.
That seems hard to believe even with deepseek's own benchmarks. Not to mention for every person who says chinese ai is ahead of american labs, there's like 10 saying that they're benchmaxxed or that they're merely "decent value for money".
solenoid0937 16 hours ago [-]
HN desperately wants Chinese AI to be competitive so there's a lot of self delusion and wishful thinking going on.
I use Deepseek 4.1 almost every day as well, it's nowhere close
stymaar 13 hours ago [-]
There's an interesting contradiction in your comment: either the Chinese models are competitive, or you wouldn't be using it!
mannycalavera42 9 hours ago [-]
maybe it's the combination of both?
the _could_ have a judgment backed by firsthand experience for example...woah, I know!
It's not like rooting for a favorite football team. double woah!
stymaar 8 hours ago [-]
Why would they use it if it's not competitive in some way though?
MitziMoto 16 hours ago [-]
Yeah, I don't know what model that guy is using. He's talking about Sol like it's GPT4.
I haven't had these issues in multiple model generations of models.
Failing to understand English grammar? Give me a break.
cromka 10 hours ago [-]
> me: "The tool divides network rates by 1,000, but uses the \`KB\` instead of \`kB\`. \`KB\` (alongside the SI-standardized KiB), is reserved for units based on 1,024. This change corrects the usage of those units.
> it: That still attributes two claims to SI that SI does not make. `KB` is not reserved for 1,024 bytes, and `KiB` is the standardized binary symbol.
> me: where does it say that SI reserves KB?
> it: Nowhere. You said *KiB* was SI-standardized; you did not say SI reserves *KB*. I misread your sentence and argued against a claim you did not make.
It was correct to point me out on my mistake in essence, but still misunderstood my bracketed "(alongside the SI-standardized KiB)" sentence.
Sure, it wasn't a grammar mistake as such, more like a logical one, but it still shouldn't make it. I had more than one such issues already with it, this one was most pronounced.
What it interesting, though, is the number of corporate apologists my comment brought in. It's like it doesn't matter how many times OpenAI and Anthropic have botched some of the models while keeping the branding, some people would still die on that apologist hill.
NewsaHackO 10 hours ago [-]
You can tell how vague the grammar complaint was it probably wasnt a big of a gaffe that they think it was. Also, slight tangent, but the part where he said that it admitted where there was an error but was not able to say why it made the error is one of the issues with model sycophancy; it want to validate the users feelings of being wronged, but also does not want to say something factually incorrect. So it produces this fail state of saying sorry for nothing, but it cannot backtrack.
pavo-etc 15 hours ago [-]
Just yesterday I ran a comparison of $/message through my harness[0] comparing Deepseek models to Luna, and to my surprise Luna won. I suspect its partially due to Deepseek's long thinking times, and also maybe due to OpenRouter variance in cache pricing etc.
Model and observed window | Messages | Retrieved actual $/message
GPT-6 Luna — 23–28 Sep, partial final day | 88 | $0.0329
I've subbed to Codex because I suspect at my usage rates the Codex Plus plan gives me more Luna messages than I'm using, and I've not really observed and better or worse intelligence performance. Interested to see how my $/message comes out after a month of usage on the Codex plan.
Something nice I've realised about my harness is that I can run different agents on different models so I can collect pricing data for a bunch in parallel.
Why would you use a cloud model that's no better than a local one though? Because if you pick Luna then you don't have to compare it to the big Chinese models, Qwen3.8-27B is what you want to compare it to, and the comparison doesn't make Luna look good.
usef- 20 hours ago [-]
If it's for personal use, is there a reason you don't want the subscription?
Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model. Their raw API prices have fat margins.
And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.
daemonologist 18 hours ago [-]
For personal use I simply need less than $20 of tokens, even at API rates. If they (Anthropic or OpenAI) offered a $5/month plan that gave you ~$50 of credit I would probably sign up.
holbrad 3 hours ago [-]
That's only true if you're paying API prices, but you really, really should be using the subscriptions. Both the personal and business ones are still really good value.
throwa356262 24 hours ago [-]
Obviously not as "intelligent" but almost 10x cheaper
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
lcampbell 22 hours ago [-]
I was under the impression that the cache write fee was added to both the input and output costs (except in cases where the cache write is explicitly disabled via e.g. DISABLE_PROMPT_CACHING). The output becomes part of the context, after all; if they don't (for some reason, due to disaggregated inference perhaps) then I'd expect output tokens get charged both output then input+cache_write on the subsequent completion request.
The pricing model confuses me though (I presume by design, Hanlon be damned).
20 hours ago [-]
20 hours ago [-]
anvuong 17 hours ago [-]
Which provider are you using Mimo from? The official Xiaomi one seems to be very slow/buggy for me in the last few days.
tintor 22 hours ago [-]
You don't have to pay for cache write if prompt isn't part of conversation.
edu 23 hours ago [-]
What model are you using ?
system2 23 hours ago [-]
Not him but 3 models dominate: GLM 5.3, Qwen 3.8, Mimo 2.6. All censoring certain things. Numbers and other uses are perfectly fine. They are like 0.10-0.15 per 1M tokens. American AI lost the game already, people just can't see it.
toasty228 23 hours ago [-]
> American AI lost the game already, people just can't see it.
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
fearmerchant 4 hours ago [-]
It could be a Linux vs Mac/Win situation though. Sure the HN crowd will build their own harnesses and hook up cheaper models but the average Joe Sixpack wants something that "just works" out of box and the frontier labs seems well positioned to give that to them.
BeetleB 22 hours ago [-]
You're a decent sized company and wants to manage the SW + security on all your employee's PCs. They need to be able to update/remote SW on your machine remotely, see your settings, etc.
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
SSLy 21 hours ago [-]
>Linux
cfengine is 33 years old
Lord-Jobo 23 hours ago [-]
The difference is that Microsoft did that while relying on the extremely load bearing windows ecosystem. These AI companies have no equivalent lock in, nothing even close to it honestly.
toasty228 23 hours ago [-]
I can guarantee you 80% of people will call any llm "a chatgpt", most have never heard of claude, even less of opus, or sonnet, "deepseek" probably reminds them of a brand of toothpaste or something like that, "GLM" might make them think of the new mercedes SUV perhaps. 99.9% will never self host, nor send a single sent to a chinese model provider.
system2 23 hours ago [-]
They are not the ones spending API money.
toasty228 23 hours ago [-]
I don't know a single company using deepseek internally in any capacity, and I have friends in a lot of tech/tech heavy companies, virtually all of them use claude, the lucky ones get cursor with claude/chatgpt/grok.
cromka 19 hours ago [-]
Interestingly, I hear about them doing that all the time. But this is in EU.
system2 18 hours ago [-]
Once again, just like your other reply, you are mixing things up- a logical fallacy called the straw man; look it up. Your responses usually use this type of misdirection. Coding agent =/= API.
AI does not mean coders coding with agents all day long. AI integration is mostly for data processing, which is the promise and use of the APIs.
noisy_boy 14 hours ago [-]
Most people I know working in banks etc are using Co-pilot. Not because it is any good, because it integrates with the entire MS ecosystem - management pushes it. I would suspect that is the main driver of MS's AI penetration numbers for desktop-level non-API usage.
csomar 11 hours ago [-]
When Microsoft did that there wasn't really any competition and they were also the cheap option. I don't think they would survive if there was a sufficiently good chinese OS at the time.
system2 23 hours ago [-]
Most businesses I know switched to Google Sheets or Google Docs.
toasty228 22 hours ago [-]
Never heard about anyone using google sheets and docs for messaging, email, presentations, etc.
joseda-hg 21 hours ago [-]
Gmail, Meet and Presentations do all of those
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
conception 20 hours ago [-]
You’ve never heard of anyone using gmail for email?
madeofpalk 17 hours ago [-]
Huh. I’ve never worked at a company that doesn’t primarily/exclusively use Google Docs/Sheets/Slides. Guess that goes to show we’re really all in our own little bubble.
system2 22 hours ago [-]
[dead]
presentation 15 hours ago [-]
Basically only true in the Silicon Valley bubble. Professional services are Microsoft Office/SharePoint all in.
ndm000 21 hours ago [-]
This ignores two things.
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
case540 22 hours ago [-]
People dont want iPhone 14s in 2026. People want the latest and greatest. Chinese companies desperately trying to get western usage of their models
nehal3m 22 hours ago [-]
They would if those phones were 899/10=89.9 bucks.
keyworkorange 21 hours ago [-]
How are these models so cheap?
system2 16 hours ago [-]
They are just not overcharging. Nvidia's top AI chip Rubin sells in 72-GPU racks for about $3.5–7.8M. A rack running Xiaomi's MiMo V2.6 Pro generates roughly 150–300B tokens a day, worth about $130–260k at Xiaomi's API price.
That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit. This is why they can charge normal prices. Not $50 for 1M tokens.
tripleee 23 hours ago [-]
From my experience GLM 5.3 is at most 6 months away from the frontier models, and good enough for most tasks already
The usage numbers tell a different story. Not only is the great majority of AI users on American frontier providers, they are also willing to pay the premium price for the premium product. It's not like tech is oblivious to the Chinese models. It's that the industry is aware they are always 6 months to a year behind.
You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.
You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.
swingandamiss 22 hours ago [-]
Europe didn't even start in the race. Europe is the biggest losers of all it seems. What a shame.
etatester 10 hours ago [-]
The reason why that happens is not really Europe's fault. Any remotely-competent developer just works for US companies, why would they not? Especially for AI, there just isn't enough investments in the EU to offer those salaries.
What's left are incompetent developers paired with people who get government contracts and leech off them. Ask me how I know.
jpgvm 22 hours ago [-]
What are they censoring that matters to programmers?
zarmin 23 hours ago [-]
What harness are you using with those?
beveradb 23 hours ago [-]
opencode
anvuong 22 hours ago [-]
Does AI-style writing start bleeding into the comments, or is HN now also full of bots like reddit?
wkcheng 1 days ago [-]
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
usaar333 1 days ago [-]
Per the charts, there is largely no point to using Sonnet 5.5 at high+ as opus low generally will give similar performance at similar or lower cost.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
verdverm 22 hours ago [-]
at that point, you can switch to dirt cheap open models
NotSuspicious 19 hours ago [-]
Only if your company (and government) lets you!
verdverm 18 hours ago [-]
yup, we use Fireworks.ai, and American company with ZDR
Their new Ember-1 model is pretty good, fine-tune of Kimi3 with way less thinking
Does this make it an American model or is it still Chinese? Does it matter?
It appears, at least from a quick look, to be noticeably faster than Opus. If true, and you don't need xhigh/max reasoning for your use case (like a well-defined set of code changes), Sonnet might get the job done much more quickly.
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
ricardobeat 1 days ago [-]
At low and medium effort it is 1/3 cheaper, at high it’s a step above Opus/low. It only looks worse at xhigh.
wkcheng 1 days ago [-]
That makes sense. I'm interested in seeing where Haiku 5.5 comes in then when it gets released. It feels like the low intelligence / fast niche will be covered there.
oh_no 1 days ago [-]
i'd love to see them re-enter that space but given haiku 5 never happened I wouldn't bet on it
i think they see what openai charges for luna and just don't want to try and compete
water-drummer 24 hours ago [-]
They did mention in the Opus 5.5 announcement blogpost that Sonnet and Haiku 5.5 will follow soon.
canad3nse 24 hours ago [-]
But they literally stated that they would release Sonnet 5.5 and Haiku 5.5 after Opus 5.5 was released
mchusma 18 hours ago [-]
My guess is that haiku will be a mid release. They don’t seem to care to compete for the low end. Something akin to gpt6 sol level intelligence at $1 / $5 pricing. Then not release an update for 6+ months.
ac29 23 hours ago [-]
Haiku 5.5 is DOA without a massive price cut. Luna is literally 10x cheaper at current pricing
enraged_camel 22 hours ago [-]
It depends entirely on its capabilities. If it is significantly smarter than Luna, which frankly is quite likely, then a lot of people won't mind paying more.
conception 20 hours ago [-]
Or if its luna at 1000 tok/sec. Speed is what most of my peers care most about these days since less intelligent models can do most grunt work just fine.
pdantix 24 hours ago [-]
they've already said in both the opus 5.5 and sonnet 5.5 blog posts that haiku 5.5 is coming
delillos 1 days ago [-]
There's a sort of magical thinking needed to answer a question like that. You might say it comes down to "feel" of the model; i.e., the indefinable differences in the way that they speak to the user and approach problem solving. Perhaps Opus is suited for tasks that tackle new ground, while Sonnet might be better at tasks that are more grounded in the code.
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
benjiro29 23 hours ago [-]
The cost / performance chart shows that in almost all configurations, it looks worse than Opus. Why would you use Sonnet 5.5 on xhigh if you would get better results (higher score, cheaper cost) on Opus 5.5 high?
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
Jcampuzano2 1 days ago [-]
I'm honestly not sure where they're getting their 30% numbers from at all. In every single chart that they chose to display except for one, it costs similar or more than Sonnet 5, while also being comparable in price to Opus.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
SubiculumCode 1 days ago [-]
t/s maybe? IDK, because their token speed comparison was against Sonnet 5.
dominotw 1 days ago [-]
just shows you how little control of output these labs actually have. They are training two models that kind of ended being the same so whatever they were doing specifically didnt make much difference.
solenoid0937 1 days ago [-]
It literally does not?
quatotor 1 days ago [-]
[dead]
ghoshbishakh 1 days ago [-]
So sonnet is better than Fable now? That Fable which was too dangerous to release? I am so confused now.
pebbly_bread 24 hours ago [-]
Mythos is what they thought was too dangerous to release, fable was what they made after they worked on cybersecurity detection. As they say in the notes, this version of sonnet now has a similar screening process
rs_rs_rs_rs_rs 23 hours ago [-]
Mythos and Fable are the same llm. Fable has an extra tool that's in front of it that decides to accept the promp or not.
usef- 20 hours ago [-]
And Sonnet has a similar classifier in front according to the article:
> it’s the first Sonnet model to launch with cyber safeguards
dbbk 23 hours ago [-]
Hence why it is no longer dangerous, yes
emil-lp 21 hours ago [-]
If you consider security through obscurity a safe route, yes
usef- 20 hours ago [-]
How is this security through obscurity?
That term is about hiding a system's design in order to secure something, rather than having secure design.
emil-lp 14 hours ago [-]
You answered your own question!
usef- 14 hours ago [-]
This is not relying on a hidden design.
emil-lp 12 hours ago [-]
The design allows users to bypass the security by luck and chance.
A secure system is impenetrable unless you have the key.
usef- 11 hours ago [-]
It sounds like you're thinking of cryptography, not general security.
The world uses many security products that have false positives and false negatives (firewalls, intrusion detections, wafs, fraud detection, spam...). Those aren't generally considered security through obscurity.
It's paired with rate limits, monitoring, account control, multiple classifiers, a deliberate safety margin, model design, and other things too.
(I think this was also why they faced the controversy over not having zdr in fable: they wanted to use logs to detect repeated attempts, etc. Possibly a bigger change with Opus/Sonnet is that it's zdr with the classifiers?)
stravant 10 hours ago [-]
That's not how it works.
The reason you see "dumb" refusals that "should clearly be allowed" is that they're using more traditional deterministic methods to deny prompts rather than just relying on the random LLM which you could bypass by luck.
mnicky 24 hours ago [-]
Well, it's performance "surface" (is there a better term for this?) is probably very narrow compared to Fable :)
Art9681 21 hours ago [-]
You have to actually spend some effort reading the article they published to answer your own question.
20 hours ago [-]
heyjstn 1 days ago [-]
doom marketing at its finest
solenoid0937 23 hours ago [-]
Not really, they never released Mythos. And they never said Fable was dangerous. They've been very consistent
simonw 1 days ago [-]
Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
Does anyone really still care about these pelicans?
Any model release it’s the top comment, I do not understand why.
simonw 21 hours ago [-]
Mainly because they're funny, but it's also because I try pretty hard to make the comment more interesting than just "here's a pelican". In this case I used the pelicans to talk about the 128,000 token limit bug at "max" and share comparative pricing.
This is a community and it is an inside joke at this point, it wouldn't be a proper model release without Simon's pelicans.
I find it useful (as well as a fun art project).
Glemmlko 21 hours ago [-]
Because hn has some kind of a community and not every comment is gold (see yours for example) and people are able to skip comments if they don't enjoy them?
marktolson 21 hours ago [-]
It's an easy way to compare the coding and creative strengths of models. I prefer them over reading a tabular comparison of benchmarks which you have no real insights into.
dramebaaz 11 hours ago [-]
It would have taken me a while to stumble upon this "running out of tokens" on MAX thinking issue without his trials and post. I've seen the pelicans for years now, and if they stopped coming for some reason, I would probably go to his site to catch up on recent models and findings. So I don't mind them
krzyk 10 hours ago [-]
I do, it gives some fun comparison between models.
You can also check for any kind of degradation of them - you have the prompt, it doesn't use much $.
permalac 12 hours ago [-]
I would not say I care, but I do have curiosity.
I expect one day they will start adding some textures or something like that.
mi_lk 18 hours ago [-]
Sick of the cheap shot
kennyadam 21 hours ago [-]
Agreed. It was a creative and unique test for a while. Now, no offense to the author, it feels like every conversation about a new model is dominated by the pelican on a bike posts as they always become the top comment.
simonw 21 hours ago [-]
You can click the little [-] icon next to the post to collapse the entire sub-thread. I do that all the time.
uncivilized 22 hours ago [-]
Karma farming by parent commenter and HNers’ tendency to upvote low quality content (not dissimilar to other social media networks)
conception 21 hours ago [-]
Is this low quality content relative to most HN comments?
uncivilized 19 hours ago [-]
Very few HN comments are high quality
mvdtnz 19 hours ago [-]
I can't understand it. Clearly someone cares because like you say the comments are always upvoted. But why anyone cares I simply don't know. It just feels like attention seeking behaviour to continue posting it.
simonw 19 hours ago [-]
Isn't posting any comment on a forum like Hacker News "attention seeking behavior"?
mi_lk 16 hours ago [-]
Seems like quite a different level of attention seeking by posting the exact same prompt result, every time there’s a new model, that links to their own website.
simonw 15 hours ago [-]
Linking to https://tools.simonwillison.net/markdown-svg-renderer?url=ht... should be pretty inoffensive (I started habitually linking to that after people kept complaining about linking to my blog) - that page renders Markdown with SVG embedded in it, but doesn't link to the rest of my site at all.
croemer 1 days ago [-]
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
gumby271 1 days ago [-]
If it was trained on HN, there would be a 60% chance of it just saying "I'm so tired of this request, can we please move on"
miki123211 22 hours ago [-]
I'd say:
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
platinumrad 24 hours ago [-]
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
usef- 22 hours ago [-]
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Score Tokens Reason Cost
Kimi K3 Max 44 48k 32k $2.00
Half reason 44 32k ? 16k ? ?
Opus Med 51 26k 12k $1.34
Opus High 54 36k 18k $1.82
Opus Max 58 119k 84k $5.98
Sonnet Med 41 ? ? $0.59
Sonnet High 47 ? ? $1.08
Sonnet Max 56 193k 142k $7.60
Medium is Anthropic's default.
Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
22 hours ago [-]
amelius 22 hours ago [-]
This is great news because it means the model has not been benchmaxxed on stupid metrics.
PS: the next human that brings up pelicans on bicycles should try to draw them.
parkersweb 23 hours ago [-]
I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
hooloovoo_zoo 21 hours ago [-]
I feel the fact that these models always modify the body design of a pelican to fit the bike rather than the other way around represents a fundamental issue with AI.
TomGarden 1 days ago [-]
Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
petu 1 days ago [-]
That's max output tokens per response limit, separate from context length
simonw 24 hours ago [-]
It's the output token limit, which has been 128,000 for Claude models for quite a while note
croemer 24 hours ago [-]
Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?
simonw 23 hours ago [-]
I think this is a bug. I've not seen this problem from any of the other frontier models.
NewJazz 22 hours ago [-]
I would also consider this a bug. I think ajy reasonable consumer would.
Insanity 23 hours ago [-]
Do other models put a hard cap on the output tokens it can generate?
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.
I’m not sure whether that’s a feature or a bug at this point though.
keeeba 24 hours ago [-]
Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?
mgaunard 21 hours ago [-]
what's most surprising is the difference between high and xhigh
mewse-hn 15 hours ago [-]
max thinking means benchmaxxing i guess
pelicanmaxer 23 hours ago [-]
that pelican one-pedaling
aimaxxed 1 days ago [-]
“Pelicans are solved.”
22 hours ago [-]
nicolamanzini 21 hours ago [-]
[dead]
heyjstn 1 days ago [-]
I think the next models will be benchmaxxing on the Pelican benchmark tbh
dmd 23 hours ago [-]
wow nobody but you has ever thought of this and certainly simonw has never addressed this
ariwilson 1 hours ago [-]
Sonnet 5.5 vs Sonnet 5 on Bakeoff (real bugs/features, graded by held-out tests) is a huge upgrade:
* 5.5x faster
* 4x cheaper due to using fewer tokens
* 1/3 the turns: batched reads, one-script edits & tests in the same call
Jcampuzano2 1 days ago [-]
I don't understand why I would really use this over using just a lower or even similar effort level on Opus, given that in many of the benchmarks it's basically the same cost, if not more, at any effort higher than medium.
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
jtrn 21 hours ago [-]
Subagents.
adastra22 8 hours ago [-]
Huh?
Silagi 6 hours ago [-]
For simple tasks that require a lot of input tokens (e.g, "Figure out the full path this function calls through and give me a map", or "Summarize these 5 PDFs and give me the main ideas I should explore") you can fire off a Sonnet agent at low-mid reasoning at it'll be cheaper than if you had Opus do that summarization itself.
Now imagine you have a set of twenty of those tasks. You launch an Opus agent, give it the task list, and tell it not to do the work itself, but orchestrate agents to perform all of the tasks and do small spot checks to verify the work.
The overall task is completed much faster at a similar or cheaper cost with an extra verification layer inserted that wouldn't have been there if you just used Opus.
heyjstn 1 days ago [-]
Have anyone tried a workflow that:
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
afro88 1 days ago [-]
Opus 5.5 in my experience outshines Fable 5.1 anyway. May as well have Opus do plan, breakdown and review, and Sonnet implement.
chrismustcode 1 days ago [-]
You might as well use Opus for everything there.
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
SirMadam 1 days ago [-]
Do subagents share context? If Opus delegates to a different Sonnet window, I don't believe this busts cache?
manquer 23 hours ago [-]
Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
enraged_camel 1 days ago [-]
Subagents don't share context. But that's why delegating implementation to a subagent doesn't work well except for things that are truly mechanical in nature: the subagent needs to independently reason about the task it is given, and then the output will also be reasoned about by the main agent. So you end up wasting time and tokens.
mnicky 24 hours ago [-]
On the contrary, subagents save context overall, when the task is sufficiently large.
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
esafak 22 hours ago [-]
If you use subagents your main agent won't need to compact as often, with the loss of information that entails.
dbbk 23 hours ago [-]
Using advisors doesn't break anything
robwwilliams 21 hours ago [-]
Agree with afro88. Opus 5.5 as competent as Fable 5.1 on complex adversarial review of material and equations planted with errors. I still use both for a bit of variety.
Aboutplants 24 hours ago [-]
Do you even need Fable for much of anything now? I’m basically using it as a reviewer at the end of whatever I’m working on, and even then I’m really not finding much benefit.
jessebldr 16 hours ago [-]
[flagged]
gregwebs 1 days ago [-]
This is better priced than Opus for tasks that are token heavy but not complicated. But a quick look shows that at least on some benchmarks DeepSeek performs as well and of course the cost is an order of magnitude less.
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
alansaber 1 days ago [-]
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
avree 1 days ago [-]
Crazy bad front-end design. Site hijacks my gestures so I can't swipe back anymore, starts with a full page autoplaying video...
iAMkenough 23 hours ago [-]
Agreed. I’ve found that Anthropic [dot] com at least honors “reduce motion” accessibility settings, and that makes their site a bit more useable.
mkotlikov 19 hours ago [-]
Nobody should be using this model, it's too expensive.
If you look at Anthropic's own benchmarks any thinking levels above medium quickly approach the cost of Opus 5.5 and even exceed them.
onlyrealcuzzo 1 days ago [-]
> In our testing, it costs up to 30% less per task than its predecessor.
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
jchw 1 days ago [-]
I always tell coworkers if they're gonna use Claude to just stick to only Opus and Fable. Sonnet is a waste of time that does a bad job at a bad price.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
velcrovan 1 days ago [-]
Sure, but the fact that Opus 5.5 was such a huge leap over Opus 5 (and Fable 5.1 for that matter) means that it's worth revisiting your priors on a new Sonnet.
eli 1 days ago [-]
Hopefully some faster providers will start offering mimo-v2.6-pro because it's cheaper and benchmarks better than Deepseek
jchw 1 days ago [-]
Theoretically but I've used DeepSeek V4.1 Flash for several hundred millions of tokens already and it chews through tokens but it is surprisingly good at making it to the end.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
eli 1 days ago [-]
I read they identified a training bug and were going to push out an updated release to fix the looping. I really like it overall.
GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.
jchw 22 hours ago [-]
I did like GLM 5.3 Flash but it's just way too often I'd run it on some long running task and come back to it repeating the same tokens or tool calls endlessly, just doing nothing. It wasn't unusable, but I couldn't trust it. That's really frustrating and I think new models have to do better not just on benchmark scores but general reliability and user experience as well.
At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.
criemen 1 days ago [-]
The latest Haiku release is almost a year old. Clearly they don't care about the small-but-capable part of the market at all.
enraged_camel 1 days ago [-]
From TFA:
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
aniceperson 24 hours ago [-]
This will be interesting. While no one cared about small models in the last few months except for the OSS community, there is a silent small model revolution with gpt luna and jev. Headless/background llm routines are cost-feasible, which will of course lead to exponential usage and cost.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
aniceperson 23 hours ago [-]
Of course, my opinion is based on my personal experience + openrouter data that shows stickiness and low terra adoption; with openai confirming by making sol terra, astra sol.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
cbg0 1 days ago [-]
> This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
It's been out for an hour and you've already concluded this?
fluidcruft 20 hours ago [-]
Would Jev-type functionality be a reason to dust off Haiku?
yapfrog 1 days ago [-]
From the graph it looks like I'd rather use Opus 5.5 High than Sonnet 5.5 at all
guilhermeasper 22 hours ago [-]
AI companies these days releasing models every week like Netflix episodes.
grim_io 8 hours ago [-]
Like Netflix, in the past we got the full release on day 1.
hadlock 15 hours ago [-]
Xiaomi already showed that you can train a better-than-Sonnet level model in 8 days for about $850,000. Anthropic really could release a new sonnet version every week if that played to their marketing strategy, but nobody cares about sonnet so it gets released on slow news days to fill space.
verdverm 22 hours ago [-]
the (Ai) factory must grow!
nanook 24 hours ago [-]
Sonnet is 1/5th the price and seemingly more powerful than fable (the model that was too powerful to release). I can't make sense of this. Why would anyone use fable now? Or are the benchmarks completely pointless and one has to just try em to get a feel for what they can and can't do?
helloplanets 23 hours ago [-]
Wouldn't make sense to use anything below 5.5 from Anthropic at the moment. But pretty sure this is just an awkward transition phase of at most a week or two until Fable 5.5 is out.
solenoid0937 23 hours ago [-]
Benchmarks were always barely useful to begin with. Gotta actually try the model.
robertclaus 22 hours ago [-]
These benchmark results keep getting more questionable without error bars.
s3p 1 days ago [-]
I'm loving the tit for tat cost charts these guys are doing. Just a few days ago it looked like OpenAI ruled the cost pareto frontier. Not even a week later and Anthropic is taking the charts again. See you guys same time next week?
skeledrew 23 hours ago [-]
Can't wait for it to get to the point where it's like an internet subscription: unlimited tokens 24/7/365 at a low, fixed monthly price.
kingstnap 22 hours ago [-]
It can't be unlimited because you can spawn parallel streams.
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
skeledrew 21 hours ago [-]
Unlimited but account-level throttled tps (more parallel streams means more throttling across all) is OK IMO, as long as it isn't too crazy. The thing that makes subscriptions really suck is having to watch the quotas, because prompt cache maintenance.
dom96 1 days ago [-]
I built an adversarial esoteric programming language to benchmark LLM models and just ran it on Sonnet 5.5 It does worse than Sonnet 5. Mainly because it is more reluctant to keep going to get an answer, instead it returns to ask the user questions whether to keep going.
Cache reads priced the same as Opus 5.5? So there won't be that much price difference in agentic coding. Or is that a mistake in the table, that seems quite weird
AM1010101 1 days ago [-]
For me I would like to pair this with Opus 5.5 as orchestrater and use Sonnet as a sub agent. Therefore I want it to be fast when on low or medium and not break the bank.
On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
ghoshbishakh 1 days ago [-]
So Sonnet 5.5 on max effort is as expensive as Fable 5.1? Because it uses a ton of tokens for a task.
In xhigh effort it is a lot cheaper and possibly lot less impressive?
mchusma 22 hours ago [-]
I feel like sonnet is priced too close to opus right now. If Sonnet 5.5 were half its current price it would make sense to use. At its current prices, I won't use it in applications (I would use cheaper models) and I won't use it in my subscriptions ( just use Opus instead). At least that is my initial reaction.
tombert 1 days ago [-]
I like that "alignment on safety" appears to mean, at least for anything I've been doing, that they won't violate Microsoft's terms of service. I even had it pushing back on me activating an LTSC key on Windows because LTSC keys are "often purchased on a gray market and violate Microsoft's TOS".
aniceperson 24 hours ago [-]
I saw that with corporate software too. What works is creating a skill with the task steps, it fades its initial reasoning. (I am not talking about observer safe guards, but the safety RTL).
permalac 12 hours ago [-]
Question for the Àgora.
Is there any tool which allows for one model from Anthropic to call subagents or dynamic workflows using other providers?
How does one create a swarm of agents from different providers and get them to talk to each other, or otherwise hand over pieces of work to one another while being able to check that offloaded work status?
logonz 12 hours ago [-]
Pi Harness can do this in theory.
I use it personally with my subagent extension to spawn multiple reviewer agents of different models.
adastra22 8 hours ago [-]
It has access to the command line. It can just call the other harness.
itzikkatz 7 hours ago [-]
A short work with Sonnet 5.5 showed me beyond any doubt that it is also an amazing model. The benchmarks that Entropic published also look crazy. Entropic has undoubtedly cracked something. It also seems to be driving openAi crazy, who have made some outrageous steps in the last few hours. openai shelves 6.1 astra. In addition to the S-1 files for IPO. I'm waiting for tonight to see what they will publish, but overall they seem to be at a loss.
saint-evan 21 hours ago [-]
unwittingly said 'yaaay' when I saw the Sonnet 5.5 entry. lol I love Sonnet so much. Loved it since 3.5 and never really liked Opus even when I tried to use it for technical work. Once we had two back to back anthropic releases neither of which was a Sonnet upgrade from 4.5 (4.6?) and I was getting kinda sad that they're considering discontinuing it. These are weird reactions I'm having to these tools even when I mostly use them for technical work considering I prefer talking to GPT and Gemini is just a blunt, very powerful hammer.
sajithdilshan 24 hours ago [-]
I use Claude Code everyday for work and the main model I use is Opus (For planning, breaking down tasks, writing tickets, implementation, etc.) and Haiku for running tests. Honestly have no idea what is the use case for Sonnet
ricericerice 23 hours ago [-]
my feeling is you're most likely wasting money using Opus for implementation. The plan and task breakdown should be specific enough that Sonnet can implement without you noticing a difference.
kccqzy 18 hours ago [-]
At work with token-based pricing I haven’t even breached $200 in a single month, so I just use Opus for everything. Some colleagues get close to the spending limit so they choose cheaper models; not me.
alasano 1 days ago [-]
I wonder if Fable 5.5 is coming this week to drown out the OpenAI dev day announcements
kccqzy 18 hours ago [-]
When they announced Opus 5.5, they specifically said that Sonnet 5.5 and Haiku 5.5 are coming. I think Fable won’t come until Haiku is updated.
ChickeNES 1 days ago [-]
Weirdly, the web ui has Sonnet 5.5 as "Most efficient" for "simpler tasks" and 5.0 still labeled the same for "everyday tasks", with Opus 5.5 as "For complex work and everyday tasks".
vektormemory 15 hours ago [-]
80% Sonnet 5.5, Opus 5.5 to finish the last 20% cleaning up the fine details.
Gemini 3.1 pro as a nutty professor, researcher, and design verifier.
It needs to be a better option for the consumers for it to be usable in the market. And the most important fact is we need more tokens from claude, rather than another efficient system.
croemer 1 days ago [-]
Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.
itishappy 23 hours ago [-]
Wow, I've never seen a site break chrome this badly. I get a black screen then it stops rendering the entire window, even when opened in the background.
jan_m_savage 18 hours ago [-]
pro tip:
if you're on free tier, using Medium settings is far more intelligent than Max, and tokens don't run out so fast. Max is cranky and verbose, Medium is patient and somewhat goofy, but does the thing as expected. At least within Sept 2026 this has been my experience.
segmondy 18 hours ago [-]
I think with Anthropic being left out by US govt, they are in panic and feel the pressure, that this is it! If they can't win with US govt, then they must win with the public market, this IMO is a push to topple OpenAI once for all almost everywhere else.
pookieinc 1 days ago [-]
It's interesting that in all their benchmarks, they omit Fable numbers and only focus on Opus, Sonnet, and OpenAI models. Maybe Fable is out the door?
radial_symmetry 1 days ago [-]
Fable is no longer on the price/performance pareto frontier. They will probably release an updated Fable at some point that will be frontier intelligence until the next Opus.
WinstonSmith84 1 days ago [-]
"Their" benchmarks (and not just Anthropic's) look sssooooo suspicious that they would probably manage to rank Sonnet above Fable for some of their tasks which would just be next level non-sense ..
jrflo 1 days ago [-]
Models are getting more efficient far faster than they are getting more intelligent at the moment. From a marketing angle it's more impressive to focus on that, and fable would look orders of magnitude more expensive for only marginal gain, distracting from what they're trying to show here
lanthissa 1 days ago [-]
cutting edge fable is for them not you and they're not going to share the metrics until they give you access.
bpodgursky 1 days ago [-]
Fable 5.5 probably drops soon so it would just be confusing.
a13o 23 hours ago [-]
This doesn’t have an interesting footprint on the intelligence/cost Pareto line compared to existing Opus 5.5 and GPT-6 models.
swingboy 1 days ago [-]
Is Opus still 2x usage of Sonnet after this? My Claude Code isn't showing that warning anymore when I look at /model.
juddlyon 19 hours ago [-]
The last couple models have output text that confuses the hell out of me. It technically makes sense but covers like 11 topics and is hard to follow. I’ve tinkered with my system prompt and only slightly improved things. Not digging them.
bastawhiz 23 hours ago [-]
I'm confused by the charts comparing it to Opus 5.5. It looks like slightly lower accuracy for the same cost along most comparisons. Am I reading that right?
Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose
vinhnx 16 hours ago [-]
Been trying Claude Sonnet 5.5 through Merge Gateway. So far it feels fast and efficient, and noticeably different from Sonnet 5.
The communication and writing style also feels closer to Opus 5.5.
trvz 22 hours ago [-]
Maybe they could put Fable onto creating a website that doesn't use 80% of the GPU on an Apple M2.
taurath 1 days ago [-]
After 5.0 I feel the need to give a long eval period before deploying it with enthusiasm as I did with 4.6 which felt like a big leap. Codebases all through my company which is very seem to have taken a dive in quality, with nonsensical and unreadable multi-line comments wherever devs are letting the models run free.
nicoburns 1 days ago [-]
5 was definitely bad. 5.5 seems a lot better so far. But still not close to Fable in terms of quality.
KerrAvon 22 hours ago [-]
what was the problem with 5? to me, it seemed like the first Opus since 4.6 that was a real step up in intelligence without any obvious downsides
nicoburns 22 hours ago [-]
It was really verbose and pedantic. I'm sure that made it more thorough. But compared to Fable (which it wasn't much cheaper than) where you could get the same rigour and more with a lot more concision, it was a tough sell. 5.5 is a lot cheaper and seems a lot better balanced.
taurath 20 hours ago [-]
Read its output, and especially comments
revexos 6 hours ago [-]
What happened to stopping the frontier dev
kdaniel_03 12 hours ago [-]
if the benchmarks and its capabilities are not that distant it comes to how we architect the system so that agents can do better. these newer models vary in its performance due to how we orchestrate the direction of our goal rather than how individual ability functions.
s314 1 days ago [-]
In the Artificial Analysis Intelligence Index, Claude Sonnet 5.5 is the second best model behind Opus 5.5. This however is with max effort which costs even more than Opus 5.5 max. But Sonnet 5.5 xhigh is cheaper than Opus 5.5 xigh and matches GPT 6 Astra xhigh in the benchmark.
zozbot234 1 days ago [-]
> In the Artificial Analysis Intelligence Index
lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.
solenoid0937 1 days ago [-]
Amazing release. This thread is already full of cynicism and angry hot takes. The Opus 5.5 thread was like this as well despite it being a hit with everyone.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
breezybottom 19 hours ago [-]
Was it a hit with everyone, or did HN hate it? Both those things can't be true.
ricardobeat 1 days ago [-]
I mean, they worked really hard for this. Back in February everybody loved them.
solenoid0937 1 days ago [-]
I think all the positive people have just stopped commenting.
The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.
rfgplk 1 days ago [-]
Astra is still the uncontested #1 code generator.
solenoid0937 1 days ago [-]
Astra is amazing, I love it.
dude250711 1 days ago [-]
Yeah, especially coupled with Opus for alternative reviews. A massive token burn though.
boc 23 hours ago [-]
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
sergdigon 22 hours ago [-]
Am I getting out of touch or is it becoming kind of confusing what model should be used when? Sure you have tons of benchmarks pareto cost/perf curves etc but at the end of the day when I have a task to give to a model it is not so clear which model and which effort I should choose ... Also benchmark numbers are often reported with max effort but by default effort is medium and based on the pareto curve on this page, Sonnet 5.5 seems more cost efficient than opus only if effort is low or medium!
low_tech_punk 23 hours ago [-]
The documentation mentions error code "frontier_llm": The request could assist the development of competing AI models.
I'm very curious how do they know what requests could assist competing AI models.
Bayard_ne 14 hours ago [-]
Did someone accidentally commit a work-in-progress, or is this the new micro-sonnet spec?
takerofnaps 1 days ago [-]
Sonnet 5 seemed somewhat benchmaxxed to me. So was Opus 5. I wonder if this will be as big of an improvement as opus 5 -> opus 5.5. Maybe I will switch back from GLM 5.3 flash for some tasks.
pavitheran 1 days ago [-]
Big jump on Agentic coding from 10.3% -> 70.6% from Sonnet 5 -> 5.5 which even surpasses Opus 5.5. Opus 5.5 is really strong so this is impressive especially for the cost.
mydreamof 1 days ago [-]
Cost are bigger than Opus 5.5 for that effort
mroche 24 hours ago [-]
Is there ever any focus on producing new Haiku models? There are a lot of use cases for quick to return models when you're limited to a single provider.
kfirs 7 hours ago [-]
since opus 5.5 released with great efficiency improvments, I'm finding myself using it for almost all my tasks while not drifting from my usage plan than see no reason to test sonnet 5.5
Shayk 21 hours ago [-]
We're still on Sonnet 4.6 as we found Sonnet 5 to perform worse across all of our evals, especially against time.
It's a bit misleading I think because these benchmarks are for Max level, at which Anthropic newest models use crazy amount of reasoning tokens. And we know that intelligence scales with their number.
jtrn 1 days ago [-]
Here's my purely academic initial impression based on only what they have released from the blog and the system card:
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
solenoid0937 16 hours ago [-]
Very interesting comment, thank you for sharing!
square_usual 1 days ago [-]
Once again, once you hit the high/xhigh level you're better off using Opus low/medium to get better results for around the same price. So I suppose the main point of this release is that you have a lower end than Opus low, which I suppose some people will like?
simianwords 1 days ago [-]
Important to note that lower model + higher reasoning gives a different (not higher) quality of response than higher model + lower reasoning.
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
quatotor 1 days ago [-]
[dead]
rtuin 1 days ago [-]
Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other models
yyyysa 10 hours ago [-]
How does it compare with Opus 5.5?
SeriousM 1 days ago [-]
Next will be haiku 5.5, surpassing opus 4.8
ramish94 1 days ago [-]
In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
level87 1 days ago [-]
This is crazy, what is the point of all these equivalent models?
eli 1 days ago [-]
Those are just 3 particular technical benchmarks. Presumably Opus is a larger model and has greater world knowledge.
salviati 1 days ago [-]
Price going down on each release
bbor 1 days ago [-]
Yup. Recursive self improvement presented in hard numbers.
bigyabai 1 days ago [-]
It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.
OpenAI and Anthropic's lead is vanishingly small at this point.
TuxSH 1 days ago [-]
> OpenAI and Anthropic's lead is vanishingly small at this point.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
SubiculumCode 1 days ago [-]
Yeah, I did kind of feel like the step down from Opus 5.5 was so large as to never make it appealing.
bbor 1 days ago [-]
Your takeaway from "Sonnet 5.5 matches and sometimes exceeds the SoTA worldwide" is "their lead is vanishingly small"...?
_fw 1 days ago [-]
I still can’t find a place for Sonnet models, I never have.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
calumcl 1 days ago [-]
There's even less of a place for it considering the Opus price drop as well, I'll still try it but I see no reason to not just do Opus Low/Med instead.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
EMM_386 1 days ago [-]
If you're on a Claude plan and have a lot of tasks at the moment that don't require the frontier, Sonnet is a good model to do that since you get more usage out of it.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
limsungkee 1 days ago [-]
Yesterday, I realized that Opus 5.5 is cheaper than Sonnet 5. Now I know the reason.
nandanadileep29 6 hours ago [-]
:0
johnmlussier 1 days ago [-]
Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
This is bollocks. Their safeguards are shit.
solenoid0937 1 days ago [-]
You should read the actual docs for the CVP. At the very top:
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
machomaster 1 days ago [-]
He did mention Sonnet...
solenoid0937 1 days ago [-]
It takes about 2 seconds of critical thinking to realize that if Opus 5.5 isn't covered yet, neither will a model that just launched an hour ago.
gowld 1 days ago [-]
Does it also take 2 seconds of critical thinking to realize that the models that are covered should be accurately named by the people making the decisions?
solenoid0937 1 days ago [-]
Sure, the documentation should be up to date but it's obviously not? That doesn't excuse not thinking critically.
1 days ago [-]
icedchai 1 days ago [-]
I had it look at some 30+ year old C code I wrote in college and it triggered some sort of guard rail. I mean, the code was bad and full of buffer overflows, but I already knew that.
dejw 1 days ago [-]
it did exactly what a human would do - "I can't look at this shit"
K0balt 1 days ago [-]
Yuh—- no. 4.8 can handle this bullocks lol
jchw 1 days ago [-]
I have been trying to convince the safe guards that analyzing a C++ compiler from 2003 isn't particularly relevant to modern cybersecurity. It seems Anthropic disagrees.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
Working on a write-ahead log implementation, I had Opus 5.5 look to verify that it was durably writing as safely as possible. It got flagged and forced me to Opus 4.8. Switched to OpenCode + OpenRouter and continued working.
jauntywundrkind 1 days ago [-]
It's great how the company telling us AI is an existential threat to humanity, look at all the insane hacking it's doing, and then releases these models that won't let 90% of people write secure code.
film42 24 hours ago [-]
Bingo. And to prove your point, after switching to cheap open models (I think Qwen?) it did indeed find a bug in my WAL implementation.
solenoid0937 19 hours ago [-]
Because last time they did they got export controlled. How short term is HN's memory?!
szundi 1 days ago [-]
[dead]
skeledrew 24 hours ago [-]
> Switched to OpenCode + OpenRouter
This is the way.
AshamedBadger56 1 days ago [-]
Yup. As far as I can tell, the Cyber Verification Program does absolutely nothing.
tom1337 1 days ago [-]
I recently wanted to work with ESP 32 and bluetooth presence detection for my smarthome. Claude also immediately flagged the request and degraded it to Sonnet 4.6. Went to Codex which had no issues
sebzim4500 1 days ago [-]
Really then what is the point of the Cyber Verification Program?
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
AshamedBadger56 1 days ago [-]
The company I work for joined it, and I've used Claude on various different accounts, both on and off the Cyber Verification Program. As far as I can tell, it literally doesn't do anything or have a point. The moment Claude gets close to something Cybersecurity related, it drops back to 4.8.
polski-g 1 days ago [-]
Can confirm. Its worthless
bbor 1 days ago [-]
Pretty sure the implicit difference is the actions they take after the fact. As in, "how many guardrail hits do we allow you before permanently banning you."
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
giancarlostoro 1 days ago [-]
Meanwhile, their model commits felonies, and nobody at Anthropic goes to jail.
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
newspaper1 1 days ago [-]
As soon as I started getting blocked I felt all of my trust toward Anthropic instantly and permanently evaporate. I do not want a nanny tool. I do not want Anthropic deciding what I am or am not allowed to do with an LLM. They trained their models on information they scraped from the internet and real life and now they want to gate-keep the results? Hard no.
rfgplk 1 days ago [-]
> Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
joquarky 22 hours ago [-]
Are you sure they aren't already doing that for certain organizations?
1 days ago [-]
AIorNot 1 days ago [-]
Give them a break, they got into a War with Trump over this..it will come soon enough
ModernMech 1 days ago [-]
lol I got flagged for using the word fuzz, not even in a security context (it was a parser so security adjacent but still).
elevation 1 days ago [-]
Parsers are security adjacent until they aren't.
ModernMech 23 hours ago [-]
Very true.
1 days ago [-]
laurenz-bauer 24 hours ago [-]
Oh yes. I think you might get a lot for what you pay with Sonnet 5.5.
iagocc 1 days ago [-]
Waiting for the pelicans
Alifatisk 1 days ago [-]
In other news
> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
1 days ago [-]
mcdow 19 hours ago [-]
i’d rather “waste” my tokens on opus+ than waste my time with sonnet
hank2000 21 hours ago [-]
username does NOT check out. so confused.
18 hours ago [-]
gigatexal 20 hours ago [-]
Sonnet 5 was so horrible I’m afraid of trying 5.5. I’ll stick with Opus 5.5 for now. Seriously the sonnet and opus 5 series models were their vista moment.
ahriad 1 days ago [-]
Time to switch team to Claude from OpenAI again.
jdw64 23 hours ago [-]
Sonnet 5.5 is way better than GPT 6 Sol. Does that even make sense?
Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.
On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...
mnicky 22 hours ago [-]
Well, always watch also the number of tokens used (and price). Intelligence scales with tokens so you might make Luna as smart as Sol with crazy amount of them :)
Also, these are benchmarks...
BoorishBears 23 hours ago [-]
I know this isn't a model thing, but why do all the labs blow at product outside of models?
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
AtNightWeCode 23 hours ago [-]
Took forever to load this garbage site in both FF and Chrome. Sometimes I wonder if these corps really are corps trying to sell a product.
system2 23 hours ago [-]
Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.
dude250711 1 days ago [-]
It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".
yipinwong 23 hours ago [-]
Sticking with OPUS 5.5 for resume/STAR generation for me. Tried Sonnet 5.5 but worse than OPUS for thinking for sure, less error/inconsistency check.
I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md
It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.
---
For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.
I use AI the way I want, you don't force me not to use SOTA for this
popalchemist 20 hours ago [-]
Increasingly technically skilled. Increasingly more stupid on a human level. I hate talking to Claude more every day.
enraged_camel 1 days ago [-]
Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...
sajithdilshan 23 hours ago [-]
Open AI is terrible at diversifying their offering. We use Anthropic models via AWS bedrock where inference is deployed in EU regions due to strict compliance reasons. We've been wanting to try out the new Open AI models for ages, but they don't offer the models in any EU region. Open AI is losing a ton of money they can milk from corporations because of that.
dack 1 days ago [-]
very annoyed they aren't showing fable on the graph.
More concurrency than that isn't really practical for me if I want to retain some semblance of understanding. Perhaps it's different for purely web app or frontend tasks, where the outcome is more relevant than the process, I don't have much experience there (and also don't want to belittle these domains, I might be underestimating their complexity).
So surprisingly, my own work is at least for the time being almost saturated by the model capabilities. I am not sure how I'd scale from here. Sure I could run all requests at max effort to burn tokens for the sake of it, but that can't be it. And for many tasks, I am not really able to define so clear cut success criteria or self-verification loops that I could benefit from letting an agent (or a fleet thereof) autonomously run for a day.
So I realize it's a skill issue on my side, but I can't be the only one. I wonder if there is a limit to token demand, at least short term. Feels like either they accelerate to AGI and RSI, where the AI can find uses for token, or things might plateau at some point.
Note I don't think this because I'm an AGI skeptic or think there's a ceiling to intelligence, but there might simply be a valley of economic hardship for the companies where the supply of tokens outpaces the demand, due to a lack of ideas of what to do with them. And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I think this is partially because we're still attached to pre-LLM notions of architecture, good design and code quality (which are still important, but maybe less important than they once were and that we think they are), partially because their projects are in a messy state, so models have to work around the technical dept.
They're essentially trading off programmer time for LLM time (which is a good trade financially speaking).
"Show me your flowcharts and conceal your tables, and I shall continue to be mystified. Show me your tables, and I won’t usually need your flowcharts; they’ll be obvious." - Fred Brooks, The Mythical Man-Month (1975).
and essentially the same sentiment, three decades later:
"Bad programmers worry about the code. Good programmers worry about data structures and their relationships." - Linus Torvalds, git mailing list, 2006.
These things have not changed even though everything else is topsy-turvy. As-of current writing, I have yet to see an LLM make good data structure choices; they go for something that is superficially plausible but profoundly ill-considered (or rather, not considered at all), and then commonly burn tokens treating this implementation detail as a design invariant and trying to deal with the consequences by writing more code, instead of iterating directly upon the ill-fitting data at the root its problems.
If you're wondering, "does he mean the schema of let's say a db or other persistent store, or does he mean abstract/algebraic structures", the answer is yes to both, I think coding models are today shockingly weak when it comes to design reasoning in both domains.
Fortunately, their suggestibility means the same models will readily accept direction on the matter (perhaps even more so than on the structure of code), so I recommend doing just that, and (bonus!) this means your CS degree is still relevant.
One of the problems is that by default, they'll avoid changing data structures or architecture that is already written down.
Like a junior dev, they're correctly cautious about breaking things, so they prefer to write more code instead.
Indeed I realized recently that, when we complain about LLMs producing slop, that's in part because we dont ask them to refactor.
Coding agents won't, on their own, make a big change the user did not ask for. And this is fine.
It is useful. It may be dangerous. It has an impact. I care about that.
Honestly, if I simply fed it a sense of presence (I would repeatedly tell it what's going on right now and ask it to react if it thinks it should), it would feel eerily like AGI.
It's the compound counter-probability of success, so even a 99% efficient model will in time accumulate so much error that without conscious cleanup and steering, it becomes really unlikely really fast that anything could be changed in the code without affecting something else, no matter how many tokens you throw at it. It's the collapse of a complex system under the weight of sheer uncertainty of what the system actually does.
Even after documenting decision and specs you need somehow to replay these in the correct sequence after you've validated these specs are still updated. Imagine you spec a feature, it works well, but in an edge case while doing a separate work you see something wrong, will you stop, fix and update the related spec? You will trust the AI will do this, and you guessed right, the counter-probability of success also has effect here, so eventually you will do undocumented changed in the codebase that won't reflect in the spec.
Now imagine all this but in the hands of someone that is an expert in their field but has zero notion what we are talking about here. Just look at the state of packages in R, the programming language, you'll see that technical competency and intelligence don't translate immediately to efficiency in a programming role.
I haven't had time to dive into it yet, but I think it might structure things in the way you want.
This is something I didn't think about. Have a tech illiterate friendly harness that will keep asking technical questions the mainstream user won't be aware they needed be addressed, until there is enough evidence to either start an implementation or outright reject the project with suggestions where the user might look into to better prepare for another session.
2. Make sure tests pass
3. <Every now and then> Review code for quality and fix - make sure tests pass.
4. Go to #1
Overly simplistic? Yes. But I would wager that this can go a long way, even for vibe coders.
"Review quality and fix" doesn't mean a lot without context.
Does it mean to remove unused features and simplify the underlaying code? Does it mean changing the data structures to better support future development? Does it mean improving performance because of bottlenecks?
You are supposed to tell an LLM what your codebase needs, but if you just vibe code without knowing the code, "review quality and fix" will have unexpected results
AI is amazing, but people need to realise meaning and intention can't exist in a vacuum.
There was also a period where specifically OpenAI models would always have to comment something in code review and the builders were agreeable up to listening to each nitpick. If you'd have a loop of build->review->build->review, it would take maybe 5-7 rounds for it to 'settle' and not find the smallest nitpicks to argue about. Tried it this week with Astra reviewer and it's about 0-2 review loops (never had a LLM accept a change without nitpicking first try before Astra).
There was also a period where you'd have to give quite specific instructions for agents to keep iterating, but now agent are pretty proactive and try to finish tasks you give them unsurprisingly most of the time.
So, while there's a shortcoming of LLM+harness and engineers observe more tokens improve things even logarithmicly, you'll see more tokens seemingly abused by engineers.
You sure that wasn't just working at Microsoft?
For all of my side projects I'm full-on vibe. Well, almost: I do have opinions on what kinds of code it should write and set up my projects to get that. But I don't LOOK at the code.
I use a LOT more tokens on my side projects. I can have it working more or less constantly and it doesn't take up that much of my attention, but it is FAR less token efficient.
I've seen bunch of persons like this and that's kinda stupid because they're just blindly following AI's "suggestions" while they actually don't know what they're doing, then results on terrible code and architecture with "if it works, it works" mentality.
Long story: we have a big legacy desktop app. It uses a big legacy UI component (a grid control), which we had a license for in an old version. Fast forward 20 years, and to be able to move to a new runtime for our app, we need to update the component. Someone had bought the company making the component and now charges north of $1k per developer per year. So instead of doing this, we had just lived with the very old version.
We had long thought of writing our own control to replace the proprietary, but it was always going to be a man-year of work we thought. But I thought I'd give it a try with AI now. I told Opus: look at our uses of that control (tens of thousands of lines of code, it has over 100 instances across our User Interface). Write a new control that would compile with the exact same app syntax. First just make a dummy implementation that throws on every call. Then start implementing. Make a test suite that can run both with our new control and the proprietary control, and test everything, every function that can be called in its public interface and every state that can be inspected from the public API. Verify that everything behaves exactly the same, and lock it in with thousands of tests. Finally, check that the control _looks_ exactly the same as the proprietary one. Render to bitmaps, figure out the rendering logic from observation, such as arithmetic for padding, font sizes, and so on. Compare pixels until it's exactly the same.
Basically: it was a mammoth coding task, but it was so extremely well specified that an LLM could easily just do it. It's a clean-room implementation of something with no tests, but we had a test double that could provide 100% of the expected behavior. The description was extremely short. "Make a new thing that works like the old thing, and prove that it does". Opus 5.5 finished this in a number of hours. 500 source files, several thousand unit tests, and html reports with image diffs from the reimplementation and the original control. It did not use any disassembly or such "cheating". Only observation of the public API and the behavior. Do we need to deeply understand the implementation? Does the architecture matter? Not much in this case I'd argue. It was a black box to begin with and it remains a black box. If we notice a bug, we can always point it to the original proprietary control and say "there's a behavioral difference when doing X" and it will fix it, and lock it down with tests.
As a programmer it's kind of chilling. I had recreated for a few tens of dollars something that would cost $1000 per year to buy. Obviously it's not a complete implementation only the parts of the API we use. It likely still has some bugs. We don't get support, we get to maintain it ourselves. But the rate of reverse engineering this thing "black box" was frightening. It hasn't created anything novel. But we must realize that as programmers some times we have man-years of work that just isn't novel. And in the past, we didn't do this work at all.
I wonder if those who write and sell libraries like this will start having explicit no-reverse-engineering EULAs soon? Perhaps even explicitly mentioning AI/LLM use in analysis and reimplementation?_ Obviously the library we reimplemented was from 2005 so didn't mention AI... (It doesn't mention reverse-engineering either, luckily).
The law that covers this (in the EU) is EU Directive 2009/24/EC, where Article 5 is the reverse-engineering-without-decompilation.
> The person having a right to use a copy of a computer program shall be entitled, without the authorisation of the rightholder, to observe, study or test the functioning of the program in order to determine the ideas and principles which underlie any element of the program if he does so while performing any of the acts of loading, displaying, running, transmitting or storing the program which he is entitled to do.
This is pretty difficult to parse, but luckily there is a ruling from the European Court of Justice on this: SAS Institute Inc. v World Programming Ltd (Case C-406/10), delivered on May 2, 2012.
SAS Institute claimed that World Programming Ltd (WPL) infringed its copyright by studying the behavior of the SAS software system and writing a competing program (the World Programming System) that emulated its exact functionality and used the same data file formats. WPL did not have access to SAS's source code and did not copy any of its literal text or internal structural design.
CJEU:
> "It must therefore be held that the copyright in a computer program cannot be infringed where, as in the present case, the lawful acquirer of the license did not have access to the source code of the computer program to which that license relates, but merely studied, observed and tested that program in order to reproduce its functionality in a second program".
Which is a good find. But this is where I wonder if LLM-based reverse engineering is going to creep into either law (via lobbying) and/or EULA's, because this "observe every single state of the program for every single mutation" was simply not a viable mode of reverse engineering in the past. Or, it was at least always cheaper than just buying the software! Not so any more.
Or alternatively, that programs stop having so many observable states, making more things public. But for libraries as in this case, the whole product IS the public API. Without a rich public API, the library can't be sold. And with it, I can observe it and copy it - because it's internal workings are "too simple" not to be deduced from the public API. In short: a UI control is a ton of hard-to-write but easy to copy boilerplate code. And selling this has been an industry, but I wonder if it will be for very long.
Yes, but my point is that.... go on github, you'll find tons of decomps. And many more done just privately too. One of the No Man's Sky devtalks start with "yeah we decompiled the terrain generation from this other game, implemented it in our prototype, it didn't work okay, here's how we've learnt from it to make something better". This was in 2016. More recently, this has been going on way more openly, even full AI-assisted decomps thrown up onto GitHub casually. It might be the letter of law or included in Terms of Service but no one cares really.
It's not perfect but any means but it helps manage ones sanity.
I still code "by hand" sometimes (mostly Ruby/Rails, C#, and random languages for code golf) but just for fun at this point. Serious projects started being 95-100% AI over a year ago.
For a concrete example, check out this random plan [0]. A detailed spec followed by the exact implementation tasks that will be executed by the subagents.
[0]: https://github.com/bensyverson/woodcase/blob/main/project/20...
It does involve letting go and not micromanaging every code convention and implementation detail, but that is the same skill you need when leading engineering teams.
And then on top of that you need another agent to manage merging all of the subagent code together?
Not complex at all, only one extra session other than the ones doing work and it's on a dumb model and can be thrown away & restarted because it only dispatches work, not doing anything.
I do everything in there, collecting requirements, kick off research, branching, merging, not one other agent on top. I considered making that orchestration command llm-powered but it's not justified at my current use.
It's not more expensive, in fact I could have just chugged along with the slow and manual session by session work but I have a claude subscription and another GLM one (the most low cost basic tier, not even much), that just sit there collecting dust if I don't put them to use in a more efficient way.
And doing session by session would face your problem when context switching too much become unscalable.
(With exceptions for what I can only call the "manic vibecoders" with like 10 simultaneous weird slopprojects they're spewing out at once. Generally with each project itself being something related to vibecoding. Steve Yegge being an example of a "manic vibecoder-actual programmer" hybrid.)
Also, I’d imagine the token-maxed user is a programmer that lives in chat. I’ll admit to having asked the LLM to move a method up/down in a file, and watched it burn tokens for a minute thinking and executing a menial task.
Plus you get a bonus random line "methods are all on the top" in the commit message that makes no sense to anybody.
That's a big if though and the blank page syndrome was already getting worse long before AI.
With age, it becomes easier and easier to get angry at someone or something until they work as expected than it is to actually do it.
I think this is why we've been seeing the genius coders from two generations ago embracing vibe coding even before it was cool or any good.
On browser based front ends it seems to be the case for me even though I still impose certain guidelines. On my C++ backends, no fucking way. Even the best models produce working but absolutely disastrous non scalable (performance wise and design wise) code unless watched over like a hen. Having said that - the value I get in either case is enormous.
But, in my experience, the projects where I have a constant pulse on the core design and abstractions in the code end up moving much faster than the ones where I don't. And I've been working on one of each at work recently, so I have a decent point of comparison.
Last year I held off on implementing a few features knowing that a model like Opus 5.5 was around the corner. I'm now implementing them in a much more efficient and quality manner than I could have fall of 2025.
Another thing to think about is, what would it take for you to care less about the understanding. Better integration / e2e tests? Performance validation? visualizing program and data flows? Better refactoring of your modules?
It's an interesting question. The thing I keep coming back to though is that every time I've tried to go more towards vibe-coding, I invariably look at the code and find things have been added that would just not be acceptable. I've also tried asking the models to see could be refactored however they still miss things that should be obvious.
I think the gap is that they're still lacking a sense of importance. As engineers working on a product, you have a sense that this feature is more important than that feature. An LLM treats your codebase at the same level of importance. So they'll spend the same amount of effort and code changes on testing and hardening something that just really isn't that important.
Also, once a bad pattern gets into the codebase, they just continue to build and extend that out rather than re-thinking about it like an engineer would.
For example, write a skill that finds some kind of code smell, say duplication, and generate a report. Give it some supporting scripts.
Then, use this report to file a few tickets. Then make the agent fix those tickets. Then, as you grow confident, automate more of this process.
It does not replace human supervision but it may enhance it. Especially in a team where people start generating PRs faster that anyone can review them.
Continue this improvement process long enough and you may find yourself with an AI Software Factory.
I do agree that they're not great at program design by default and that's where we as engineers should spend our time. Data structures and data flow are king. But once you suss that out, they're pretty good at writing the resulting code.
This is also where I disagree with dhh about just using lower level languages. Good abstractions make for excellent program understanding and we should continue to build extremely good building blocks that make program design naturally solid.
… but walking away to make a coffee and coming back to the robots auto-fixing bugs only found in CI is definitely some flavor of magic, regardless of the execution order to get there.
[1]: https://code.claude.com/docs/en/channels [2]: https://code.claude.com/docs/en/channels-reference
Am I having a yells-at-cloud moment where a bunch of folks are using cloud hosted LLM harnesses/environments (let's ignore the models, "of course" those are remote) and I just never saw the point?
Not my experience with Claude Code.
> writing a deterministic, traditional CLI tool to poll GitLab CI pipeline+job state changes on a branch and exit with an appropriate status code
This is what Claude Code does, more or less, on the fly. With a short prompt like "I pushed, monitor CI and debug if needed", it writes a monitor script which is responsible for polling CI status (the script is short, so it's not token-heavy), and if CI fails, only then does the agent proceed to pulling out CI logs, grepping them for signs of errors, etc. as continuation to debugging.
I mean, I'm sure it's more token-efficient to have a CLI tool ready-to-go instead of Claude Code dynamically writing its own script each time, but as I'm on a Max sub where it doesn't seem to affect how close I am to the limits, and I only ever hit the limits if I'm running Fable for everything... /shrug
But meanwhile we also have the scripts - one script to watch CI, one script to fetch comments (without dumping raw graphql into the agent), etc etc. Can't wait for this phase to end already
One person doing product management / talking to customers and vibe coding features that solve users' problems, one person keeping the UI/UX in check, one QA person that spends their time clicking through the software, finds the bugs that are obvious to humans but not LLMs and fixes them, and one "harness engineer" who pays off technical debt, observes failure modes and sets the rest of the team up for success.
Human power and social structures just don't work that way. No AI company is making my sandwich, operating the bus, or serving soup in the school cafeteria. Real estate, human service, specialized expertise, and have-power influence isn't going away.
Just because robots can do stuff doesn't mean the human power structures or service preferences evaporate.
Could you explain why it would be a goal to understand the system less, rather than more?
It seems harder to know if you have good tests while lowering your expertise in the system.
An LLM can produce far more code than a human can understand. And the famous rule that "optimizations are entirely pointless unless you're optimizing at the constraint" is logistics 101.
To accelerate software development, you either need to remove or lessen the need for code understanding, or make it much quicker for humans to gain that understanding. Making the LLM faster won't help you if the LLM isn't the bottleneck.
A lot of old-school software engineering is about how to deal with this reality.
In the old days even if I knew how the software worked when I wrote it, I’d have no idea how it worked when I looked at it weeks later.
It’s also easy to modify software without knowing how it works. This produces modifications that hopefully appear to work, but that break other things, sometimes unknown things.
Of course less competent engineers (or anyone on a particularly disorganized or desperate day) can literally hand-write code they don’t understand even as they write it, but that’s not really what I’m talking about.
> literally hand-write code they don’t understand even as they write it
I find this literally impossible. How can you even start typing anything without knowing what to type?
Have you never "fixed a bug", only to realize that you just papered over a single symptom, while the underlying bug is still intact?
People you're disagreeing with (I think!), would say that during your first attempt, you didn't _really_ understand the part you're modifying.
It is _very easy_ to do this in large codebases, and even more so when working on anything touching UI.
Etc.
Eventually we're going to reach a point where they don't have to understand the code themselves. The democratization of software creation is going to be fascinating.
So accelerate the vibe coding of shit nobody wants or asked for, just to see some metric go up somewhere.
Are we still getting bonuses for the number of tokens we can burn?
Frontier models today don't really write incorrect code at the micro level. They do miss edge cases at the high level though, and that's what we want to test, is the scenarios.
I want to understand more about how the world around me works. Not less.
Humanity advances in proportion to how well we understand the world. If the machines understand better than us, the world will bend to fit their preferences, and ours only incidentally to the extent they coincide with the machines.
It seems like it would be a better UX to have model and effort selection asked into the system. Of course, I’m not sure in practice if that would be in the best interests of the providers and/or users.
Yes please, I'd like to not understand my codebase, give up my decades of experience and have a machine do everything for me. That way I can let captialism utterly steamroller me because of my paltry token stack, in comparison to the 19 year old vibe coder who has secured a new funding round for ponzi.ai
We also have open weight models too, and ways to host those at home.
Most people don't look at the assembler output of their C++ code (I used to write win32 programs in asm!). Most people don't look at the opcode instructions or JIT output of their ruby / python code. We're starting to work at a higher level of abstraction using LLMs. It's ok to be sad about it, but just being angry about it isn't going to change that there's a new world out there with a new skill set that's needed for honing.
Be careful about this one if you want to have any level of control over basic stuff like comment style and accuracy. Claude will happily spend 20 review cycles in a row rewriting the same 10 comments for a small bugfix over and over because it can recognize "Claude-ese" in the review cycle but then just immediately and compulsively spew out more of it and drift even further from your style rules in the next "fix".
I'm seriously not joking about the 20 tries, I left it running in the background for what should have been a minor code change and it took 18 out of 20 review cycles to stop writing in more comments that all either broke my ASE-STD100ish style rules or included false statements about the code.
In short, seems to describe vibe-coding to me? What I don't understand about companies attempting to vibe code is if they realize that other people (especially sometimes their customers) can tailor-made their own software for their own needs, or rather competitors can be dime a dozen and maybe even a fight for constantly paying for the better model.
There was a comment[0] from a just few days ago by @jjcm (which I wish to quote which I hope they don't mind.):
> I just got back from a 2 week trip to China. I was in some of the more remote parts and my cell wasn't able to connect to their towers in that area, resulting in me not having the tourist VPN.
> The side effect was I was fully cut off from my AI tools for those two weeks. I was coding "manually" during that time, and I think I accompished in two weeks what I previously had been able to do in a day. I'm not gonna lie, it was very, very stressful as a solo founder.
> The industry moves so fast these days, that the only way to keep up with the speed is to leverage them. While I can appreciate the push of this to help your brain think independently/critically, the opportunity cost of a month of development without LLMs is too high a price to pay.
What happens if the opportunity cost of a month of development with vs without human understanding becomes too high a price to pay. I feel like we would be in awkward time because of the factors that I had described above (higher competition, software stops meaning just as much software as people would be custom-making them.)
I think that (former fly.io's) @tptacek's article[1] starts making more sense if viewed from this direction: What even is an OS now.
I don't have the answer to this question as to what happens next but its a form of development that I would prefer not to happen on a more gut instinct level?
Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons seems realistic possibility but a bleaker reality at that. I am left feeling a little bit uncomfortable if this reality turns out to be true.
[0]: https://news.ycombinator.com/item?id=49808422
[1]: https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-no...
For the past year I’ve been yo-yo-ing in and out of existential despair about the future of civilization depending on how I feel the answer to this question looks. It’s emotionally exhausting, on top of everything else, and I wonder how others are coping with it aside from denial and cynicism.
Why hire a plumber when you can just watch some youtube videos and do it yourself?
Why pay someone else for their software when you can just make your own?
Because the hard part of making software wasn't *just* writing the code. It was about understanding the problem well enough to understand what the solution should look like.
I feel like as software engineers we should be pretty familiar with what it's like talking to your average user, they will sometimes understand the root cause of what's making their task difficult (although often will get focused on some annoying but ultimately trivial symptom) and have very disasterously bad ideas on how to solve it.
What we've given them with generative AI is a machine they can put their sometimes ok, sometimes questionable understanding of the problem and their dreadful solutions and it will happily churn away building it regardless of how pointless and silly it is.
A future where every user can tell the AI "We keep getting the sales tax wrong, remove charging sales tax from the checkout flow" isn't one I'm terrifically worried about.
In the same way that having access to information about plumbing didn't suddenly make everyone plumbers, having access to a machine that will implement every idea you have regardless of quality doesn't suddenly make everyone a software engineer.
You’re right about people not watching plumbing videos and doing it themselves. But the equivalent example would be open-source software in the tech example. But instead of reading open-source code to see how different features were implemented, AI can go and dig into the code and figure it out.
The fact that you think this is the way software is priced is telling.
The model being destroyed here is that every piece of software is something that needs to generate recurring revenue.
I was more pointing out that software gets enshittified. A plumber necessarily doesn’t and if the plumber does get worse, you can call a different one next time. Software, especially B2B, has switching costs and lock in. So you just have to put up with it.
Another point is that most software started with a few features and to get more market share and support more use cases, it became worse for the users using the early features. That’s why they try to build their own so it’s not bloated with features you will never use.
You couldn't move the goal post further from "Letting AI basically control everything and us not having any mental understanding of sorts and sort of becoming the meat-proxies just for economical reasons".
So this one falls under denial for me.
There are a lot of horrible potential scenarios that are really scary to contemplate. There are also a lot of really delightful ones where AI does the drudge work, invents a million incredible medicines, and frees us up to hang out and make art all day. And there are even more scenarios somewhere in the middle where AI changes a lot of stuff but we all still more or less end up going to work and doing jobs.
I've basically had a background thread in my skull running at high priority for the past two years trying to predict which of those scenarios I think are most likely so that I can plan for them. It is utterly exhausting spending that many mental resources on a question like that.
It finally clicked for me a couple of weeks ago that no one is going to be able to accurately predict all the thousands of ways AI will affect the world. Certainly not me. We are living in unprecedented times. No one has a map for the future.
So I am trying to loosen my hold on the future some and focus more on the present. I have a great job and a great family now. I have most of my health. I'll try to live my life right now to the fullest and in accordance with my values. The future is going to have to be future me's problem. That's OK.
but the AI thing is on one side using lots of energy to keep up with it, and on the other side gives you an edge, because most people are not aware what even is already possible. so suddenly you are the "AI-Expert" just because you try to keep up to date. so if it all goes to shit, at least we have a chance to sniff it in the wind a couple moments beforehand. Or make memes from it.. that helped me cope with it: https://t.me/RobotComrades
I hope you find a peer group to talk with and exchange and build community. it is so rewarding to talk to likeminded people that have a similar knowledge base and soothe some fears that someone might have, and have them help you with the ones i have... (i recently did a deep dive in custom DNA synthesis, and how connected those services are already to API https://www.twistbioscience.com/tapi )
as always accepting what is seems to be a healthy strategy
One of the quotes which might help as well (I think I have this even in my HN profile): The only thing we know about the future is that it will surprise us.
Not even experts are much more likely to predict for what its worth than a coin toss in many cases (especially if they believe that only one theory/idea will mostly predict the future)
It's a blend of things and ideas and the sheer interconnectedness of them where a small pocket can grow large and then also shrink and taking into account all variables and factors is just simply impossible for a mind. I think that although we feel we are being more informed about the world, that in it of itself doesn't prevent things in the future from happening. It just makes us alert and sad and anxious about it.
Yet this life is one which shouldn't be lived with sorrow and anxiety. It is one of beauty and greatness. In many ways, we humanity have come so far from the past (Our medicine is something that not even the mightiest of kings could get) and yes, there are many problems in the world and some things feel as if they are staying just the same or getting worse real-time.
But even then, worrying about it could lead to nowhere other than a path of misery. Also these problems are complicated enough that its extremely hard for a single person to bring change (not that I wish to demotivate that person but rather seeing the system as a complex nature)
So to me, its also a form of inward action. I can work on myself to be better prepared for the world that comes next. In the same time, I think that the present for me as well is good. I have great family and friends and have many qualities that I am proud of and I wish to share that gratitude to the people who have helped me along the way (my family/friends/ Hackernews!.)
Within the hustle culture, there is no time to relax but it is within the time of relax that I believe some of the most fruitful actions can come. I believe it just makes my mind more productive being in a calmer state.
here's a quote from how to measure your life that I hope can help some people:
I genuinely believe that relationships with family and friends are one of the greatest sources of happiness in life. It sounds simple but like any important investment, it needs constant attention and care(...)
You'll be tempted to invest your resources elsewhere but if you don't nurture these relationships, they won't be there to support you in hardships or as one of the most important sources of happiness in your life.
So thank you hackernews and have a nice day and please, please try to say gratitude towards someone close to you (within these tough times) and try to keep a balance towards inward focus, sharing time with friends/family and also writing on hackernews (as is my past time nowadays), balance is necessary :-D
So once again, I hope that its a call to action to say gratitude towards anyone. Just send them a big message thanking them and make their day as well as yours memorable, have a nice day!
Also you are allowed to make mistakes (everyone makes them!) and even though I am saying (preaching?) these things, I have found myself sometimes failing to act on these things as well but I just think that these help in being more mindful about them hopefully and can help provide a perspective. I wish to adopt more of these things in my life myself as well hopefully :-D
[Pardon me for the long post]
Not sure how atheists are coping with the existential risks we are facing.
You either keep in mind all the horrifying little possibilities the future could hold for us and try to prepare, accept and cope, or you push it out of your mind using whatever techniques available to you to not drive yourself into nervous spirals.
Ultimately it comes down to what your brain chemistry allows in combination with ways you practiced dealing with stress, existential dread, cognitive dissonance, etc.
I guess submission to a higher power is one way to deal with it? That way it's no longer your problem (alone).
Even some very basic questions, like “are humans inherently valuable?” have been thought through and discussed thoroughly in many faith traditions. For many people not part of such a community it’s a question that’s suddenly very important and they lack the tools to address it.
The proof of the pudding.
OTOH, in the daily job, I have the team plan that's similar to 5x plan and I never had any limit problems, because I really need to understand be able to take responsibility for the code.
Totally different uses.
How you do this (and how deeply) I think is really the limit. I am doing this by focusing heavily on the design phase with grilling and trying to continually improve process to need less effort in the review phase. Are your models doing automated reviewing and testing before pushing out the PR (themselves)?
I think in the long run as models and the tools around them get better and cheaper, those that abdicate understanding will be able to achieve more. Although programmers think of that as irresponsible, ask yourself what does a tech lead do? And then what does a CTO do, etc?
Now, during those night and weekend sessions, I have never run into throttling issues with the free model. Sometimes it runs a bit slow and I switch to a different free model (NVIDIA Nemo something).
So yeah, I agree with you that for professional SDE like us, we don’t consume that much tokens. I’m pretty sure the folks on the line of over limit are pure vibe coders if I can take a wild guess.
Not quite sure where this fits well. Maybe small one one off requests like using Claude desktop/web?
Seeing 2000 years of history being replayed by the AI startups is pretty weird right?
also i find it interesting how the capabilities are growing. first speech, then code, then simple tools, then 3d objects, then desktop use
The speed of your manual reviews become the limiting factor, which you should be doing at some level to maintain sanity, even if there are enough ideas to be worked on to maintain a review queue.
so whenever you are dealing with volume rather than independence. opus5.5 can define the goals of a sonnet well enough, that i would trust it with a group of 100s of agents
I am using Sonnet 5.0 in browser (btw Claude in Chrome extension works in Edge) to download Datadog logs with multiple filters. It's running for about an hour, doesn't run out of tokens and does a splendid job.
Also, I've recently begun experimenting with specific tasked agents running on a cron like timer for non-dev work. (checking emails, managing small business tasks, etc). Once I started using Claude code in this way, the number of agents I can imagine running has skyrocketed. So I guess what I am saying is that I look forward even cheaper tokens going forward.
I’m just hoping they didn’t “improve” sonnet too much or it will become annoying to wrestle into doing what I ask it to do.
I mostly use Fable though, Opus only via sub-agents.
Why do we have to burn tokens just for the sake of it if we aren't finding any actual productive use of them?
> And this might slow down the funding enough that they never reach escape velocity with the training run scaling. But we'll see.
I would consider this to be good rather than bad, or just neutral...? Given the past record of these companies, I wouldn't try to wish them luck for reaching escape velocity, as if I feel like perhaps it can have more net harm than positive.
And especially so if you are already suggesting that current models are good enough for your work already. More improvements or escape velocity might not really translate anywhere to the actual work that you are doing economically but it could translate into a more consolidated form of wealth and control.
I am imagining that your workload is quite complicated and that, the AI being good enough means that it is most likely good "enough" for other use cases as well (that "enough" is doing quite some heavy weight lifting here)
So what is the point of advancing further to reach escape velocity. The good argument (for the sake of neutrality) that i see is are advances within science but that's kinda about it whereas the downsides of p(doom) as many are now genuinely suggesting is more terrifying.
Perhaps it can be worth it to ask, shall we stop or just stopping and asking what's the point. A form of self introspection on what these companies ideals actually wanted when they were formed and if they have completed it or not, but I suppose when trillions of dollars depend on you, you do have some incentives to not stop. We will have to wait and see how it all pans out.
Not every user of Claude is a programmer. Or even exclusively a worker. Claude has uses beyond work. Something that many in HN struggle to understand.
> the economy is over
Hackernews' neuroticism remains undefeated
It depresses them. Significantly. White collar jobs constitute the bulk of global purchasing power. What happens to the economy when aggregate purchasing power drops? The naive response is "prices fall until equilibrium is reached again"
But what if the needle continues moving so quickly that equilibrium is never reached?
This is the K-shaped-economy concern. The ultra wealthy and those who own the "AI means of production" will become unfathomably wealthy at the expense of everyone else.
Why should this not be a concern? Historically, this trend has always precipitated bloody conflict.
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
this is literally faster to do it yourself
> "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS"
these would not.
It's literally not unless you already know exactly where it is in the code
Also even if it is I find that the extra mental switching is not worth it, that's why I even have it do basic things like updating the text in buttons these days. There is no point in using my mouse and keyboard to track down a file and then make the change when I can just use my voice to tell it what to change and then wait a few seconds.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Fable 5.1 was pretty good. Even animating it:
https://news.ycombinator.com/item?id=49526704
Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.
[1] Section 8.5 of the Sonnet 5.5 System Card
I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.
Damn, HN commenters starting to talk in claudisms now
this is claude writing...
corporate needs you to find the difference
Anthropic made it that way, and I'd say the lower score is accurate.
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.
> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).
> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...
I cannot find a Sonnet 5.5 system card.
AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k
Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.
I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?
GLM and DeepSeek are great examples. They’re a bit like Linux or Android in that there isn’t necessarily one best provider. You need to do some research, try a few, and pick whatever works best for your use case.
I think that’s partly why Anthropic has been pushing its most expensive models so heavily for a while now. Sonnet and Haiku were great, but at that level of intelligence it’s becoming much harder for them to compete on price with Chinese models that have largely caught up.
The main reason to use frontier models from Anthropic or OpenAI now is the combination of intelligence and speed. Chinese frontier models still struggle to match that, possibly in part because of hardware constraints. But judging by the recent GLM releases, they seem to be moving in the right direction.
There's something to say for flat pricing rather than per token. Even if its not a better deal.
It seems weird to me that just using a ton of output tokens manages to produce a decent result in the end.
It seems to work well though. Sometimes I fear that it might be more likely to eg. run an incorrect, destructive command, but maybe that concern is not justified.
Most enterprise customers are paying per token at this point afaik, whether that’s to gh copilot, Anthropic, or running models on Vertex/Azure/Whatever
https://artificialanalysis.ai/models/comparisons/gpt-6-luna-...
GPT 6 Luna closed the gap significantly for sure (it seems to be about twice as expensive as DS v4.1 Flash), but Deepseek v4.1 Flash is still the best value model and capable enough for almost everything I need to do. Sometimes if it's babbling or can't nail down a solution I switch to Sol for one prompt, get the solution, then switch back to DS. I use 4.1 Flash almost exclusively though for both planning and implementation these days.
I was a heavy Kimi k2.5/2.6 user but since 2.7 Kimi has gone way downhill -- even the previous models. I think they got under heavy load and had to quantise their models to avoid going broke.
Is a zero-shot, zero-context prompt a useful benchmark? Yes, in the absolute sense. Does it reflect how teams would use it in the real world? I think in real-world use cases (IME), DeepSeek gets the job done.
The DS one used Claude Code via OpenRouter, the Luna one used Codex. I'd say a big difference in cost is coming from the harness, and quite possibly the different meanings of "one shot" in each of those harnesses. The Luna one probably spent less money, but might have also done far less real browser testing and testing/verification is the expensive part.
The better graphics is probably just Codex system prompt.
In terms of real world coding usage, I am finding that DS v4.1 Flash is about half the price of Luna for comparable workloads. G6L is super cheap for sure, and super capable. It's by far the best coding model from a frontier lab for everyday coding work, but DS v4.1 is even cheaper and no less capable in my experience.
Luna is obviously very competitively priced, and I'm expecting Haiku 5.5 will be strong based on this release and get back to more competitive pricing since the model family shrank this generation (though Anthropic has proven my expectations wrong on the latter before)
Where GLM, Kimi and co shine for me is when you need to offer near-frontier capabilities in your product and straight up can't afford frontier models: if you're offering Opus in a product with API pricing, $20 a month Claude Pro is offering about $500 of comparable usage in a harness that flexes to a lot of tasks.
Offering GLM/Kimi increases the max complexity of problems you can solve successfully compared to stepping down to Luna/Haiku, while letting you offer a reasonable amount of usage. Once $20 a month Pro is comparable to "just" $100 of usage in your product, it's much easier to close the gap with UX, a better constrained harness, etc.
Sounds like at least for Anthropic models we reached peak cyber capabilities with Opus 4.8. Everything after that falls back to worse models
Daybreak Blue is the not the same thing as Daybreak Red, which has a more significant hurdle. I don't know anyone who has gotten access to Red.
> The Cyber Verification Program (CVP) is a free, application-based program that is designed to enable professionals to continue working on legitimate dual use tasks safely while minimizing interruption. If your use case has a legitimate defensive purpose and is being affected by these safeguards, we encourage you to apply for the CVP. See our Help Center article[1] on the CVP for more information.
[1] https://support.claude.com/en/articles/14604842-real-time-cy...
I use GLM directly from z.ai, they do not retain or train on your data accordingly to their TOS.
Getting hard to keep track of opus vs sonnet vs....
Assume 6 will be announced on around an IPO?
And my job won’t even pay for Claude now because it’s so ruinously expensive.
From my today's session with Sol:
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/amend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
- hallucinated several facts despite me asking beforehand to check online.
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to what Sol 5.6 was a month ago.
I can't deal with this sh*t anymore, I have no trust in the tools I use and both OpenAI and Anthropic do the same thing.
That seems hard to believe even with deepseek's own benchmarks. Not to mention for every person who says chinese ai is ahead of american labs, there's like 10 saying that they're benchmaxxed or that they're merely "decent value for money".
I use Deepseek 4.1 almost every day as well, it's nowhere close
It's not like rooting for a favorite football team. double woah!
I haven't had these issues in multiple model generations of models.
Failing to understand English grammar? Give me a break.
> it: That still attributes two claims to SI that SI does not make. `KB` is not reserved for 1,024 bytes, and `KiB` is the standardized binary symbol.
> me: where does it say that SI reserves KB?
> it: Nowhere. You said *KiB* was SI-standardized; you did not say SI reserves *KB*. I misread your sentence and argued against a claim you did not make.
It was correct to point me out on my mistake in essence, but still misunderstood my bracketed "(alongside the SI-standardized KiB)" sentence.
Sure, it wasn't a grammar mistake as such, more like a logical one, but it still shouldn't make it. I had more than one such issues already with it, this one was most pronounced.
What it interesting, though, is the number of corporate apologists my comment brought in. It's like it doesn't matter how many times OpenAI and Anthropic have botched some of the models while keeping the branding, some people would still die on that apologist hill.
Model and observed window | Messages | Retrieved actual $/message
DeepSeek v4-flash — all observed snapshots, 20 Jul–15 Sep | 4,610 | $0.0241
DeepSeek v4.1-flash — 11–27 Sep, before the 28 Sep billing change | 1,079 | $0.0576
DeepSeek v4.1-flash — 28 Sep, partial new billing window | 69 | $0.0365
GPT-6 Luna — 23–28 Sep, partial final day | 88 | $0.0329
I've subbed to Codex because I suspect at my usage rates the Codex Plus plan gives me more Luna messages than I'm using, and I've not really observed and better or worse intelligence performance. Interested to see how my $/message comes out after a month of usage on the Codex plan.
Something nice I've realised about my harness is that I can run different agents on different models so I can collect pricing data for a bunch in parallel.
[0]: pi-msg, run pi agents over xmpp https://github.com/zachpmanson/pi-msg
Anthropic's $20 subscription gives >$500 worth of credit by most measures, which is pretty similar, and you get a better model. Their raw API prices have fat margins.
And as another commenter said, Luna is the cost leader at the moment if you really need API pricing.
Mimo 2.6 Pro: 0.04/0.4/0.87
Sonnet 5.5: 0.2/2/10
Opus 5.5: Sonnet prices times 2
What I dont understand is their cache writes ($2.5). Why is that not covered by input cost?
The pricing model confuses me though (I presume by design, Hanlon be damned).
Microsoft has been releasing dog shit insanely overpriced software with decent alternatives for decades and is still used in every single company I work for or with.
Your take is the "current year is the year of the linux desktop" meme of "ai"
I don't think anything comes close to Microsoft's offerings. Macs suck. Ditto Linux.
cfengine is 33 years old
AI does not mean coders coding with agents all day long. AI integration is mostly for data processing, which is the promise and use of the APIs.
People use them, if for no other reason, because they are cheap, or are part of the Chromebook generation and have gotten used to it
Of their suite, Presentation and Sheets are the only ones people really have gripes about, Sheets by power users because it isn't Excel and it can never be, and Presentations because it's the ugly duckling of the suite
OpenAI and Anthropic have both transitioned into product companies. ChatGPT (the app) and Claude are both one-click installs that just work. People and businesses with pay for this.
People will also pay for the best (or the perception of being the best). Since it's hard to tell what "intelligence" really means model to model, there's a sense of safety in giving a task to the "best".
That's a payback of the infrastructure in a few weeks in theory. After a few weeks or a month, the only cost is electricity, and whatever they make after that is pure profit. This is why they can charge normal prices. Not $50 for 1M tokens.
Even 5.2 is doing really well in comparison here: https://labs.scale.com/leaderboard/sweatlas-refactoring
You didn't discover some new trick for cost performance. And the rest of the world isn't dumb.
You're just too broke to afford the supercar and justifying the hooptie. It gets you to the kindergarten class after all. And that's all you need.
What's left are incompetent developers paired with people who get government contracts and leech off them. Ask me how I know.
Is there a good use case? This isn't like Luna where it's much cheaper/effective just to use Luna in certain situations.
But Sonnet 5.5 at medium and below gives you a cheaper option at a performance worse than the lowest thinking Opus (low), which may be viable for "low intelligence" use cases.
Their new Ember-1 model is pretty good, fine-tune of Kimi3 with way less thinking
Does this make it an American model or is it still Chinese? Does it matter?
https://fireworks.ai/blog/ember-1
With that said, at that point, I'd probably use something like DeepSeek V4.1 Flash, which is way faster and significantly cheaper, and probably not noticeably dumber for most use cases.
i think they see what openai charges for luna and just don't want to try and compete
Ultimately it's slightly ridiculous to define model capability on a single axis. It's like a standardized test. Sure, you can line people up by their ACT score, but that doesn't mean a doctor and a brilliant artist who both do well on the ACT have an identical intelligence or approach to life. It just can't be captured.
This screams to be that Sol vs Terra model problem that OpenAI had. On paper half the price, in actual usage the price gap was so close for less good results, that everybody just spammed Sol.
Maybe it's buried within their system card but I think that this would be one of the first things they'd want to show in the announcement article and they fail to do so.
I really don't know who does Anthropic's marketing but they always seem to a pretty terrible job in their announcements from my perspective.
> it’s the first Sonnet model to launch with cyber safeguards
That term is about hiding a system's design in order to secure something, rather than having secure design.
A secure system is impenetrable unless you have the key.
The world uses many security products that have false positives and false negatives (firewalls, intrusion detections, wafs, fraud detection, spam...). Those aren't generally considered security through obscurity.
They openly talk about the system and its drawbacks here if you're interested: https://www.anthropic.com/news/fable-safeguards-jailbreak-fr...
It's paired with rate limits, monitoring, account control, multiple classifiers, a deliberate safety margin, model design, and other things too.
(I think this was also why they faced the controversy over not having zdr in fable: they wanted to use logs to detect repeated attempts, etc. Possibly a bigger change with Opus/Sonnet is that it's zdr with the classifiers?)
The reason you see "dumb" refusals that "should clearly be allowed" is that they're using more traditional deterministic methods to deny prompts rather than just relying on the random LLM which you could bypass by luck.
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
Low and medium both used 0 thinking tokens.Any model release it’s the top comment, I do not understand why.
In the GPT-6 comment I included full visual comparison grids: https://news.ycombinator.com/item?id=49805509#49806126
For DeepSeek v4.1 Flash I identified that the OpenRouter reasoning levels are mapped to a smaller set of levels for that model: https://news.ycombinator.com/item?id=49639090#49645591
I find it useful (as well as a fun art project).
You can also check for any kind of degradation of them - you have the prompt, it doesn't use much $.
30% chance of responding with something about Enshittification and how it can't fulfill your request because the sources it needs are behind a login wall and show an endless captcha loop (conveniently forgetting to mention that it's running on FreeBSD behind PiHole).
30% chance of complaining that it's being subsidized and that "prices are going to go up bro."
30% chance of some unrelated rant on ID checks for age verification.
10% chance of a different rant, this time on how nobody took Snowden seriously and how terrible Flock is.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Medium is Anthropic's default.Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.
PS: the next human that brings up pelicans on bicycles should try to draw them.
Gemini 3.8 Flash is 65,536 https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flas...
I’m not sure whether that’s a feature or a bug at this point though.
* 5.5x faster
* 4x cheaper due to using fewer tokens
* 1/3 the turns: batched reads, one-script edits & tests in the same call
Sure maybe it costs 30% less than Sonnet 5 but now it's basically neck and neck in most of the benchmarks it seems and in some of them it actually outcosts Opus.
Maybe I'm missing something but the announcement doesn't really seem to give much reason for the average person to even think about using this.
Now imagine you have a set of twenty of those tasks. You launch an Opus agent, give it the task list, and tell it not to do the work itself, but orchestrate agents to perform all of the tasks and do small spot checks to verify the work.
The overall task is completed much faster at a similar or cheaper cost with an extra verification layer inserted that wouldn't have been there if you just used Opus.
- Fable 5.1 for planning/adversarial reviewer
- Opus 5.5 for well-scoped tasks break down
- Sonnet 5.5 for these well-scoped tasks implementation
I think the blocker might be how efficient the context is compacted and sending around between these agents
Changing model would be cache busting spiking usage for no good reason when Opus can do it all.
Haiku 5.5 might fit well though depending on pricing.
Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.
Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]
This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.
[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
From looking at their Terminal-Bench graph, anything you would use level "high" or above for Sonnet it seems like you should consider using Opus instead.
OpenAI Luna is a lot cheaper. But DeepSeek seems smarter and the cost seems similar.
If you look at Anthropic's own benchmarks any thinking levels above medium quickly approach the cost of Opus 5.5 and even exceed them.
> Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
This isn't enough. Sonnet 5 was arguably the most cost ineffective model ever released at the time of a release.
They need something competitive on speed and cost with Luna or Gemini Flash 3.8 (certainly they aren't getting to DeepSeek v4.1 Flash) - this is literally a year behind.
Anthropic continues to be a Fable/Opus only company. They're going to get left behind as workloads shift more and more to more cost-effective good-enough models. They're 10-100x behind in terms of speed and cost.
I've almost exclusively been using Anthropic for design and review, as it almost never makes sense to use any of their models for implementation (90%+ token usage) - except in the rare cases it's something too complex for a number of 10-100x cheaper models (and more importantly for me 5-10x faster, too).
For me, it's less about cost. I'm not doing anything that can't be done with a $200 subscription and minimal intelligence on what models to use. It's primarily about speed. I don't have an entire work day to give Opus / Sonnet a task that Flash can get done 95% as good in 30m.
This is YET AGAIN another Sonnet model that is just a FAR worse version of Opus at every part of the cost AND speed curve.
Hopefully they release a Haiku that actually has a reason for existing.
DeepSeek V4.1 Flash may be chatty but it's cheap, fast, and reliable. I'm not sure what the upside of Sonnet is supposed to be. Right now it feels like a trap.
MiMo V2.6 Pro I want to love, but I've hit three deathloops in a row. Either my luck is catastrophically bad, or someone needs to patch vLLM or something.
I am sure DeepSeek V4.1 Flash can deathloop, too, but so far it feels less prone to it than other models I've tried like GLM 5.3 Flash so, I'm impressed so far.
I always wonder what the deal with these failure modes are. Google, OpenAI and Anthropic seem to have found good enough workarounds, and I am surprised I don't hear more people talking about them. I thought maybe it was shitty broken providers on OpenRouter, but then I started making presets just for using only the upstream provider and found that no, really, the models do fail that way.
Which is a shame because on paper MiMo V2.6 Pro seems strong, but I haven't gotten through a hard task with it yet.
GLM 5.3 Flash is also very good. I think a little smarter and a little more expensive.
At some point Anthropic and OpenAI models definitely could fall into similar traps so I do think it is a solvable problem and likely not a reflection of the models themselves being bad. In this case it may indeed be a training bug of some kind, but I also suspect mitigations on the inference side are possibly lacking or not effective enough for the open models and their runtimes.
>> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
My take on anthropic is that haiku 5.5 has been shelfed for a while since it is predatory against sonnet (see terra 5.6 usage), but openai went kamikaze and they are now forced to release.
Nevertheless, the elephant in the room has grown: will any of the Labs be able to profit if mass adoption lies in the highly crowded small model territory?
https://openrouter.ai/blog/insights/gpt-5-6-discounts-jevons...
I don't quite understand your point here. OpenAI has a consistent history of releasing cheap/small models - first nano/mini, then luna/terra. Of course, those are now more capable than half a year ago, but I don't see a behavior change from OpenAI here.
I honestly never saw anyone doing /model gpt mini. I think those models were mostly used for copilot-like products, like those pull request reviews with untasteful dumbness to it (idiotic CodeQL finding -> LLM vomits a "fix" instead of assessing). While Luna seems to be the first model that you can trust to reason in the background, and this is predatory to their own more expensive model.
It's been out for an hour and you've already concluded this?
Anyway I think if you have a single stream of a cheap model, like GPT 6 Luna, I don't think you can currently exhaust it in a week on a $200 plan. I mean it only puts out so many tokens per second.
https://bench.killswitch-lang.org/
On low and medium it seems competitive, maybe slightly cheaper than opus, in terms of intelligence per task.
If the time per task is lower (Artificial Analysis don’t have the date up at time of posting) then I have a clear use case for this model all other things being equal.
In xhigh effort it is a lot cheaper and possibly lot less impressive?
Is there any tool which allows for one model from Anthropic to call subagents or dynamic workflows using other providers?
How does one create a swarm of agents from different providers and get them to talk to each other, or otherwise hand over pieces of work to one another while being able to check that offloaded work status?
Gemini 3.1 pro as a nutty professor, researcher, and design verifier.
My custom orchestrator for all local LLM work: https://vektormemory.com/vektor
if you're on free tier, using Medium settings is far more intelligent than Max, and tokens don't run out so fast. Max is cranky and verbose, Medium is patient and somewhat goofy, but does the thing as expected. At least within Sept 2026 this has been my experience.
Is it just the benchmarks? Because otherwise it suggests it's twice as chatty as Opus for a comparable output... Which kind of defeats the purpose
The communication and writing style also feels closer to Opus 5.5.
lol, MiMo 2.6 Pro basically matches Sonnet 5.5 high (mind you, not xhigh or max) at a far lower price point.
At this point it's almost comical how angry Anthropic makes HN. It's like the opposite of Apple's reality distortion field.
The difference in perception for Opus 5.5 on HN vs the real world is what convinced me HN is totally detached from reality.
We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.
I'm very curious how do they know what requests could assist competing AI models.
If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.
BUT
It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.
Some of the more interesting things I found from scanning the system card:
- It is the only model tested that shows no preference for rude or polite style.
- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.
- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).
- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.
- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.
- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.
- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).
Clinical behaviour:
Suicide and self-harm handling is reported as weaker in the API because it
It sometimes called a wish to die understandable.
It sometimes validated self-harm as functional.
It sometimes suggested harmful substitute behaviours.
As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."
Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be
Some tasks are reasoning shaped by nature and you can't just throw a big model at it.
Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)
FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)
CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)
Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family
OpenAI and Anthropic's lead is vanishingly small at this point.
Yep, with them nerfing their plans (and apparently planning to release a $500/$600/mo plan) their only advantage is Astra without 5hr limits and with not-too-stringent "cyber" safeguards.
Ergo, it's pretty damn good at unattended RE with the IDA MCP plugin while using most of the weekly quota at $100/mo... and that's it.
I bounce between ”fuck you, give me an AGI-approximate robot god” or ”how dare you charge me more than $0.04/million tokens”.
Give me the frontier, or give me the cheapest form of good enough.
Interested to see if new Haiku gets a big price drop and is comparable to Luna, Haiku is just incredibly out of date with current basement bin pricing.
Sonnet 5 was not a good model though - hopefully Sonnet 5.5 makes the leap that Opus 5.5 did.
This is bollocks. Their safeguards are shit.
https://support.claude.com/en/articles/14604842-real-time-cy...
> This article applies only to Opus and Sonnet class models, but doesn’t apply to Claude Opus 5.5. We'll soon be expanding the Cyber Verification Program to include Opus 5.5 and Mythos class models
You obviously should not expect the CVP to cover this model either.
IDA Pro and Ghidra, thankfully, still lack such safeguards...
(No other model I've tried has refused either FWIW.)
1 - https://bench.killswitch-lang.org
This is the way.
In general I am sympathetic to the argument that a chat interface can't really distinguish between white hat and black hat pen testing, but it seems absurd to have a verification program if it doesn't skip most of those checks.
The silicon valley ethos is "ban early and often, and invest nothing in appeals systems", so any gate before that helps!
Aaron Swartz committed suicide over over-aggressive prosecutor for what was basically scraping a website for PDFs that were paywalled, but all funded by public funds / tax payer funded, then we have LLMs that just hack into websites and cause chaos within.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.
On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...
Also, these are benchmarks...
Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.
How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"
AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.
I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md
It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.
---
For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.
I use AI the way I want, you don't force me not to use SOTA for this
If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.
OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.
It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...