Rendered at 17:41:09 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ashleyn 22 hours ago [-]
If you were wondering the same thing I am - it's not about skills loss, quality, and less about money spent. It's more about frontier AI shops dogfooding their own models.
nateglims 21 hours ago [-]
If it’s anything like AWS there’s hundreds of people making bespoke software factory setups, enhanced interfaces for ai tools, spinning up 10 parallel review agents with the best model available, etc because the budget is basically unlimited.
devin 18 hours ago [-]
The other part that is kind of comical is that the software factories keep growing quality gates and automated tasks that need to run for every commit, PR, deploy, etc. Then someone goes "oh, but now it's slowing us down", so a new thing gets added which decides when it's appropriate to run the gate, and on and on endlessly until it's a gigantic soup of actions that are running which provide negligible, and more frequently negative value over the software lifecycle. People are making big expensive messes of agents and acting like it's galaxy brain stuff.
mepiethree 14 hours ago [-]
There are a couple people I know who actually clearly ship faster with AI, but most projects take just as long as ever
etempleton 5 hours ago [-]
The biggest bottleneck in most organizations is not the production of said work, but the bureaucracy surrounding it. AI isn’t going to solve that. So it really doesn’t matter if AI makes programmers 20, 30, 40 percent faster, they will still be facing the same bottleneck at the end.
devin 3 hours ago [-]
It's also frequently the same old bottleneck/problem that we already consciously decided to move away from. We used to produce big binders of specifications. Then a lot of us switched over to "agile", or a more flexible, JIT'd style where you iterate as you go, because surprise, you learn a lot as you go. This remains as true as ever, even with agent help.
nradov 14 hours ago [-]
Are those projects delivering more functionality or same amount?
pjerem 12 hours ago [-]
Most of the times, the bottleneck is mainly the company’s legacy. People are using AI on an old, rigid codebase and are still following processes and team rituals dating back from before the IA.
SlightlyLeftPad 17 hours ago [-]
Ye’ Olde Pendulum be swingin’ back again
Muromec 16 hours ago [-]
If the money is free, wht not.
gretzquee 6 hours ago [-]
[dead]
joshstrange 4 hours ago [-]
> If it’s anything like AWS there’s hundreds of people making bespoke software factory setups
Yep, this is happening in lots of companies. I know my small company has about ~1/3rd of the developers working on one (different ones, all semi-personal projects).
I sometimes wonder if we are drawn to building these because it feels like one of the remaining challenges and a desire not to be an "LLM text shuttle". Additionally/alternatively, you can go as fast as you want with building a software factory and you aren't held back by the "bottlenecks" (PR review, QA, etc).
Working on my software factory is the closest to a "flow state" I've been able to achieve since LLMs turned the corner earlier this year and replaced the vast majority of our code-writing work.
xpct 18 hours ago [-]
I find it a bit comical. AI tools are so easy to pick up, and it's unlikely you've pushed some highly critical features which couldn't have waited for a few months.
iJohnDoe 15 hours ago [-]
You make a good point. That’s why there is AI fatigue. People are moving at speeds that aren’t normal and it’s mentally taxing.
If it’s not already a thing, it will be soon, the mental health aspects of all professions moving at unsustainable speeds and what that will do to people.
pdimitar 14 hours ago [-]
No stakeholder will care, they're high on the illusion they can drain a batch of people and then replace them with fresh meat.
Nobody will take care of us. Ever.
ponector 9 hours ago [-]
But it is not illusion. Everyone knows about conditions at Amazon but they never had an issue to hire people.
pdimitar 3 hours ago [-]
Oh I agree about there (and many other places). I was talking about programmers. We started getting treated like warehouse workers, gradually.
red-iron-pine 4 hours ago [-]
they pay
locknitpicker 13 hours ago [-]
> You make a good point. That’s why there is AI fatigue. People are moving at speeds that aren’t normal and it’s mentally taxing.
I don't think any vague claim of "speed" has anything to do with fatigue.
I do think the parallel world monitoring and validating N work fronts, followed by rounds of fixes where you lay there waiting for the agents to converge, is the defining factor.
Instead of you hunkering down and hammering out a single task where your full attention can be focused 100% on a problem, your mindset is a kin to juggling N balls and hoping to not let any of which to fall.
So it's not really about speed but throughput, and keeping up with the throughout rate is mentally taxing.
RataNova 3 hours ago [-]
Reading, understanding and verifying code has always required more cognitive effort than typing it out. Running agents in parallel just forces you to spend 100% of your day doing the most taxing part of the job
inSumErgoCogito 9 hours ago [-]
I guess a shop can burn away his whole crew, no newcomers- and the old guard is burned out and quits- and the profession suddenly is just AI and managers yelling at AI realizing the job leads to burn out fast.
gretzquee 6 hours ago [-]
[dead]
18 hours ago [-]
Jach 17 hours ago [-]
Budget unlimited + upper management claiming that engineers should be spending at least their salary equivalent on token costs or they're being ineffective. So many possible misaligned incentives...
seanmcdirmid 19 hours ago [-]
Token maxing really is a thing if your token budget is unconstrained.
claysmithr 18 hours ago [-]
No, because not all work is productive
onion2k 16 hours ago [-]
Tokenmaxxing doesn't say anything about productivity. It's only about using as many tokens as possible; there's nothing about using them usefully.
red-iron-pine 4 hours ago [-]
no joke when we were trying test and baseline token use they had us do anything w/ the AI.
ask it for recipes, weather, directions, explain baroque art, etc.
arguably still a common use case...
kingstnap 42 minutes ago [-]
All these things are sort of irrelevant amounts of tokens. Even if the answers are worthless.
Like in my experience the real burn is if there is some sort of feedback edge where an agent is producing stuff that eventually it has to reconsume.
Like if there is a dag of agents doing something its almost fine.
But if there is a loop somewhere without a huge amount of damping then you have an issue.
If you have ever been a human debugging in the middle of an agent it produces like so many commands and tests for you to do. Like way more than a human would.
I'm pretty sure they do the same thing to each other. Like the loop will amplify the yapping.
seanmcdirmid 12 hours ago [-]
If you can use as many as you want but only 1% productively, thats still potentially infinite productivity. Of course, the efficiency is not that great, but more tokens = more productive work. This is only a thing if your token budget is unconstrained, which isn’t sustainable and never lasts.
gradus_ad 19 hours ago [-]
Tokenmaxxing is just the new pr-maxxing. Proxies truly are the root of all evil.
UltraSane 20 hours ago [-]
That sounds both very fun and very stressful.
wolvoleo 17 hours ago [-]
Microsoft is not a frontier AI shop at all. They just resell others' models. They don't have anything of their own worth mentioning. They missed the boat and tried to acquire OpenAI and failed, and now they pivoted their strategy to be a model-agnostic middleman.
And meta is only arguably so. Not really in the same league as OpenAI and Anthropic. More second tier like Google and xAi. (and of those Google is pretty close to the top two at times)
locknitpicker 14 hours ago [-]
> Microsoft is not a frontier AI shop at all. They just resell others' models. They don't have anything of their own worth mentioning.
Microsoft's MAI-Code-1.1-Flash is on par with OpenAI's Luna line of models. If not for OpenAI's recent radical change of heart on Luna's pricing to hastily slap a 50% discount, MAI-Code-1.1-Flash could very well be the dominant cheap model.
insane_dreamer 13 hours ago [-]
> MAI-Code-1.1-Flash
I see MSFT hasn't gotten any better at product naming :/
Talking to my friends working there, they said they were surprised when that happened and they didn't find Claude quality to be that much better than Gemini when using agy. So maybe after all the issue is not mainly the model, but harness and custom tooling.
kelvinjps10 21 hours ago [-]
They give claude usage on antigravity and aren't them a investor on antrhopic?
nwhnwh 21 hours ago [-]
What in the world is happening?
CamperBob2 21 hours ago [-]
Ask Claude
nwhnwh 20 hours ago [-]
We broke up, ask it yourself.
netsharc 17 hours ago [-]
I wonder if there's a couple (2 humans) out there where his name is Claude and her name is Alexa... Or Siri.
nwhnwh 4 hours ago [-]
I wonder what they would talk about?
actionfromafar 20 hours ago [-]
Life uhhh... finds a way.
BearOso 19 hours ago [-]
I dunno, it could be about money, too. $100,000 budget per employee, per month. Why were they willing to spend that much on AI usage? I hope said employees also make that much in salary.
iririririr 18 hours ago [-]
it's circular revenue. they can claim tokenmaxing on one side, and huge revenues on the other.
486sx33 18 hours ago [-]
[dead]
ActorNightly 21 hours ago [-]
I wonder why they allow it at all.
Like its a no brainer to force your employees to use your own models, then RL train them to be better.
p1necone 20 hours ago [-]
Only if all you care about is developing models. I assume the rest of the business would rather just use whatever's best in class regardless of who made it, so I'm sure it's not that straightforward of a decision.
pimeys 12 hours ago [-]
I would say the rest of the business world is looking for price and ROI rather than the best in class.
p1necone 12 hours ago [-]
Those things fall under the 'best in class' calculus for me.
locknitpicker 13 hours ago [-]
> Only if all you care about is developing models.
Is security no longer a concern?
When you are using a third party model, you are literally feeding it not only your current software but also all the context and work fronts under development.
This is way more than granting a third party access to your internals. This is feeding it in advance updates on all their operations in real time.
NewLogic 8 hours ago [-]
Rule 2 security went out the door months ago. Nobody seems to give a shit that all unmonitored agent runs hit the trifecta.
InsideOutSanta 21 hours ago [-]
Or let them use Claude, track everything, and use that data to train your own models.
18 hours ago [-]
user43928 21 hours ago [-]
If you assume that noisy general usage data enables good RL, particularly compared to curated RL training sets.
I am not convinced that's the case.
mawadev 21 hours ago [-]
I think if you talk to LLMs and give feedback or openly say what works and what doesn't, you are essentially solving a captcha and produce accurate training data, while you pay for the token spend. I'd be a bit nervous with this lol.
Just one unsanitized input and you leak info. Or one hidden character and code may or may not belong to you anymore. Its very odd on many levels
user43928 21 hours ago [-]
Accurate training data?
At best you produce some noisy signals that are going to have a tiny impact if even that.
And that's on a personal plan where you didn't opt out of sharing usage data.
Business plans offer zero data retention. This is a non-issue.
plasticchris 20 hours ago [-]
These are companies famous for following the rules when it comes to handling other people’s data and IP after all, totally a non-issue and they would never violate contract law
user43928 20 hours ago [-]
It would be a low quality training set that you theorize is a goldmine and worth illegally stealing from your customers.
I think this data is likely worthless compared to curated RL tasks.
19 hours ago [-]
20 hours ago [-]
Semon132 12 hours ago [-]
`Business plans offer zero data retention. This is a non-issue.`
- Not all, I checked Anthropic has a min 30 day data retention policy for Enterprise Plan.
21 hours ago [-]
tinza123 22 hours ago [-]
Microsoft?
ihuman 22 hours ago [-]
Copilot
NewJazz 22 hours ago [-]
That's not a model.
ihuman 21 hours ago [-]
True, but its not pure OpenAI GPT. If the point is dogfooding, then they'd use Copilot instead of using OpenAI's models directly
98codes 21 hours ago [-]
They do.
Zambyte 18 hours ago [-]
They had phi for a while. Interesting that they haven't really continued with anything like that.
wolvoleo 17 hours ago [-]
Phi was more of a research project. Not trained on real world data but synthetic input. That's really good to control if you go for strict limits but it also loses a lot of originality.
But it was never a SOTA model.
therein 21 hours ago [-]
If you ask Microsoft, it is a lifestyle.
glerk 21 hours ago [-]
Dogfooding their own model ls and not letting their competitors use their data to train their.
syngrog66 17 hours ago [-]
all of the above. plus less likely to be ripped off by competitor. (yes I know they have all been ripping off the general public)
trueno 21 hours ago [-]
if AI never happened there's like zero chance I would've ever used or noticed the usage of the word "dogfooding" lmao i hate this timeline
andybak 21 hours ago [-]
That's odd. I've been aware of that word for decades.
Terr_ 21 hours ago [-]
Ditto, it's been around for a decade or three, especially if we include longer phrase "eating your own dogfood" and not just the verbification.
andybak 20 hours ago [-]
I can't be sure but I vaguely recall it being associated with the big Microsoft antitrust court case? Along with the delightful phrase "knifing the baby"
21 hours ago [-]
fragmede 20 hours ago [-]
Same as "load-bearing", but I guess that depends on everyone's non/pre-swe background.
andybak 20 hours ago [-]
Or house flipping TV shows
compiler-guy 20 hours ago [-]
Wikipedia has the introduction of the term "dogfood" to the corporate world as a Microsoft internal memo....
MS was famous for that. Seems since around .NET they have forgotten about it. Delivering and promoting toolkits to devs they don't use themselves.
latentsea 16 hours ago [-]
anyone who works in software knows this term
dleslie 19 hours ago [-]
> Within Microsoft’s cloud and AI sector, monthly AI spending limits have reportedly been slashed from $100,000 per employee to approximately $10,000 in the majority of instances.
They were allowing $100,000 per month per employee?!
That's incredible. Just a few years ago that would have been zero.
cpncrunch 17 hours ago [-]
I'm curious if they have any data showing whether it benefits productivity. I get by on the free tier of ChatGPT, and do 100% of the coding myself. It has never been a bottleneck, and I like to keep my skills sharp. AI helps in analyzing problems, getting quick answers to technical queries, reviewing code, etc. But running my own business I'm incredibly careful about expenses.
pizzafeelsright 3 hours ago [-]
We have some data and hard numbers.
The question is does it scale properly? Answer: it depends.
What is the value of a project going from six months to six hours? You can look at saving the man hours but how long can that go?
What about a project that takes $100k in tokens and generates $200MM? The difficult part is determining how much was AI is the multiplier.
We are finding both to be true while quantifying requires new accounting.
cpncrunch 2 hours ago [-]
Does six hours factor in maintenance? Also, I think it depends on the project. I was able to get a a number of quantum computing programs knocked up in braket with the help of various AI tools to run some simulations It would have taken me months on my own, but I got knocked out in a few prompts with free AI tools. It's a real eye opener.
However, that's very different from building and maintaining a product that will be used for live customers, where you have to consider security implications. This is what we're talking about here for Meta and Microsoft, and it's what I do for my actual coding job.
Quothling 12 hours ago [-]
The whole world is working on this right now, but in some cases it can be extremely money saving. We had a couple of engineers who build a web portal to control some IoT devices because there wasn't one in the market. We've done similar things like that in the past, where we'd buy development from the outside. I can't give you exact costs, but getting it developed is around 5% of the price in AI and HUMAN hours, and then another 10% (bringing it to around 15%) for human hours rewriting it to run securely (and maintainable) in our cloud.
We're working on bringing the last 10% down by giving non-developers various AI skills such as a compliance module that can be pushed to specific Entra groups which will make sure their vibe code madness will stick to external dependencies we've approved. (There are still hard checks later).
For other tasks it's running on guesses, and sometimes you can save help people €1000 a week by teaching them how to use the AI better.
cpncrunch 4 hours ago [-]
Yes, that sounds like a reasonable use case. But very different to the 100k a month by meta and microsoft.
leesec 16 hours ago [-]
This is so ridiculous lol. not paying 20 bucks for a sub has to be a handicap
cpncrunch 13 hours ago [-]
Not for me. I end up using chatgpt about once a week for my actual work, and maybe once a day for non work related things. Certainly if you need the subscription it's worth $20/month or whatever for a personal sub, but I don't seem to need that at the moment. I'm more curious about the $10k or $100k/month being value for money with people using it to do their entire coding.
kanwisher 7 hours ago [-]
the $100 and $200 plans can basically give you enough usage you can never write code by hand again. If you cant save yourself 25% of your time at least isnt worth $200 then maybe you should rethink this business
cpncrunch 4 hours ago [-]
Coding is much less than 25% of my time, and as i said above I want as much coding as I can get. I do very well thank you.
pizzafeelsright 3 hours ago [-]
We have projects that would burn through a subscription in a matter of hours. The scale of some people are doing is impressive.
not_a_bot_4sho 12 hours ago [-]
It was the first limit impressed but majority of employees went nowhere near it.
Then they tightened it down again and again. Right now, it's at $2k/month or $24k/year and the model restrictions are and to go wider.
Still, most employees are not hitting the limit but enough are that there's an exception process
0cf8612b2e1e 19 hours ago [-]
Just a cool million+ annual expense per headcount.
ryanhecht 16 hours ago [-]
Notably, the limit per employee != actual average usage per employee!
blitzar 10 hours ago [-]
I presume failing to hit your target (limit) saw you on a performance improvement plan, or even worse not invited to the pizza party.
BLKNSLVR 14 hours ago [-]
Is that the sound of a deflating bubble I hear?
internet101010 18 hours ago [-]
Microsoft bloat really knows no bounds.
MisterMunchkin 22 hours ago [-]
My company took away my Claude because it’s too expensive. I feel like there is a reckoning coming. The accountants are finally realising the cost of token maxing.
vablings 22 hours ago [-]
That's pretty stupid. Most people who are incurring significant costs are just tokenmaxxing rather than being efficient with usage. You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
Tsarbomb 22 hours ago [-]
There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't.
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
weinzierl 21 hours ago [-]
I get it but it goes against the grain for me. Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us?
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
jjav 9 hours ago [-]
> Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us?
Nothing ironic about it, it's basic engineering efficiency optimization.
Just today I was calculating that if money was not a factor at all, we could simply use Mythos to handle all of our continuous security scanning needs, at a cost of about $15 million a year.
Well, my budget is far (far, far, far, far) below 15M a year, so that's not going to work. So I must find compromises to make it work within budget. The single most important engineering constraint is always budget. Everything would be easier with infinite money, but there is never infinite money.
weinzierl 8 hours ago [-]
It is ironic because choosing the right model is what the models themselves are good at. If there were no business incentives it'd be easy to have just a single frontend that routes to a suitable
model. Instead I have to develop the decision skills which is a complete waste of my time which I could better spend in the problem domain for tasks I'm better at than the AI.
remus 20 hours ago [-]
> In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
ForHackernews 21 hours ago [-]
I dunno, using tools and resources effectively is arguably the essence of good engineering.
Gigachad 20 hours ago [-]
This is much like how devs got grilled for creating expensive test VMs on AWS. Someone has to pay for all this at the end of the day.
RataNova 3 hours ago [-]
If someone burns through two grand in api tokens over a month and still can't move tickets across the board, the bottleneck definitely isn't a lack of compute.
jjav 9 hours ago [-]
> There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't.
Very much this. There is a vast range of effectiveness and combined with so many models and pricing tiers, it can be tricky for some. You really need to treat it like partly a programming language, but also partly as a management delegation exercise (do I delegate this task to the intern (cheapest model) or to the principal engineer (frontier model) based on complexity).
I do coworking sessions with most in my team to see how they are using it to understand and coach for effectiveness. Just using the most expensive model for everything isn't going to cut it anymore in the post-tokenmaxxing age.
n4r9 22 hours ago [-]
What do sorry points mean anymore.
SOLAR_FIELDS 22 hours ago [-]
Did they ever have meaning? It's always been a nebulous feels term
Terr_ 22 hours ago [-]
Attempting a serious but not-a-certified-whatever answer: "Points" do have meaning when properly used as a kind of moving-average tool for forecasting within a particular context.
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
fdsajfkldsfklds 21 hours ago [-]
For forecasting what, if not man-hours?
t-writescode 21 hours ago [-]
Effort. Which is a very nebulous term, I agree.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
Terr_ 20 hours ago [-]
To play with the math analogies, imagine a black-box function:
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
n4r9 7 hours ago [-]
Honesty isn't the issue, in my experience. The issue is unknowability. You usually can't know what clients actually want until you've given them a prototype and gotten their feedback. You often can't know the rough performance impact until you've done the first implementation. Sometimes you just have to work through a list of candidate solutions S1, S2, etc... and you can't know which one will eventually work. All of which means that forecasting beyond a couple of weeks is wasted effort, except for rote factory-line work.
bitwize 19 hours ago [-]
Sorry.
My wife, despite loving all things French, just doesn't "do metric". She wants my height in feet and inches, my weight in pounds, boom done. So while I know my mass in kilograms, she needs the conversion done before she can even begin to have a reference point.
Upper management is the same way. They have forecasts that they need to make, deadlines and budget goals that they need to hit. They only deal in the units of hours and dollars (or local currency). Every software engineer is accountable for their work in those units only. The conversion needs to be done before the management chain has a reference point.
One easy way to do this is to have each engineer estimate the time it takes to fulfill a story after it's been pointed; then, upon completion, record their actual hours spent. Their estimated vs. actuals tend to stabilize over time, so even if they misestimate a task, you can arrive at a good guess at the time it will actually take.
Terr_ 17 hours ago [-]
> > Problems arise when people try to perma-peg [points to man-hours]
> Upper management [...] only deal in the units of hours
I feel you're mixing up different operations here. You can always express unfinished work as likely to require a certain number of team-sprints, which are convertible to theoretic man-hours. The key is that the conversation rate is only valid for a moment, and technically that moment was the prior sprint.
That's very different from management thinking (or worse, declaring) that points have a permanently fixed proportion to man-hours.
> [...] and dollars
If your management deals in international currency, then perhaps that would be a useful analogy to them: Points and Man-Hours are different sides of FOREX, and they fluctuate based on different conditions.
When Engineering predicts a group of tasks is 54 points, that's like a foreign company signing a long-term contract in €100 EUR instead of USD. You can estimate that you'll receive ~$112 USD in a year, but the actual dollars will likely be different because the exchange-rate will continue changing before that happens.
n4r9 7 hours ago [-]
> estimated vs. actuals tend to stabilize over time
Have you measured this stabilisation?
senko 22 hours ago [-]
I love the typo.
btown 22 hours ago [-]
Claude, vibe code me an entire startup, the actual product doesn't matter, but it should all be based on the incredible pun "turn 'sorry' points into story points."
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
fragmede 20 hours ago [-]
Ah not to worry, you'll make it up in volume!
21 hours ago [-]
onehair 20 hours ago [-]
rookie numbers. in one of the top companies, i know someone who tokenmaxed so hard they ended up spending $50000
proxyscore 21 hours ago [-]
So you are going to blame this one that dev?
Define productivity, and while at it, quality, maintainability , modularity and so forth.
kragen 18 hours ago [-]
What are the biggest pitfalls devs on your team fall into?
PunchyHamster 21 hours ago [-]
It's funny, the thing that makes effective prompt also makes effective documentation/communication.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
ctkhn 15 hours ago [-]
The guys that made clown issues are still using AI to make clown issues, they just look like plausible specs now that claude wrote it more thoroughly. At least pre claude I could tell that they had missed something earlier on, ask a question to clarify, and get a real answer they had to type themselves. Now most of my stories are rehashed after I raise PRs or even after these same guys approve the PR and then realize they forgot a requirement.
LtWorf 19 hours ago [-]
You mean that now rather than taking 30 seconds to write a ticket, because we all know what we're talking about, we must take 10 minutes to give all the background information to the AI?
cyanydeez 22 hours ago [-]
there's a manifold to what "effective" means. The problem is once you get into the vibe flow, it's really difficult to eject yourself into the other realms of vscode or IDE or whatever it is you normal do because the vibing provides no anchor to what you're doing.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
mainmailman 21 hours ago [-]
I’m trying really hard to keep my skills up but it doesn’t feel productive when I’m using it to write code. It feels like I’m slowing down the AI to the point that it’s not as effective as just letting it go. But I don’t get all the learning that comes from that time along the way.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
cyanydeez 17 hours ago [-]
I'm primarily using local models. I turn on thinking visible and try to review at speed what it's doing.
Other than that, no. I drag myself out to poke around occasionally because at times the local models get to far into context and refuse to do simple tasks.
hexapus 17 hours ago [-]
I'm not exactly one-shotting software, but because of the pace I'm expected to keep, I place a lot more trust in the bot that I'm actually comfortable with. I need the job for the time being, so I just keep hitting the button and letting it do its thing until the tests are green.
I'm just working in DevOps though, so it's writing IaC, not application code (save the odd Lambda function or python script). Still, even when I spend an entire day conversing with Claude and watching "bot go brrrrr", I'm one of the lowest users in our company. I have no idea what the devs who regularly hit their limits are doing.
cogman10 21 hours ago [-]
That's what happens when token usage becomes a performance metric. As has been done at my company.
VCFundedGenYer 15 hours ago [-]
That’s just not how any of this works.
You’re conceptualizing. In reality, AI is expensive, power consuming, planet destroying, and overall productivity killing.
hatthew 21 hours ago [-]
I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
LeBit 19 hours ago [-]
That is always amazing to me.
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
californical 18 hours ago [-]
I’m a different person and definitely not a genius but my experience today goes even beyond theirs.
I had Opus trying to simplify a query for me which was slow - it ran for maybe 30 minutes, including writing and running tests, and came up with a refactor across 9 files with a couple hundred lines changed. I was looking through the output before moving onto the next step, and noticed something a little fishy- I said “why does it do x, isn’t that a more complex y?”
Opus thought for another 20-30 seconds then output “Actually that would make the majority of the diff irrelevant, if we do that change it is just these 4 lines in this single file instead.
So then I had it do that. 5-10 minutes of writing and testing and that was done.
So my company spent $25 in tokens and I spent probably an hour in total for a 4 line change that, in the days before Claude, I probably could have found the correct file and thought through the problem, understood the solution, and written the 4 lines of code myself. Probably in the same amount of time.
So basically there was no benefit at all for my time, an extra cost to the company of $25, and now I understand our codebase a little bit less instead of more if I had done all the work.
As good as Claude is at building greenfield projects it still struggles a lot at complex ones
NichoPaolucci 15 hours ago [-]
This is a recurring issue for us. Very small team. We move fast and pretty loose.
Dev + AI spend 3-4 hours on a project plan, there's a "wait a minute" moment, and finally they spend another hour dialing it back to a solution that could have been built, tested, and deployed in 2 hours.
Example: Someone was setting up a dev environment with multiple DB migrations from different branches - AI planned this wild 8 phase solution with a pretty fancy cutover event.
In review I essentially said... "Wait, isn't this a dev environment? It doesn't need 0 downtime, why not just destroy and recreate the DB" and it turned into a <1000LOC script.
Technically the original plan would have worked, it would have been more robust, but it would have taken a good deal more time to implement.
Some of this falls on the devs to know what fits our team well, what's realistic, what's obviously overengineered, etc... But some of it feels like AI just defaults to the most complex version of a thing. I catch it SUPER frequently. (And unfortunately some devs think that more complexity means it's a better solution)
californical 12 hours ago [-]
Thanks for posting this!
I feel like I’m going crazy, using all of the best models, spending time to have excellent prompts, configuring tools and skills… and still getting overly complex solutions with mediocre results.
Like it’s still impressive how far we’ve come, and undoubtably cool technology. It’s made a bunch of personal projects possible that I never would’ve started.
But for a business I’m struggling to see the ROI. Sometimes there’s a big benefit and sometimes it’s net negative. Not saying we won’t get there but I’m trying to stay grounded in the reality of today rather than the hopes of where the technology could get to
NichoPaolucci 2 hours ago [-]
It's brilliant tooling, and made the mundane/tedious parts of my work a lot easier.
But, I fear that the "facade of complexity" makes the output SEEM better. Someone might read a 10 page Codex generated plan with fancy diagrams / charts and have a feeling that because it is so complex, it must be good!
Same deal with text output in general. Oh, it uses a lot of big words and there's a LOT here, it must have done a lot of work to get to that point.
I find, in reality, that it takes much more effort to get to the simplest solution.
As they say, any old fella can build a bridge that stands, but it takes an engineer to build a bridge that barely stands...
RataNova 3 hours ago [-]
Classic tradeoff. Generating syntax is cheap, maintenance is expensive. RN models heavily optimize for the former at the expense of the latter :)
strange_quark 16 hours ago [-]
Matches my experience to a tee.
And don’t forget the company also spent a bunch of money in tokens for the initial author to implement the thing poorly.
californical 12 hours ago [-]
Thanks for the reply, I appreciate the validation :)
hatthew 18 hours ago [-]
It's not so much that I can code fast, it's that it takes a significant amount of time to tell the model what exactly I want the script to accomplish, and at that point I might as well write the script myself. And often, figuring out what I want the script to do is that hardest part, so it doesn't really matter whether writing the script takes 10% or 20% of the total time.
jjav 9 hours ago [-]
> I feel like anything that I'd trust haiku to handle isn't worth doing in the first place.
The trick is making the workload manageable by the cheapest models, or costs will destroy you. We seek to make all repetitive tasks be effectively done by cursor composer model which is the cheapest.
Doing recurring tasks with anything more expensive than that will burn the budget in no time.
onehair 20 hours ago [-]
in my company there are a few who keep sharing screenshots of reaching limits on 3 separate subscriptions, 2 of them their personal on top of the company subscription
esseph 19 hours ago [-]
Wonder when subscription-hopping attacks will become more often (jumping from a personal model to injecting instructions into the business account and exfiling data)
nozzlegear 18 hours ago [-]
> You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
You can get the same jobs and work done with Kimi and GLM (ZDR on OpenRouter) for a fraction of the price too.
poisonborz 10 hours ago [-]
No large companies I know allows chinese models, and OpenRouter is abysmal for enterprise usage. Don't confuse personal/yolo startup usage with corporate use.
funnym0nk3y 22 hours ago [-]
Sorry, but that is nonsense. Compared to opus haiku doesn't cut it most of the time.
usef- 20 hours ago [-]
I think they mean the new Haiku, which is mildly above Luna now . If you have a plan written by a smarter model (so the hard parts are solved) they can be great at implementation.
latentsea 16 hours ago [-]
I use Qwen3.8-27B as a daily driver, and for things I know will be quite hard I tend to get ChatGPT to do the planning. Works very well.
usef- 16 hours ago [-]
It's impressive for its total size if you need local inference.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
latentsea 11 hours ago [-]
> Or you want independence from the labs (fair enough)
This.
perching_aix 22 hours ago [-]
What on earth do you even do with these models?
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
dpkirchner 21 hours ago [-]
I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code.
steve_adams_86 20 hours ago [-]
I equate most LLM work to squeezing a glue bottle
It's just glue code
It's not complicated. Someone just has to be there to squeeze the bottle
perching_aix 21 hours ago [-]
It's possible it's my role distorting my perception, cause technically I don't write software, I work an SRE role. None of my items come pre-chewed or paced, it's all good luck and god bless.
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
qlte 19 hours ago [-]
For the last few weeks I've been using Luna with Medium Reasoning for routine debugging, installing/updating dependencies, test creation/fixes, CLI/config miscellaneous problem solving, etc and it's been solid. I can let it churn away for a half hour and it barely moves the needle on remaining usage on my Plus plan. Previously I had been using Sol Low/Medium and would frequently hit the 5 hour limit doing those kind of tedious automation and routine problem solving tasks.
AIblemblio 21 hours ago [-]
Our ai basic analysis for SRE / k8s based platform is haiku and its surprisngly good.
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
perching_aix 21 hours ago [-]
I wish we had such a platform (and was properly adopted). Maybe then the necessary context would be properly organized, and these lesser models could be effective here as well, especially if combined with harnessing integrations too. I still have a hard time accepting that Haiku/Luna tier models can be effective there even then, but I'll just have to take your word for it I suppose.
usaar333 22 hours ago [-]
> You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment.
Optimally? Opus will pay for itself if you save just 10% of your time
qznc 20 hours ago [-]
Only if all money is equal. Budgets in big enterprises work differently.
geodel 22 hours ago [-]
True. I always Opus to pay for itself if it wants to get used by me.
AIblemblio 21 hours ago [-]
the poster did mention "if it saves 10% of your time".
So be less snarky?
compiler-guy 20 hours ago [-]
That's funny. In a meeting with my manager recently, they specifically called me out for not spending enough money on tokens. It's not like I didn't use it, just apparently not enough, and apparently on too low a setting.
Since then I've had fable cranked up to 11 for even the most trivial of tasks.
aix1 13 hours ago [-]
That's insane. What type of company are you at? (Or even the actual company if you're OK sharing.)
compiler-guy 3 hours ago [-]
One that is all in on AI, and wants to be sure its engineers are too.
233mhz 20 hours ago [-]
I know plenty of people working at very well known and very large companies who's CEOs were boasting about "not hiring anyone anymore", "all the code will be generated in 3 months", etc. they all went from "unlimited budget per dev" + public dashboard with ranking to flex how much credits everyone was burning to hard caps at $500-$1000/month/employee real fast. Some are even not allowing their devs to use the more expensive models
password54321 22 hours ago [-]
You realise the subject here is Meta, which is all in on this stuff? Of course they are going to use Muse Spark over Claude.
>Great Depression style collapse and all the current AI companies go bankrupt.
Oh this is just a 33 day old doomer account.
righthand 21 hours ago [-]
And yours is a 4 year old hype account?
password54321 21 hours ago [-]
Nah, I just respond to a few things here and there.
morgoo 7 hours ago [-]
My company gives all employees $2500/month. I tokenmaxxx to make the most of it but have a gut feeling I could get as much work done with 10% of the budget spent on DeepSeek V4.1 Flash and similar models.
tty456 21 hours ago [-]
Do you know the details of the Claude Code plan you and your company are using (if not part of some enterprise deal)? Does your individual capacity out run something like Claude Max 20x ($200/mo)?
woah 21 hours ago [-]
$200 a month is too expensive yet they employ human developers?
cpncrunch 21 hours ago [-]
It's unclear how much OP's company was spending. The article gives a figure of $100k/month per employee.
But even $200/month is worth shaving if it doesn't generate value.
chrisweekly 21 hours ago [-]
$100k/month per employee?
um.
pimeys 12 hours ago [-]
It is called paying per token.
14 hours ago [-]
platinumrad 21 hours ago [-]
Corporations pay API rates.
proxyscore 21 hours ago [-]
So what, if they're cutting it, it has propagated to the last bean counter that the roi isn't there.
This takes some doing and now is the time where it's dawning on the finance departments.
233mhz 20 hours ago [-]
> So what,
So it's not $200 a month but it can easily reach $200 a day, and unless you're a startup playing with monopoly money the maths don't work
Chris_Newton 15 hours ago [-]
As a point of reference, I tried an experiment with Claude Code and the latest Opus the other day. It was work for my company, so this was using API tokens and not the individual user plans that have non-commercial terms.
A simple task, migrating a typical password reset flow as part of updating a long-lived web application from legacy libraries and software architecture to modern equivalents, apparently cost roughly the same as 2 months of Pro subscription, over the equivalent of about half a working day in wall time.
It produced code of decent quality at a small scale, but it wasn’t always on point architecturally. It also had a tendency to drift off topic and try to tangle up other changes it decided should be made with the main change we were supposed to be working towards. So even for a routine task, based on a plan developed using the harness first and with the agents working under close supervision, a near-SOTA model is still producing quality on par with a decent mid-level developer but substandard for anyone senior+ in this case.
Moreover, based on a direct comparison with other migration tasks of similar complexity that I’d already done by hand, it was actually a bit slower overall to work this way. I had to babysit Claude throughout and review everything it proposed carefully, both to avoid subtle errors (it would have made several) and to prevent drifting off track. I also had to spend a significant amount of time cleaning up its final output to an acceptable standard after the session. Those two overheads more than cancelled out the much faster code generation an LLM offers under favourable conditions.
So for now, I remain sceptical about these high multiples of improved productivity that I keep seeing claimed online from people who are apparently writing almost everything using AIs now. I could certainly have achieved a multiple of my normal productivity by YOLOing everything without reviewing it in detail and then accepting the output code without tidying anything up. However, I doubt this codebase would still have been good enough for normal human developers to work on it reasonably after even 10 or 20 AI-led sessions like that. The architecture would have degraded significantly and the test suite would have been large and largely pointless. And again, this wasn’t rocket science in this experiment, it was completely unremarkable maintenance of a relatively small and simple web application.
233mhz 10 hours ago [-]
Definitely, according to my omp stats I spent the equivalent of $650 opus 5.5 tokens at api price, this model isn't even a month old, I'm exclusively using the $20 plan and I'm not a power user at all: no backround loop, no automated anything, just using it for development and I stay in the loop the whole time
pimeys 12 hours ago [-]
Yes. Hundreds of dollars per day is easy to burn with Opus if you pay the API rates.
And now if you get the task done faster with the same quality. And the cost is 80 cents, considering DeepSeek and your own server starts to starts to be relevant...
IshKebab 21 hours ago [-]
I use Opus 5.5 heavily but only spend around $800/week at API rates. I mean, I say "only"... That's a lot in absolute terms, but trivial compared to my salary and EASY worth it.
CamperBob2 21 hours ago [-]
And someday I hope to understand why they do that. CEO: "Let's see, I can pay $200/month for Bob's tokens, or I can pay $2000/month or so, and then hope he doesn't screw up and rack up a seven-figure bill. The service is the same either way. Hmm."
endless1234 20 hours ago [-]
It's not like they can't decide to pay API or subscription rates. Subscriptions aren't possible for >150 person companies.
CamperBob2 19 hours ago [-]
Is there something in the ToS that says you can't give each of your employees a personal $200/mo subscription... and if they happen to use it at work, well, that's their business?
Even if so, it might be worth spinning up a separate LLC for each division, given what they are charging for the API.
endless1234 9 hours ago [-]
>Is there something in the ToS that says you can't give each of your employees a personal $200/mo subscription
Surprisingly, yes.
pimeys 12 hours ago [-]
You can also ask your employees to bring their own pirated Adobe Photoshop and avoid paying for the license..
CamperBob2 1 minutes ago [-]
Nobody's talking about pirating anything.
20 hours ago [-]
AIblemblio 21 hours ago [-]
Besides that the article states quite high numbers, budget is budget in these companies.
You had budget for your normal salaries, for externals and now suddenly you have a few millions additional.
What do you do? You compensate.
Business people doing business things.
inferniac 20 hours ago [-]
taking away sounds insane, we had basically unlimited tokens (inference bought from aws) and they moved us back to the $100 sub to save money
lbreakjai 18 hours ago [-]
I've got a 100$ openAI sub through work and I never even get close to reaching the limit. I really wonder what the hell is everyone else doing that they could even get close to spending 4 figures a month in tokens.
locknitpicker 13 hours ago [-]
> My company took away my Claude because it’s too expensive.
My company didn't took away any expensive model, but it did imposed a tight budget and gave my team hardware to run local models. The budget works mainly as an influence on the decision process of which model to run. We can still run expensive budget-burning models if we really want to, but there is a clear incentive to use options that allow us to save our token budget for a rainy prompt.
Surprisingly, I feel this had a positive impact on results. Instead of succumbing to extremely expensive one-shot prompts with the most expensive model in the news, chaining agents running cheap models equiped with context and specialized skills and scripts ends up having a better and more reliable output. And faster too.
I think there is a lot of propaganda, perhaps even astroturfing, on how only the most expensive models churned out by US companies are worth using. Nowadays the cheapest models get the job done, and local models can already handle most tasks as well specially as part of orchestration chains.
epolanski 19 hours ago [-]
My company (me, I'm self employed) did so as well.
Since this summer coding on Opencode Go + Codex for a total 28$/month gives me more intelligence and token than 400$ did in may.
Also, SOTA models are increasingly useless for anything even barely tangential to security work.
lenerdenator 21 hours ago [-]
We're just getting put on a budget.
Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.
Gigachad 20 hours ago [-]
Does all this velocity translate to increased income to the business though. At some point if we are releasing 20x more features, more games, more music, who is actually buying it all?
lenerdenator 19 hours ago [-]
At this point we're solidifying a lot of stuff for the product: disaster recovery, getting workflows to scale so we can actually sell it to more people, paying off the Fort Knox of tech debt, so, at least in my case, yeah. YMMV.
NichoPaolucci 15 hours ago [-]
I'm in a similar boat. When I joined my current company the tech was this 700KLOC legacy behemoth stuck in 2005. It was pretty much all tech debt.
With agents I've been able to make a decent sized dent in it. Lots of this gain is due to AI - but the fact of the matter is we still have 50+ bigger projects we could work on, and non-developers building with AI has increased that number.
Busier than ever, because of AI.
Gigachad 11 hours ago [-]
That’s not the original question though. Are there more subscribers to the product now?
Certainly more change is getting done, but is the market paying more for it?
NichoPaolucci 6 hours ago [-]
No, our sales are up slightly YoY. The system is more stable than it has ever been. Customers are more happy.
Some of our AI tooling IS “generating revenue”, but not most of it.
Much of it is workflow efficiencies and general improvements.
One of our big initiatives is cutting out the CRM we use. It’s going to save about 30K a year to do that in house.
Our AI bill is going to be sub 30K USD this year, and if you asked the CEO if there was a return on that investment I believe they would be positive.
But, I am 100% sure there are multiple companies wasting money on AI and dialing it back because they are NOT getting a return.
233mhz 20 hours ago [-]
> Our velocity
Is this the new buzzword for the quarter? Last quarter was "granularity", I didn't get the memo yet
lenerdenator 19 hours ago [-]
Is that not a widely-known word in project management circles for scrum and other methodologies? I've heard it for years. Basically how much work you get done through the lens of how long it took to get done via a pointing system.
233mhz 19 hours ago [-]
I didn't know people were using these non ironically.
Might as well go back to counting the number of lines or number of commits. It doesn't sound as good as "granularity" and "velocity" though. The good thing with velocity is that it doesn't care which way you're going as long as you're going there fast, so you can never be wrong
jpleyden98 16 hours ago [-]
> The good thing with velocity is that it doesn't care which way you're going as long as you're going there fast, so you can never be wrong
Velocity famously being a vector with a direction. Go in the wrong direction you'd have negative velocity.
Your description would be better for the concept of speed (IE the magnitude of velocity or directionless velocity).
18 hours ago [-]
yeahBoiii 21 hours ago [-]
The reckoning started years ago when we did the equivalent to token maxxing hiring coders for everything to crank LOC
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
AIblemblio 21 hours ago [-]
The reckoning will be throwing out all external help, then reducing team sizes.
righthand 21 hours ago [-]
My thoughts were the reckoning would come when Infra teams started offloading AWS usage to LLMs and ended up token maxing and deploy maxing.
sergiotapia 22 hours ago [-]
which is quite sad because opus 5.5 is really good. i say this as an anthropic hater. i wish I could move away to other models like 6.1 sol or deepseek or whatever, but they just all lack something. i _trust_ opus 5.5
i hope other labs catch up, especially chinese labs.
fn-mote 16 hours ago [-]
> i _trust_ opus 5.5
Just passed the Turing test.
dude250711 22 hours ago [-]
Not sure it's even possible to catch up by distilling.
dyauspitr 21 hours ago [-]
I mean, we’re not far from a situation where instead of how many story points you completed per sprint the metric to optimize is going to be what was your efficiency? How many story points did you complete while minimizing your token usage. In fact, that’s a pretty good idea. I’m going try and implement it at work with some sort of complexity normalization function
user43928 21 hours ago [-]
Good luck to the accountant that tries to tell leadership to slash AI usage.
I'm sure investors will love it.
Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.
You think now is the time they're going to cut the spend?
vld_chk 22 hours ago [-]
If true, it is a huge blow to Anthropic’s revenue stream. IIRC it was reported that the quarter of their revenue comes from just two clients and as the ex-Meta guy who left this July, I am convinced that Meta must be one of the two.
simonw 22 hours ago [-]
That was 2025. In 2025 a quarter of their revenue came from two customers, and those customers were GitHub Copilot and Cursor.
In 2026 their revenue has gone up by a factor of more than 10x, and they no longer have just two whale customers.
I heard a rumor recently that customers spending less than $100m/year aren't even considered their "top tier" now.
shakna 21 hours ago [-]
We know the 2025 revenue numbers from Anthropic's leaked prospectus.
Where are you getting the 2026 figures from? The Bloomberg article that was "on track to generate" and just prediction?
> The company's financial trajectory already shows how quickly that equation is changing. Anthropic's revenue run rate was about $9 billion at the end of 2025, according to the company, before rising to more than $47 billion by May. Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, on track for its first quarterly operating profit of $559 million.
> The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter.
These are leaked and self-reported numbers, but no matter how much skepticism you pile on them it still looks likely that Anthropic in 2026 have had some of the fastest revenue growth of any company in history.
weakfish 18 hours ago [-]
I don't have access to the article linked, but I am default skeptical of any self-reported numbers. They have a financial incentive to play it up.
If you're able to link something non-paywalled, I'd be happy to read thru.
They may be self-reported numbers, but they're consistent across multiple reporting sources. If Anthropic are lying to their investors about these numbers they will be in very real trouble with the SEC come IPO time.
And even if they're using funky accounting tricks to exaggerate their profits and downplay some of their losses, I expect the numbers for 2026 will still be really impressive.
weakfish 18 hours ago [-]
I'll check w/ archive.ph in a bit, thanks for the tip.
Is the hope for them that they'll outpace their run with a massive revenue growth? I believe that they have that, but numbers like a net loss of $46b in 2025[0] seem hard to overcome.
Re: self-reporting, I'd be inclined to think they're using non-GAAP practices to make it sound more favorable, but I am by no means an expert, so willing to believe I am wrong.
I'll read the article you sent; this is just my off-the-cuff thinking.
simonw 18 hours ago [-]
They claim to have been profitable in both Q2 and Q3 (under whatever their definition of profitability is).
I really don't think the 2025 figures are interesting at this point. Everything changed for them in 2026 - nobody was spending $1000/month/employee in 2025, there wasn't enough interesting token-heavy stuff to do with the models.
I think Anthropic might actually make it to profitability - they're earning more revenue than OpenAI and they've spent significantly less, too.
weakfish 5 hours ago [-]
I think that's a fair prediction, and you're right that there was a curve in 2026.
I do think their definition of profitability is likely skewed, but you're right that 2025 is out-of-date.
They're limiting it to $10,000 per employee per month.
loeg 16 hours ago [-]
That's Microsoft, not Meta.
binlog 22 hours ago [-]
Meta was #1 by far
22 hours ago [-]
watwut 22 hours ago [-]
Well that explains the limitation.
gfrecvh 17 hours ago [-]
What do you mean, "if true"?
Doesnt matter if Microsoft and meta are pushing employees towards their own models today. What matters is that it's the economically sensible direction, and so it will happen if it's not today.
gloryjulio 20 hours ago [-]
There is a conspiracy theory that the tokenmaxxing period was anthropic/openai ipo play. They do the tokenmaxxing to revenuemaxxing first, then time the ipo window to show they had the the huge growth to justify their price tag.
However Elon was able to ipo before them which took a lot liquidity of the market. Their financials are exposed. The market condition and sentiment now is in the gutter. It would be very interesting to see how these would pan out
lbreakjai 17 hours ago [-]
That's my theory, which I've seen being echoed here a few times. Encourage increased usage, suddenly raise prices, capture the photo-finish in the short lapse between prices going up and usages adjusting, and present that goldilock point-in-time to the investors as the new normal.
throwitaway222 21 hours ago [-]
> monthly AI spending limits have reportedly been slashed from $100,000 per employee...
what. I can see a team of 20 costing 100k per month (but rare), but per person?
nkrisc 21 hours ago [-]
Considering that’s a healthy portion of a salary for an additional employee per person, the fact they’re slashing spending sure makes it look like AI wasn’t even a 2x multiplier at minimum.
legulere 21 hours ago [-]
You could only conclude that if they would have dropped LLMs completely. It just means they see the benefit/cost optimum at a lower point than 100k/programmer.
I also doubt a longterm 2x multiplier for most developers.
pinkmuffinere 20 hours ago [-]
I don't think you need the LLMs to be completely dropped to indicate this. Even entry-level faang programmer salaries are in the 300k range, so the fact that LLMs aren't worth 100k/programmer does tell us something about the marginal usefulness. If the drop was motivated purely by price-per-output, it would indicate the LLMs are less useful than adding another 1/3 engineer. However, as discussed in other comments, there are other motivations as well, notably dogfooding homegrown tooling.
legulere 6 hours ago [-]
Let’s say you could get 10x speed up with 10k and 11x speed up with 100k. The 100k option still wouldn’t make sense.
It’s a contrived example, but to me it’s not improbable that a lot of the spend comes from tokens that are used very inefficiently. After all developers were pushed to try using LLMs without looking at costs.
I’m not saying the others are wrong, just that you can only tell something about marginal cost effectiveness, not total cost effectiveness of LLMs by measures to reduce AI spending
loeg 16 hours ago [-]
The marginal utility over $10k/mo/employee isn't super high now, but maybe there is some future where it is. shrug
ivantop 10 hours ago [-]
Not even close to 300k
pinkmuffinere 10 hours ago [-]
Sorry you’re right, I said entry level but in my mind I was thinking after one promotion, not sure how I messed that up. L5 software dev total comp was about 300k when I left Amazon two years ago. Levels.fyi reports it’s lower now [0], but I am suspicious of that. Maybe it’s brought down by non-HCOL locations.
Even if it was a net negative, you still might want to titrate the exit a bit just to prevent the total chaos of broken workflows. Can't really draw the conclusions you want to draw.
flatline 21 hours ago [-]
We're sitting on a year of more or less capable coding models and I have not exactly seen a revolution in new software being released. AI is likely a force multiplier for specific subsets of individuals and workflows. Coding is not a bottleneck for a lot of systems.
I suspect that SaaS is silently being eaten alive as an industry.
legulere 21 hours ago [-]
You often pay for regulatory compliance and liability with SaaS. Would you risk running your own vibe-coded payroll or accounting system for your company?
weakfish 18 hours ago [-]
Yep. The SaaS-is-doomed crowd are missing this part. It's the Red Hat model (debatable how well they do it, but they exist, so..)
IsTom 20 hours ago [-]
I think SaaS will be fine. How the field works might change, but in the end marginal cost of software is still zero and when you've got a quality, well-tested and well-designed product it's going to be obviously better than things vibed together.
holoduke 20 hours ago [-]
Talk to the emulator community. Huge leaps being made there. Or binary decompiling topics. Huge leaps there. There is more than web applications
kragen 18 hours ago [-]
Can you link to some of the huge leaps in emulators and binary decompilation that have caught your attention? (I hope that doesn't sound like a challenge. I'm just interested.)
That's the thing though. My hobby projects have been doing better than ever. My job is more or less the same it's always been, just with more unit tests (that cover...who knows) and slightly better tooling/CI because it's now less tedious to fix that shit when it gets mangled.
AI is very good at creating small codebases. It still sucks shit at maintaining huge ones where someone needs to understand how the codebase works.
I've seen what happens when companies just vibe code an entire mega codebase and it's really bad. AI is fucking terrible at grand architecture decisions.
IshKebab 21 hours ago [-]
That's monthly, not yearly.
nkrisc 19 hours ago [-]
Oops, yikes.
grebc 21 hours ago [-]
Very likely a negative multiplier.
oldmanhorton 21 hours ago [-]
It’s mostly from people using their personal accounts to run LLM services that serve a larger team or organization. At least at Microsoft, it’s still impressively hard to get access to an LLM for service usage with high enough rate limits to be useful, making running services on dev boxes much more appealing (despite the countless drawbacks that few people seem to care about around security, compliance, reliability, etc).
Benard-dev 22 hours ago [-]
This is not surprising. Every big lab blocks competitor tools for internal use; it is a data governance thing, not a quality statement.
jvanderbot 22 hours ago [-]
Right! it seems obvious why: Both these companies want to dogfood their own coding models and stop paying competition.
You can also read this as diminishing returns / AI isn't good enough, etc, but the simplest explanation is that they don't want to send money to Anthropic.
fasterik 22 hours ago [-]
This, and in addition to dogfooding, incentivizing employees to be more effective with the cheaper models. A lot of problems don't need anything fancy, but it takes more brain power and engineering effort to make that work. By default humans will take the path of least resistance if it's available.
mattm 22 hours ago [-]
This is likely the future as well. Down the road, every company will have their own internal coding models.
NewJazz 22 hours ago [-]
Does MS have coding models?
dymk 18 hours ago [-]
yes, but there's a reason nobody talks about them
WaltPurvis 22 hours ago [-]
But Microsoft and Meta are not blocking competitor tools for internal use, they're merely trying to reduce costs and divert a fraction of use to their own technologies. Microsoft and Meta are both still spending nine figures a year on Claude, and the article does not state or imply they're even considering a complete halt.
AIblemblio 21 hours ago [-]
Microsoft has full and unlimited access to OpenAI models. That was some agreement when they invested originally.
octoberfranklin 22 hours ago [-]
So nobody except AI labs is allowed to do "data governance"?
cmiles8 14 hours ago [-]
Most large companies are doing similar things, including:
1) Giving tight budgets for AI use after a few years of a total free for all, and
2) creating internal API marketplaces where one can access all the models including open weight and “Chinese” models.
This is increasingly sending tokens to the lowest bidder, which is a terrible setup for the big labs and their hopes of being worth trillions.
The big labs are tying to spin up “enterprise sales teams” like the traditional big SaaS players and trying to secure large contracts, but it’s broadly blowing up in their face and turning into these realtime token marketplace models.
ehnto 14 hours ago [-]
Will the productivity expectations drop with the AI budgets I wonder?
cmiles8 13 hours ago [-]
Most enterprises say that while the tools are useful and here to stay in some more limited form, the productivity improvements have been broadly underwhelming. So it’s more the inverse, because the impacts observed have been limited companies are reining in spending. When CEOs were hoping AI was going to result in massive productivity improvements CEOs were handing out blank cheques.
Beached 11 hours ago [-]
We see the amount of time employees write code decrease. But that time has been taken up by double checking the code, having that code reviewed by 1 or 2 code reviewing platforms and vuln management systems, having claude rework the code based on those results, dev testing and deployment testing, product and feature design work, etc.
A lot of developers spend significantly little time coding. So yeah, their coding time has decreased, but they were never coding for 6 hours a day every day to begin with.
tyre 12 hours ago [-]
Models are getting better all the time and either keeping prices or lowering them.
What you could achieve with $10k a year ago vs today is quite different.
rangledangle 22 hours ago [-]
We've reached the era of "good enough" ai, it seems. The truth is you don't need the best model in most cases.
combyn8tor 21 hours ago [-]
I feel like "good enough" was reached around Opus 4.6 - 4.8. All I wanted after that is improved speed, continued tweaks to the tooling to get the most out of it and quality of life features added.
efields 21 hours ago [-]
That's where I am too. Opus 5.5 but increasingly faster is a future I'm hoping for.
IshKebab 21 hours ago [-]
It depends on your domain. In mine we didn't reach good enough until Opus 5.5/Astra.
Beached 12 hours ago [-]
As a Claude fan, and user, I must say I very much dislike the enterprise options available. I really do want something similar to the Pro, max 5x and max 20x plans but for enterprises.
Our accountants want that, our devs want that, but we cant have that. we would even be willing to pay a premium of 2x or 3x the price to have the enterprise benefits of private data that isnt retained. We would even be OK with all of the bottom of the barrel capacity and availability QoS too. We also would be ok Signing 3 year and 5 year lock in agreements.
We would be ok having a delay to the latest model by 30 days if that helps.
If anyone in Anthropic is reading. PLEASE give enterprises something like this. We want the Pro, Max 5x, max 20x plans. Charge us 3x the price for those exact plans but with the enterprise data privacy.
sublimefire 19 hours ago [-]
Just to sprinkle more details here about MSFT. The usage was trimmed down even more for some orgs and some specific accounts (I think those were vendors and the like). But if you are an FTE and somewhere in CoreAndAI then your limit is like 1M GitHub credits ($10k as article mentions) that could be spent on any model you have access to, e.g. Opus 5x or GPT-x. It is hard to pin down how many people use Claude vs how many just use internal harness over the Opus though. Importantly, Anthropic models are more expensive (not even talking about Fable here as it is a special case) as seen in usage metrics. Peeps can check their usage which has a good breakdown that shows how many tokens were burned on which model and how much that "costs" (the prices as you can imagine are not consumer ones), so IMO this change will drive more people to optimize based on their usage patterns if they are heavy hitters. The easiest path is to just switch to something like Sol/Astra and have similar perf for lower price.
greggoB 19 hours ago [-]
> Peeps can check their usage which has a good breakdown
Must be nice. At Sony we use Copilot (as an MS enterprise customer) and restrictions to token usage circa June have resulted in a hunger games of who-can-use-them-all-first, with the pool usually drying up by the 10th day of each month.
As for usage breakdown: we got a manually-published table most days just showing per-user token usage, until eventually we were told even that was somehow too difficult to provide.
Between that and learning how much the whole of Sony Corporate Group spends on Copilot (spoiler: it's hilariously little), I've sort of run out of disbelief.
NichoPaolucci 15 hours ago [-]
I'm so glad I work at a company that isn't looking at our token usage. I mean, we're encouraged to use it but hearing about leaderboards and things is wild - but we also aren't big enough to need API costs, so power users get a flat rate 200/Month subscription or whatever.
I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
greggoB 5 hours ago [-]
To clarify: we don't have leaderboards (thankfully), our issue is more in the opposite direction, with people who do want to make use of the tools (optimally or not) majorly constricted from doing so.
> I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
It is indeed surreal. I am not bullish about this company's technological capabilities going forward tbh.
acedTrex 19 hours ago [-]
Ten days of pure slop then bliss for the rest of month sounds like a hilarious amount of whiplash to get monthly.
greggoB 5 hours ago [-]
The fear-and-loathing generated by those colleagues who no longer can or want to synthesize anything themselves after the token-fueled high does indeed make for good popcorn viewing.
_ink_ 19 hours ago [-]
> Importantly, Anthropic models are more expensive (not even talking about Fable here as it is a special case) as seen in usage metrics.
Not sure where you see that. The API prices of Opus are cheaper than Astra and there was a study suggesting that the Anthropic abos are offering more bang for the buck:
Its internal, you can see the breakdown of cache/in/out and prices, but it does not mean this is the actual price MS pays.
pluc 21 hours ago [-]
You have to use Windows, Teams and Copilot; how many people did I just lose?
matsemann 21 hours ago [-]
Wouldn't mind if people at Microsoft actually dogfed themselves Windows. It's as if the designers and PMs all use a Mac and screw up everything.
Windows 11 shipped with a broken task bar, couldn't even ungroup items. No power users was ever involved in this.
Beached 11 hours ago [-]
I cant find it, but I think I recall a windows dev tweeting his frustration when they launched windows 11. He was upset because they didnt want to make the design choices they made, and the UI designers forced them to do it because it would modernize the OS and be the next great thing...
Does anyone else remember this?
leptons 20 hours ago [-]
The first things I install on Windows 11 are OpenShell and ExplorerPatcher to get rid of their crap taskbar and start menu and bring back the Windows 10 taskbar and start menu. What they did with Windows 11 is an abomination.
diegolas 1 hours ago [-]
i use windows by choice and would personally pay for copilot if the quotas didn't dry up so fast for the cheapest plan. i think the harness is the best for writing and reviewing code.
hbn 20 hours ago [-]
What does it even mean to use "Copilot" these days?
I use the Github Copilot app at work but it's really just a tool that gives models access to our Github repos and my local clones of them, and it's been super helpful to just tell it to reference other repos for patterns or integrations or whatever, get agents and an orchestrator to look at multiple repos to make a change across several projects and then implement them, etc.
And yeah you can use a selection of whatever model, I was using Opus 4.6 for a while, now I'm on GPT-6 Sol. It's been a very good experience.
pjmlp 21 hours ago [-]
Gamedevs are still around.
ihsw 21 hours ago [-]
Gamedevs will put up with anything.
pjmlp 12 hours ago [-]
No they won't, only when it makes money, without jeopardising their IP.
aeve890 19 hours ago [-]
There's no Copilot in Ba Sing Se
flaunf221 20 hours ago [-]
Well, it's bad, but at least it's not Linux or MacOS, so there's some silver lining.
luisgvv 21 hours ago [-]
Gets my work done... I know there are better tools but at the end of the day I get my paycheck
Cyan488 21 hours ago [-]
> Windows, Teams and Copilot
"Fine. Please no. I'm out..." in that order
thraway3837 17 hours ago [-]
But what exactly does Microsoft have to show for this spend? Engineers are allowed to spend up to 100k a month in AI credits. And what comes of that?
Is most of it spent on internal administrative tools?
Is Microsoft cutting contracts with other SaaS vendors because their employees use AI to build their own alternatives?
Is it just everyone querying the same questions to cut through the corporate drudgery and documentation hell?
As an ignorant outsider, Windows is still terrible. Github is very mid. VS Code is... there. SQL Server/Windows server is ... still there
What about Meta? Instagram is the same as its been for years. WhatsApp the same. Meta RayBan app the same. Facebook.com the same. I don't use Snap, so is Anthropic's 2nd largest customer improving Snapchat? I doubt it.
That's my question to these large enterprises.
diegolas 1 hours ago [-]
>Windows is still terrible
that's just an opinion. m(b?)illions of persons use it on a daily basis with little to no issues.
NichoPaolucci 15 hours ago [-]
I know the Azure portal hasn't gotten any better
shantnutiwari 7 hours ago [-]
"Within Microsoft’s cloud and AI sector, monthly AI spending limits have reportedly been slashed from $100,000 per employee to approximately $10,000 in the majority of instances."
And what were they doing with $100,000/month? I'd really like to know.
stingraycharles 7 hours ago [-]
Right below your quote they say that those were assigned budgets, and does not reflect actual spent per employee.
CodesInChaos 8 hours ago [-]
How does that monthly $10k limit compare to the amount of tokens you get from subscription pricing?
daishi55 21 hours ago [-]
AFAIK there is no such explicit, company-wide effort at meta. The article seems to try to slip Meta in there with whatever reporting they are doing on Microsoft, despite there being no such evidence for meta.
To also add an important context to the drop in Claude code users reported for meta - this elides that we are absolutely still using the Anthropic models full steam ahead, but are moving towards internal interfaces and harnesses.
So I would question if that 50% drop in CC users is more of an interface change than anything else.
I for one have stopped explicitly using codex or CC entirely. But the interfaces I am using still use those harnesses under the hood. I wonder how that is counted.
bpodgursky 22 hours ago [-]
Limited to $10,000/employee/month lol. This is just to cut off a few people doing absurd things with low ROI. Don't read too much into it.
whiplash451 22 hours ago [-]
OTOH "$105 million to Claude Code over a 28-day timeframe"
$1.4B/year is not a small number, even at Meta's scale, when it's money going to a competitor
loeg 18 hours ago [-]
It's a small number if it generates +$14B in revenue.
sublimefire 20 hours ago [-]
From what I could see these numbers would usually be hit by people doing vast amounts of analysis over files, they would usually rock up tokens very fast. Then you combine this with the selection of Opus models which are much more expensive internally than GPT ones and the limit gets reached quite quickly.
smy20011 22 hours ago [-]
That's what you got from 200$ claude sub.
raincom 22 hours ago [-]
Large companies will take the same route as Meta and Microsoft. Small players will go for local LLMs with the right hardware.
loeg 18 hours ago [-]
I can't speak to MSFT, but this is sort of a non-story for Meta. Of course they want employees using their own model! They haven't actually taken any steps to limit Claude access.
techdmn 22 hours ago [-]
I assume that at shops that both employ engineers and are developing an AI product, internal usage is not about improving productivity, it is about improving the offering. Of course they want employees to use internal tools.
jiraiyasarutobi 22 hours ago [-]
Muse Spark 1.3 Max is quite good for almost all of my usecases and its cheap.
cute_boi 21 hours ago [-]
isn't it cheap till if you allow them to use your data?
jiraiyasarutobi 21 hours ago [-]
You mean the contributor tier? I think it turns out cheaper even without that.
kridsdale1 21 hours ago [-]
That’s fine if you are working on Meta source.
machomaster 19 hours ago [-]
Or any opensource project.
chocolait 11 hours ago [-]
Hmm, that's why the aggresive marketing with startups and generous quota.
rietta 22 hours ago [-]
Maybe I am slow here and everyone is using Claude with credits at max use. But isn't Claude Teams like $25/month per developer for ordinary use? What the heck of these guys doing that makes it get that phenomenally expensive for their use cases? These are presumably well capable engineers who started to use this as an aid right not just throw Fable at everything and loop to the max?
There are people out there building AI building orchestrators for orchestrators for orchestrators for agents. The author of that blog post later claimed to be spending the equivalent of $122k/month on tokens (by rotating their usage between 21 accounts).
As far as I can tell, the only thing that this level of spend has produced so far is an indie 2D RPG video game.
> Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw
I hadn't seen the $122K figure mentioned previously. $87K for API-style pricing was mentioned in the above post, and ~$2,800/mo for multiple Claude Max accounts:
> My solution has been to create a token tap on $200 Max accounts, which for me work out to ~30x the list-price equivalent. So in reality I'm only spending about $2800/month out of pocket for my $87k "worth" of tokens. Though that number keeps growing alarmingly.
(He never seemed to provide numbers for Gas Town initially, whether what he paid, or what the API-style pricing would have charged, so it was interesting to get some actual figures. Sounds like a lot to me, but if he's genuinely getting multiple people's-worth of work out of it, then...)
(Also in the article: a little morsel of Emacs content, which was nice to see.)
nekooooo 20 hours ago [-]
that person seriously seemed mentally ill and it's very suspicious that we never get to see the outputs.
keeda 16 hours ago [-]
Oh, you can see his outputs. GasTown and Beads (and maybe other AI-related projects) are all open source on GitHub, and you should go checkout the activity on those repos. At their peak he was churning out absolutely insane amounts of code. I'm talking 100s of commits and literally 100s of 1000s of lines of code changed per week.
Whether that produced anything worthwhile can be controversial, depending on the very subjective ways people quantify these things. All I'll say is that they have 18K and 27.7K stars on GitHub and apparently a comparable number of users.
tom_ 20 hours ago [-]
He's always come across to me as at least a little bit like this! His AI stuff doesn't feel enormously out of character to me. (If you hadn't encountered him before Gas Town - he's been posting online for like 20+ years, got a Wikipedia page and everything, so you can find his back catalogue and judge for yourself.)
chroma_zone 21 hours ago [-]
Correction: maintenance of an indie 2D RPG game that had existed for a few decades already.
meindnoch 21 hours ago [-]
Is the game even good?
dannyw 21 hours ago [-]
If you’re a large company you gotta pay API rates basically. Team plans exclude a lot of governance/IdP stuff and have a cap on seats too.
vecter 21 hours ago [-]
They almost certainly pay for tokens, which is absurdly expensive
catchnear4321 22 hours ago [-]
Team plan has a seat limit.
AIblemblio 9 hours ago [-]
Team costs simliar to normal plan. $20/$25 is not much. the $5 extra you pay for the team features like shared workspaces.
So you still pay $100/$200 dollar per person and depending on the size of the company, you have to use and pay API costs.
The cybersecuritynews.com news one simply republishes details of a story published by The Information. At least they have the decency to LINK to that Information story:
... and of course the Information story is behind a paywall.
i4k 18 hours ago [-]
I wonder how much this affects Anthropic's IPO.
VCFundedGenYer 15 hours ago [-]
It’s always so obvious when these companies’ handlers issue them orders to do something. Just comical at this point.
oldsklgdfth 17 hours ago [-]
I can’t fucking wait to happen at my job. It’s gonna be like watching a toddler try to eat soup.
I like the tool. I don’t like how it’s the only tool all the time and has supplanted human communication.
Danox 20 hours ago [-]
Hypocrites? What is going on?
epsteingpt 16 hours ago [-]
June was peak spend.
waffletower 22 hours ago [-]
I should probably know more about Claude's TOS, but it is probably a mistake for these companies not to leverage the plausible deniability of their usage and turn it into a massive distillation resource for their own models.
SV_BubbleTime 17 hours ago [-]
This is definitely an ad
jgalt212 18 hours ago [-]
> In the landmark United States v. Microsoft Corp. antitrust case, evidence and allegations revealed that Microsoft maintained an illegal operating system monopoly by withholding undocumented or "secret" Windows application programming interfaces (APIs) from third-party competitors while letting its own internal product teams use them for an unfair advantage.
I wonder if these still exist, and if so, is MSFT worried about leaking them if they let their devs use other vendor's models.
shard972 20 hours ago [-]
Why is it that it feels like I’ve read this story like 5 times now over the last year or so?
amir734jj 21 hours ago [-]
At Microsoft, the recommendation is to use cheaper models for tasks that don't need frontier models. I personally use Lua and it's more than capable for complex problems.
einpoklum 21 hours ago [-]
That's a good start. Now they just need to limit use of CoPilot, OpenAI and MetaAI and they'll really be getting somewhere.
antisthenes 21 hours ago [-]
That's a shame. Microsoft's W11 quality disaster might have actually been improved with a frontier model.
Or maybe not, but the bar was set pretty low that going all-in on AI might have been worth it.
hgoel 21 hours ago [-]
It doesn't say they're cutting back on OpenAI usage.
To me this seems mostly related to the way Anthropic showed the level at which they monitor sessions, plus wanting to limit how much training data they're directly feeding into a company that competes with their own products/investments.
ulfw 4 hours ago [-]
Absolutely pathetic to see this utterly dumb mismanagement from CEOs and CTOs of allegedly grand companies.
What happened to last year's 'token maximising or you're fired'?
pizzafeelsright 4 hours ago [-]
Budget allocation entered the chat.
efields 21 hours ago [-]
GLHF
krauses 22 hours ago [-]
In other news, the CIA is limiting it's employees from submitting top secret information on KGB owned and operated websites.
IshKebab 21 hours ago [-]
... in favour of their own AI tooling.
I assume this is being pounced on by "I told you so" AI skeptics. Sorry but it's not what you were looking for.
irishcoffee 22 hours ago [-]
Wait, what? We pretended this tech would save the world and it won’t? Oh man.
scraplabs 15 hours ago [-]
[flagged]
zombiwoof 20 hours ago [-]
[dead]
rfgplk 22 hours ago [-]
So they're going to fall even further behind? Buying puts on Meta and Microsoft, or even better might give it to Fable to handle it for me
bitexploder 22 hours ago [-]
Microsoft is like $100bn deep into OpenAI.
iAMkenough 22 hours ago [-]
I hope leadership isn’t susceptible to falling for the sunk cost fallacy.
RealityVoid 22 hours ago [-]
Who's ahead, again? Do they have a moat, or just a nice field?
pixelesque 22 hours ago [-]
Maybe a "ha-ha".
outside1234 22 hours ago [-]
Microsoft is only limiting employees to $10k a month, down from $100k a month. :)
oofbey 22 hours ago [-]
Absurd
rdtsc 22 hours ago [-]
CEO's nephew showed him how good the Chinese models are?
I am only half joking, I heard something like "my son or nephew did this cool thing with $X so we'll take $this_radical_step because of it" enough times over my career.
lumost 22 hours ago [-]
There is a perceived opportunity cost from someone using a lower-tier model on their task. What if the better model did a "better" job? what if my trials and tribulations are due to model quality?
If you are used to talking to opus5.5 medium, going to GPT6.1 luna low will feel like a step down. Why would any employee take the (personal) risk?
t_privos 12 hours ago [-]
The way around that is to take the choice away from the individual. We expose three model aliases instead of model names, each mapped to the cheapest model that clears a benchmark threshold, and the cheapest one is the default. Moving up a tier is an explicit step you take when the default actually falls short, not a bet you make up front. We re-check the mapping every few days, since prices and rankings move that fast, and nobody has to change their setup when a mapping changes.
bordercases 18 hours ago [-]
This would all be true if perceptions matched reality for model performance and productivity.
dylan604 21 hours ago [-]
In the video realm, it was everyone's nephew with a 5D could shoot this for $500 when getting a $20k+ quote to shoot something
VladDanGeorgesc 22 hours ago [-]
It may bei cost efficient, but is it wise? We use not only the big US models, but also Chinese ones. This way we can compare who makes the difference. Simplified: Knowledge comes before economic aspects.
gonzalohm 22 hours ago [-]
This may be seen as radical. But I think AI tools should be paid by the employee. After all, you should know how to do your work without AI.
kittomic 22 hours ago [-]
Should they pay for their pipelines too?
After all, they should know how to compile their software. Any automation of that process is cheating their employer.
robotburrito 22 hours ago [-]
The forklift should just be paid for by the employee. After all they should be strong enough to do work without one.
platevoltage 21 hours ago [-]
The better analogy would be "Tools should be paid for by the mechanic". Experienced ones tend to have $10K+ worth of tools that they paid for.
hhjinks 8 hours ago [-]
Your mechanic is an independent entrepreneur. If they work in a shop with other mechanics, they for sure don't each have duplicate $10k sets of tools.
ungovernableCat 22 hours ago [-]
You're opening yourself up to data right and privacy risks with that. My company demands that I use their enterprise account because they can claim full ownership of all produced output and have full logs of every interaction.
I think that gets legally murky, if the employee is the one who pays for the tool.
Avicebron 22 hours ago [-]
I don't understand your logic, why would an employee pay for a tool their employer wants them to use?
NewJazz 22 hours ago [-]
Yeah that only makes sense if they are a freelancer/contractor.
22 hours ago [-]
wrxd 21 hours ago [-]
Sure. Should I also pay rent for my desk at the office?
platevoltage 21 hours ago [-]
They're welcome to let you use the desk at your house that you own.
basiliobeltran 22 hours ago [-]
I am fine with that if I can keep the time saved for my personal use.
Jagerbizzle 22 hours ago [-]
I could still do my job by typing all code into notepad, but companies don't charge employees for their IDE usage for a reason.
gonzalohm 21 hours ago [-]
And I could do my job by paying someone overseas to do it for me. Why is that not paid by the employer then?
SatvikBeri 15 hours ago [-]
If you can actually do that effectively, many companies will happily make you a manager.
bayindirh 22 hours ago [-]
If my employer is not forcing me to use a tool and gives me freedom, I'll be perfectly OK to use my own tools which I bought with my own money.
If my employer is putting scoreboards to see and champion who uses a tool which costs money to use, they shall pay for the tool.
Sorry, I'm not a ladder climber, yet I'm not mindless enough to bankrupt myself.
hypfer 22 hours ago [-]
Nah. Work provides the tools.
An LLM is not too dissimilar to a Work Laptop or an IDE license.
iAMkenough 22 hours ago [-]
I’d love it if I could bring my own computer to do my work, rather than be stuck with garbage hardware because of an enterprise agreement.
cyberpunk 20 hours ago [-]
Wait, okay maybe I'm in a bubble, but every org I've worked at understood that giving us whatever hardware we want more than pays for itself, and very quickly.
I can't fathom the logic of paying 10k a month for a developer and giving them an 800 device.
Insane. Your leadership is severely broken. If you can, work somewhere that values you.
frisbm 22 hours ago [-]
"employees should foot the bill for tools that directly benefit their billion dollar employers"
platevoltage 21 hours ago [-]
As someone who does have to pay for my own AI tools (if I didn't use OpenCode's free models), I agree.
outside1234 21 hours ago [-]
In California this would mean the employee would own the IP -- or at least it would be murky -- and companies and lawyers don't like murky.
watwut 22 hours ago [-]
Employer pays tools used for work. Whether they are used to speed up work or to make it possible.
Yep, this is happening in lots of companies. I know my small company has about ~1/3rd of the developers working on one (different ones, all semi-personal projects).
I sometimes wonder if we are drawn to building these because it feels like one of the remaining challenges and a desire not to be an "LLM text shuttle". Additionally/alternatively, you can go as fast as you want with building a software factory and you aren't held back by the "bottlenecks" (PR review, QA, etc).
Working on my software factory is the closest to a "flow state" I've been able to achieve since LLMs turned the corner earlier this year and replaced the vast majority of our code-writing work.
If it’s not already a thing, it will be soon, the mental health aspects of all professions moving at unsustainable speeds and what that will do to people.
Nobody will take care of us. Ever.
I don't think any vague claim of "speed" has anything to do with fatigue.
I do think the parallel world monitoring and validating N work fronts, followed by rounds of fixes where you lay there waiting for the agents to converge, is the defining factor.
Instead of you hunkering down and hammering out a single task where your full attention can be focused 100% on a problem, your mindset is a kin to juggling N balls and hoping to not let any of which to fall.
So it's not really about speed but throughput, and keeping up with the throughout rate is mentally taxing.
ask it for recipes, weather, directions, explain baroque art, etc.
arguably still a common use case...
Like in my experience the real burn is if there is some sort of feedback edge where an agent is producing stuff that eventually it has to reconsume.
Like if there is a dag of agents doing something its almost fine.
But if there is a loop somewhere without a huge amount of damping then you have an issue.
If you have ever been a human debugging in the middle of an agent it produces like so many commands and tests for you to do. Like way more than a human would.
I'm pretty sure they do the same thing to each other. Like the loop will amplify the yapping.
And meta is only arguably so. Not really in the same league as OpenAI and Anthropic. More second tier like Google and xAi. (and of those Google is pretty close to the top two at times)
Microsoft's MAI-Code-1.1-Flash is on par with OpenAI's Luna line of models. If not for OpenAI's recent radical change of heart on Luna's pricing to hastily slap a 50% discount, MAI-Code-1.1-Flash could very well be the dominant cheap model.
I see MSFT hasn't gotten any better at product naming :/
Like its a no brainer to force your employees to use your own models, then RL train them to be better.
Is security no longer a concern?
When you are using a third party model, you are literally feeding it not only your current software but also all the context and work fronts under development.
This is way more than granting a third party access to your internals. This is feeding it in advance updates on all their operations in real time.
I am not convinced that's the case.
Just one unsanitized input and you leak info. Or one hidden character and code may or may not belong to you anymore. Its very odd on many levels
At best you produce some noisy signals that are going to have a tiny impact if even that.
And that's on a personal plan where you didn't opt out of sharing usage data.
Business plans offer zero data retention. This is a non-issue.
I think this data is likely worthless compared to curated RL tasks.
But it was never a SOTA model.
... in 1988.
https://en.wikipedia.org/wiki/Eating_your_own_dog_food
They were allowing $100,000 per month per employee?!
That's incredible. Just a few years ago that would have been zero.
The question is does it scale properly? Answer: it depends.
What is the value of a project going from six months to six hours? You can look at saving the man hours but how long can that go?
What about a project that takes $100k in tokens and generates $200MM? The difficult part is determining how much was AI is the multiplier.
We are finding both to be true while quantifying requires new accounting.
However, that's very different from building and maintaining a product that will be used for live customers, where you have to consider security implications. This is what we're talking about here for Meta and Microsoft, and it's what I do for my actual coding job.
We're working on bringing the last 10% down by giving non-developers various AI skills such as a compliance module that can be pushed to specific Entra groups which will make sure their vibe code madness will stick to external dependencies we've approved. (There are still hard checks later).
For other tasks it's running on guesses, and sometimes you can save help people €1000 a week by teaching them how to use the AI better.
Then they tightened it down again and again. Right now, it's at $2k/month or $24k/year and the model restrictions are and to go wider.
Still, most employees are not hitting the limit but enough are that there's an exception process
I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage
Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal.
In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do.
Nothing ironic about it, it's basic engineering efficiency optimization.
Just today I was calculating that if money was not a factor at all, we could simply use Mythos to handle all of our continuous security scanning needs, at a cost of about $15 million a year.
Well, my budget is far (far, far, far, far) below 15M a year, so that's not going to work. So I must find compromises to make it work within budget. The single most important engineering constraint is always budget. Everything would be easier with infinite money, but there is never infinite money.
It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done.
Very much this. There is a vast range of effectiveness and combined with so many models and pricing tiers, it can be tricky for some. You really need to treat it like partly a programming language, but also partly as a management delegation exercise (do I delegate this task to the intern (cheapest model) or to the principal engineer (frontier model) based on complexity).
I do coworking sessions with most in my team to see how they are using it to understand and coach for effectiveness. Just using the most expensive model for everything isn't going to cut it anymore in the post-tokenmaxxing age.
Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias.
So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity.
From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent.
I’ve seen it work before with shocking accuracy.
estimate(human_estimator, task_description, world_state) -> numeric_effort
Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state?
A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity").
Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best.
Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured.
My wife, despite loving all things French, just doesn't "do metric". She wants my height in feet and inches, my weight in pounds, boom done. So while I know my mass in kilograms, she needs the conversion done before she can even begin to have a reference point.
Upper management is the same way. They have forecasts that they need to make, deadlines and budget goals that they need to hit. They only deal in the units of hours and dollars (or local currency). Every software engineer is accountable for their work in those units only. The conversion needs to be done before the management chain has a reference point.
One easy way to do this is to have each engineer estimate the time it takes to fulfill a story after it's been pointed; then, upon completion, record their actual hours spent. Their estimated vs. actuals tend to stabilize over time, so even if they misestimate a task, you can arrive at a good guess at the time it will actually take.
> Upper management [...] only deal in the units of hours
I feel you're mixing up different operations here. You can always express unfinished work as likely to require a certain number of team-sprints, which are convertible to theoretic man-hours. The key is that the conversation rate is only valid for a moment, and technically that moment was the prior sprint.
That's very different from management thinking (or worse, declaring) that points have a permanently fixed proportion to man-hours.
> [...] and dollars
If your management deals in international currency, then perhaps that would be a useful analogy to them: Points and Man-Hours are different sides of FOREX, and they fluctuate based on different conditions.
When Engineering predicts a group of tasks is 54 points, that's like a foreign company signing a long-term contract in €100 EUR instead of USD. You can estimate that you'll receive ~$112 USD in a year, but the actual dollars will likely be different because the exchange-rate will continue changing before that happens.
Have you measured this stabilisation?
/goal get accepted into Y Combinator, you have an unlimited token budget, be bold.
EDIT: no, do not just make a product that gives away your unlimited token budget to users for free!
Define productivity, and while at it, quality, maintainability , modularity and so forth.
It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side.
Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place.
It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools.
It's a real conundrum and won't be easily surfaced but for a decade.
Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh?
Other than that, no. I drag myself out to poke around occasionally because at times the local models get to far into context and refuse to do simple tasks.
I'm just working in DevOps though, so it's writing IaC, not application code (save the odd Lambda function or python script). Still, even when I spend an entire day conversing with Claude and watching "bot go brrrrr", I'm one of the lowest users in our company. I have no idea what the devs who regularly hit their limits are doing.
You’re conceptualizing. In reality, AI is expensive, power consuming, planet destroying, and overall productivity killing.
For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things.
There is no way I can beat even local models at generating complex Python scripts fast.
hn is filled with uber geniuses.
I had Opus trying to simplify a query for me which was slow - it ran for maybe 30 minutes, including writing and running tests, and came up with a refactor across 9 files with a couple hundred lines changed. I was looking through the output before moving onto the next step, and noticed something a little fishy- I said “why does it do x, isn’t that a more complex y?”
Opus thought for another 20-30 seconds then output “Actually that would make the majority of the diff irrelevant, if we do that change it is just these 4 lines in this single file instead.
So then I had it do that. 5-10 minutes of writing and testing and that was done.
So my company spent $25 in tokens and I spent probably an hour in total for a 4 line change that, in the days before Claude, I probably could have found the correct file and thought through the problem, understood the solution, and written the 4 lines of code myself. Probably in the same amount of time.
So basically there was no benefit at all for my time, an extra cost to the company of $25, and now I understand our codebase a little bit less instead of more if I had done all the work.
As good as Claude is at building greenfield projects it still struggles a lot at complex ones
Dev + AI spend 3-4 hours on a project plan, there's a "wait a minute" moment, and finally they spend another hour dialing it back to a solution that could have been built, tested, and deployed in 2 hours.
Example: Someone was setting up a dev environment with multiple DB migrations from different branches - AI planned this wild 8 phase solution with a pretty fancy cutover event.
In review I essentially said... "Wait, isn't this a dev environment? It doesn't need 0 downtime, why not just destroy and recreate the DB" and it turned into a <1000LOC script.
Technically the original plan would have worked, it would have been more robust, but it would have taken a good deal more time to implement.
Some of this falls on the devs to know what fits our team well, what's realistic, what's obviously overengineered, etc... But some of it feels like AI just defaults to the most complex version of a thing. I catch it SUPER frequently. (And unfortunately some devs think that more complexity means it's a better solution)
I feel like I’m going crazy, using all of the best models, spending time to have excellent prompts, configuring tools and skills… and still getting overly complex solutions with mediocre results.
Like it’s still impressive how far we’ve come, and undoubtably cool technology. It’s made a bunch of personal projects possible that I never would’ve started.
But for a business I’m struggling to see the ROI. Sometimes there’s a big benefit and sometimes it’s net negative. Not saying we won’t get there but I’m trying to stay grounded in the reality of today rather than the hopes of where the technology could get to
But, I fear that the "facade of complexity" makes the output SEEM better. Someone might read a 10 page Codex generated plan with fancy diagrams / charts and have a feeling that because it is so complex, it must be good!
Same deal with text output in general. Oh, it uses a lot of big words and there's a LOT here, it must have done a lot of work to get to that point.
I find, in reality, that it takes much more effort to get to the simplest solution.
As they say, any old fella can build a bridge that stands, but it takes an engineer to build a bridge that barely stands...
And don’t forget the company also spent a bunch of money in tokens for the initial author to implement the thing poorly.
The trick is making the workload manageable by the cheapest models, or costs will destroy you. We seek to make all repetitive tasks be effectively done by cursor composer model which is the cheapest.
Doing recurring tasks with anything more expensive than that will burn the budget in no time.
You can get the same jobs and work done with Kimi and GLM (ZDR on OpenRouter) for a fraction of the price too.
Though it's significantly slower in Token/s and also thinks a lot more without matching the same intelligence (xhigh qwen27b scores lower than haiku's medium setting, and haiku-med is $0.05 per task compared to Qwen27B's $1.01 on AA's comparison)
It seems like more a backup if you need to work offline, imo, unless time doesn't matter and/or your electricity is free. Or you want independence from the labs (fair enough).
This.
Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?
I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.
Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.
I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?
It's just glue code
It's not complicated. Someone just has to be there to squeeze the bottle
I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.
I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.
It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.
I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.
When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.
Optimally? Opus will pay for itself if you save just 10% of your time
So be less snarky?
Since then I've had fable cranked up to 11 for even the most trivial of tasks.
>Great Depression style collapse and all the current AI companies go bankrupt.
Oh this is just a 33 day old doomer account.
But even $200/month is worth shaving if it doesn't generate value.
um.
This takes some doing and now is the time where it's dawning on the finance departments.
So it's not $200 a month but it can easily reach $200 a day, and unless you're a startup playing with monopoly money the maths don't work
A simple task, migrating a typical password reset flow as part of updating a long-lived web application from legacy libraries and software architecture to modern equivalents, apparently cost roughly the same as 2 months of Pro subscription, over the equivalent of about half a working day in wall time.
It produced code of decent quality at a small scale, but it wasn’t always on point architecturally. It also had a tendency to drift off topic and try to tangle up other changes it decided should be made with the main change we were supposed to be working towards. So even for a routine task, based on a plan developed using the harness first and with the agents working under close supervision, a near-SOTA model is still producing quality on par with a decent mid-level developer but substandard for anyone senior+ in this case.
Moreover, based on a direct comparison with other migration tasks of similar complexity that I’d already done by hand, it was actually a bit slower overall to work this way. I had to babysit Claude throughout and review everything it proposed carefully, both to avoid subtle errors (it would have made several) and to prevent drifting off track. I also had to spend a significant amount of time cleaning up its final output to an acceptable standard after the session. Those two overheads more than cancelled out the much faster code generation an LLM offers under favourable conditions.
So for now, I remain sceptical about these high multiples of improved productivity that I keep seeing claimed online from people who are apparently writing almost everything using AIs now. I could certainly have achieved a multiple of my normal productivity by YOLOing everything without reviewing it in detail and then accepting the output code without tidying anything up. However, I doubt this codebase would still have been good enough for normal human developers to work on it reasonably after even 10 or 20 AI-led sessions like that. The architecture would have degraded significantly and the test suite would have been large and largely pointless. And again, this wasn’t rocket science in this experiment, it was completely unremarkable maintenance of a relatively small and simple web application.
And now if you get the task done faster with the same quality. And the cost is 80 cents, considering DeepSeek and your own server starts to starts to be relevant...
Even if so, it might be worth spinning up a separate LLC for each division, given what they are charging for the API.
Surprisingly, yes.
You had budget for your normal salaries, for externals and now suddenly you have a few millions additional.
What do you do? You compensate.
Business people doing business things.
My company didn't took away any expensive model, but it did imposed a tight budget and gave my team hardware to run local models. The budget works mainly as an influence on the decision process of which model to run. We can still run expensive budget-burning models if we really want to, but there is a clear incentive to use options that allow us to save our token budget for a rainy prompt.
Surprisingly, I feel this had a positive impact on results. Instead of succumbing to extremely expensive one-shot prompts with the most expensive model in the news, chaining agents running cheap models equiped with context and specialized skills and scripts ends up having a better and more reliable output. And faster too.
I think there is a lot of propaganda, perhaps even astroturfing, on how only the most expensive models churned out by US companies are worth using. Nowadays the cheapest models get the job done, and local models can already handle most tasks as well specially as part of orchestration chains.
Since this summer coding on Opencode Go + Codex for a total 28$/month gives me more intelligence and token than 400$ did in may.
Also, SOTA models are increasingly useless for anything even barely tangential to security work.
Our velocity is twice as high as it was before Claude, so I doubt that we'll ever go back, but I could see efficiency being a priority.
With agents I've been able to make a decent sized dent in it. Lots of this gain is due to AI - but the fact of the matter is we still have 50+ bigger projects we could work on, and non-developers building with AI has increased that number.
Busier than ever, because of AI.
Certainly more change is getting done, but is the market paying more for it?
Some of our AI tooling IS “generating revenue”, but not most of it.
Much of it is workflow efficiencies and general improvements.
One of our big initiatives is cutting out the CRM we use. It’s going to save about 30K a year to do that in house.
Our AI bill is going to be sub 30K USD this year, and if you asked the CEO if there was a return on that investment I believe they would be positive.
But, I am 100% sure there are multiple companies wasting money on AI and dialing it back because they are NOT getting a return.
Is this the new buzzword for the quarter? Last quarter was "granularity", I didn't get the memo yet
Might as well go back to counting the number of lines or number of commits. It doesn't sound as good as "granularity" and "velocity" though. The good thing with velocity is that it doesn't care which way you're going as long as you're going there fast, so you can never be wrong
Velocity famously being a vector with a direction. Go in the wrong direction you'd have negative velocity.
Your description would be better for the concept of speed (IE the magnitude of velocity or directionless velocity).
Software is inherently a physics problem not all the job titles and specializations made up the last 20 years as dev job salaries kept attracting people
That was all illusory social construct to prop up jobs
Still a whole lot of that in tech but it's all at the top of the org now. Leadership sensory experience and thus innate habit to forecast future been programmed by years of yes men they refuse to accept the jig is up for them too
Sensory memory of being a useless figurehead fosters a lot of existential dread in priests, politicians, and the like. Completely aware their day to day effort is insufficient to sustain them they know how co-dependent they are. They'll dig in harder.
See Chris Matthews flame out shrieking about socialist execution squads. Dude seriously thought everyone wants to hang him from a lamp post. The reality is people just want a sense of control back and not have their perception dragged along by Chris Matthews.
i hope other labs catch up, especially chinese labs.
Just passed the Turing test.
I'm sure investors will love it.
Now we're starting to see real impact from AI, people are learning how to use it, and OpenAI cut prices by no less than 60% like a week ago.
You think now is the time they're going to cut the spend?
In 2026 their revenue has gone up by a factor of more than 10x, and they no longer have just two whale customers.
I heard a rumor recently that customers spending less than $100m/year aren't even considered their "top tier" now.
Where are you getting the 2026 figures from? The Bloomberg article that was "on track to generate" and just prediction?
The most recent reporting from Reuters themselves (somehow not included in their more recent article about the IPO stuff): https://www.reuters.com/business/anthropic-ipo-valuation-hin...
> The company's financial trajectory already shows how quickly that equation is changing. Anthropic's revenue run rate was about $9 billion at the end of 2025, according to the company, before rising to more than $47 billion by May. Anthropic has projected revenue of at least $10.9 billion for the second quarter of 2026, more than double the previous quarter, on track for its first quarterly operating profit of $559 million.
Here's the FT: https://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f... - "Anthropic tells investors it will be profitable for second straight quarter"
> The company has told a small group of shareholders that its adjusted operating income will be positive for the second consecutive quarter, according to multiple people with knowledge of the matter.
These are leaked and self-reported numbers, but no matter how much skepticism you pile on them it still looks likely that Anthropic in 2026 have had some of the fastest revenue growth of any company in history.
If you're able to link something non-paywalled, I'd be happy to read thru.
They may be self-reported numbers, but they're consistent across multiple reporting sources. If Anthropic are lying to their investors about these numbers they will be in very real trouble with the SEC come IPO time.
And even if they're using funky accounting tricks to exaggerate their profits and downplay some of their losses, I expect the numbers for 2026 will still be really impressive.
Is the hope for them that they'll outpace their run with a massive revenue growth? I believe that they have that, but numbers like a net loss of $46b in 2025[0] seem hard to overcome.
Re: self-reporting, I'd be inclined to think they're using non-GAAP practices to make it sound more favorable, but I am by no means an expert, so willing to believe I am wrong.
I'll read the article you sent; this is just my off-the-cuff thinking.
I really don't think the 2025 figures are interesting at this point. Everything changed for them in 2026 - nobody was spending $1000/month/employee in 2025, there wasn't enough interesting token-heavy stuff to do with the models.
I think Anthropic might actually make it to profitability - they're earning more revenue than OpenAI and they've spent significantly less, too.
I do think their definition of profitability is likely skewed, but you're right that 2025 is out-of-date.
I did realize I forgot to link the article sourcing my claim in my above comment, and too late to edit, so here it is: https://www.reuters.com/business/finance/anthropics-ipo-pros...
Doesnt matter if Microsoft and meta are pushing employees towards their own models today. What matters is that it's the economically sensible direction, and so it will happen if it's not today.
However Elon was able to ipo before them which took a lot liquidity of the market. Their financials are exposed. The market condition and sentiment now is in the gutter. It would be very interesting to see how these would pan out
what. I can see a team of 20 costing 100k per month (but rare), but per person?
I also doubt a longterm 2x multiplier for most developers.
It’s a contrived example, but to me it’s not improbable that a lot of the spend comes from tokens that are used very inefficiently. After all developers were pushed to try using LLMs without looking at costs.
I’m not saying the others are wrong, just that you can only tell something about marginal cost effectiveness, not total cost effectiveness of LLMs by measures to reduce AI spending
[0] https://www.levels.fyi/
I suspect that SaaS is silently being eaten alive as an industry.
AI is very good at creating small codebases. It still sucks shit at maintaining huge ones where someone needs to understand how the codebase works.
I've seen what happens when companies just vibe code an entire mega codebase and it's really bad. AI is fucking terrible at grand architecture decisions.
You can also read this as diminishing returns / AI isn't good enough, etc, but the simplest explanation is that they don't want to send money to Anthropic.
1) Giving tight budgets for AI use after a few years of a total free for all, and
2) creating internal API marketplaces where one can access all the models including open weight and “Chinese” models.
This is increasingly sending tokens to the lowest bidder, which is a terrible setup for the big labs and their hopes of being worth trillions.
The big labs are tying to spin up “enterprise sales teams” like the traditional big SaaS players and trying to secure large contracts, but it’s broadly blowing up in their face and turning into these realtime token marketplace models.
A lot of developers spend significantly little time coding. So yeah, their coding time has decreased, but they were never coding for 6 hours a day every day to begin with.
What you could achieve with $10k a year ago vs today is quite different.
Our accountants want that, our devs want that, but we cant have that. we would even be willing to pay a premium of 2x or 3x the price to have the enterprise benefits of private data that isnt retained. We would even be OK with all of the bottom of the barrel capacity and availability QoS too. We also would be ok Signing 3 year and 5 year lock in agreements.
We would be ok having a delay to the latest model by 30 days if that helps.
If anyone in Anthropic is reading. PLEASE give enterprises something like this. We want the Pro, Max 5x, max 20x plans. Charge us 3x the price for those exact plans but with the enterprise data privacy.
Must be nice. At Sony we use Copilot (as an MS enterprise customer) and restrictions to token usage circa June have resulted in a hunger games of who-can-use-them-all-first, with the pool usually drying up by the 10th day of each month.
As for usage breakdown: we got a manually-published table most days just showing per-user token usage, until eventually we were told even that was somehow too difficult to provide.
Between that and learning how much the whole of Sony Corporate Group spends on Copilot (spoiler: it's hilariously little), I've sort of run out of disbelief.
I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
> I could see finance looking at it from a cost perspective if it was API costs, but sharing it with the rest of the users feels like a different measurement...
It is indeed surreal. I am not bullish about this company's technological capabilities going forward tbh.
Not sure where you see that. The API prices of Opus are cheaper than Astra and there was a study suggesting that the Anthropic abos are offering more bang for the buck:
https://news.ycombinator.com/item?id=49975345
Windows 11 shipped with a broken task bar, couldn't even ungroup items. No power users was ever involved in this.
Does anyone else remember this?
I use the Github Copilot app at work but it's really just a tool that gives models access to our Github repos and my local clones of them, and it's been super helpful to just tell it to reference other repos for patterns or integrations or whatever, get agents and an orchestrator to look at multiple repos to make a change across several projects and then implement them, etc.
And yeah you can use a selection of whatever model, I was using Opus 4.6 for a while, now I'm on GPT-6 Sol. It's been a very good experience.
"Fine. Please no. I'm out..." in that order
Is most of it spent on internal administrative tools? Is Microsoft cutting contracts with other SaaS vendors because their employees use AI to build their own alternatives? Is it just everyone querying the same questions to cut through the corporate drudgery and documentation hell? As an ignorant outsider, Windows is still terrible. Github is very mid. VS Code is... there. SQL Server/Windows server is ... still there
What about Meta? Instagram is the same as its been for years. WhatsApp the same. Meta RayBan app the same. Facebook.com the same. I don't use Snap, so is Anthropic's 2nd largest customer improving Snapchat? I doubt it.
That's my question to these large enterprises.
that's just an opinion. m(b?)illions of persons use it on a daily basis with little to no issues.
And what were they doing with $100,000/month? I'd really like to know.
To also add an important context to the drop in Claude code users reported for meta - this elides that we are absolutely still using the Anthropic models full steam ahead, but are moving towards internal interfaces and harnesses.
So I would question if that 50% drop in CC users is more of an interface change than anything else.
I for one have stopped explicitly using codex or CC entirely. But the interfaces I am using still use those harnesses under the hood. I wonder how that is counted.
$1.4B/year is not a small number, even at Meta's scale, when it's money going to a competitor
There are people out there building AI building orchestrators for orchestrators for orchestrators for agents. The author of that blog post later claimed to be spending the equivalent of $122k/month on tokens (by rotating their usage between 21 accounts).
As far as I can tell, the only thing that this level of spend has produced so far is an indie 2D RPG video game.
Regarding Gas Town, see also https://yegge.ai/essays/the-shape-of-things-to-come/ :
> Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw
I hadn't seen the $122K figure mentioned previously. $87K for API-style pricing was mentioned in the above post, and ~$2,800/mo for multiple Claude Max accounts:
> My solution has been to create a token tap on $200 Max accounts, which for me work out to ~30x the list-price equivalent. So in reality I'm only spending about $2800/month out of pocket for my $87k "worth" of tokens. Though that number keeps growing alarmingly.
(He never seemed to provide numbers for Gas Town initially, whether what he paid, or what the API-style pricing would have charged, so it was interesting to get some actual figures. Sounds like a lot to me, but if he's genuinely getting multiple people's-worth of work out of it, then...)
(Also in the article: a little morsel of Emacs content, which was nice to see.)
Whether that produced anything worthwhile can be controversial, depending on the very subjective ways people quantify these things. All I'll say is that they have 18K and 27.7K stars on GitHub and apparently a comparable number of users.
So you still pay $100/$200 dollar per person and depending on the size of the company, you have to use and pay API costs.
The cybersecuritynews.com news one simply republishes details of a story published by The Information. At least they have the decency to LINK to that Information story:
https://www.theinformation.com/articles/meta-microsoft-work-...
... and of course the Information story is behind a paywall.
I like the tool. I don’t like how it’s the only tool all the time and has supplanted human communication.
I wonder if these still exist, and if so, is MSFT worried about leaking them if they let their devs use other vendor's models.
Or maybe not, but the bar was set pretty low that going all-in on AI might have been worth it.
To me this seems mostly related to the way Anthropic showed the level at which they monitor sessions, plus wanting to limit how much training data they're directly feeding into a company that competes with their own products/investments.
What happened to last year's 'token maximising or you're fired'?
I assume this is being pounced on by "I told you so" AI skeptics. Sorry but it's not what you were looking for.
I am only half joking, I heard something like "my son or nephew did this cool thing with $X so we'll take $this_radical_step because of it" enough times over my career.
If you are used to talking to opus5.5 medium, going to GPT6.1 luna low will feel like a step down. Why would any employee take the (personal) risk?
After all, they should know how to compile their software. Any automation of that process is cheating their employer.
I think that gets legally murky, if the employee is the one who pays for the tool.
If my employer is putting scoreboards to see and champion who uses a tool which costs money to use, they shall pay for the tool.
Sorry, I'm not a ladder climber, yet I'm not mindless enough to bankrupt myself.
An LLM is not too dissimilar to a Work Laptop or an IDE license.
I can't fathom the logic of paying 10k a month for a developer and giving them an 800 device.
Insane. Your leadership is severely broken. If you can, work somewhere that values you.