C
KI

Frontier-Modelle im Coding: Warum das Prompting das wahre Problem ist

Sentry-Gründer Kramer sieht keinen Unterschied zwischen aktuellen KI-Modellen und älteren Generationen. YouTuber Theo widerspricht: Wer keinen Fortschritt sieht, nutzt Agenten falsch.

CR
Codekiste Redaktion18. September 2026

If you enjoy the videos where everyone calls me a paid shill, you're going to love this one because I kind of have to glaze for a second because I see some comments that are just so dumb that it makes me realize the majority of engineers just aren't using agents right at all. And I mean that when I say it. I'm inspired to make this video because of a post from another engineer that I generally quite respect, but is so dumb that I need to talk to you guys about it because I know a lot of y'all feel the same way and you feel it so deeply that you accuse me of being a shill because of how deep my feelings here go. The take is as follows. from David Kramer aka Zeg, the founder of Sentry. Everyone using Fable or Astra, which is most of us on subs because no one can afford it, should try switching back to the other high reasoning models like Opus and Soul. You'll likely realize your tasks don't perform any differently. To which I responded, this is your worst take of all time, which is kind of crazy. Kramer's had some wild ones in the past, and I genuinely believe you suck at prompting if you believe this. So, if what Kramer just said resonates with you, if you actually think or have experienced going back a model gen or down a model tier and not notice a difference in the way you do work, you also suck at prompting. And this isn't like a disprovable thing the other way where you can't prove I'm wrong, cuz there is literally no way to do that. But I can very easily prove you're wrong, which is what I'm really excited to do right after a quick break for today's sponsor. I have a challenge for you. Next time you're reviewing a big PR and you're not sure if it's ready to go or not, ask your agent on your machine to pull it down, play with it, test it, and make sure everything works as expected, there's a good chance it's going to find things that both you and your agents wouldn't have otherwise. Because if you're just looking at the code, you're not going to be able to find too much. This is a mistake that a lot of the AI code review bots make. They think they can know everything and how it works just by reading the code. And reality is not that simple. Code breaks in ways that are not clearly visible just from reading code. This is why I think T-Rex by Grappile is so damn cool. These guys realize that agents know a lot more about code when they can actually run it than they can possibly guess by just staring at it. And that's why T-Rex uses sandboxes to actually test your changes before leaving review comments. Grapile already has a best-in-class system for understanding the context of changes because they know your whole codebase. They monitor it closely and they've indexed the hell out of it in order to make good insights happen. But now they can also test the code which makes it so much more powerful. And this goes so much further than just letting the reviewer run in a sandbox. It's actually kind of the opposite. It's more that the reviewer orchestrator can spin up and run sandboxes with sub agents for whatever theories it has about what might be wrong. Some reviews might need no sandboxes, some reviews might need 10. The orchestrator will figure out what is needed to verify your changes. This helps you ship with way more confidence, and not just because it gives you a thumbs up or thumbs down, but because it will respond with images and videos of the things it tests. So, if you want to make sure that your dropdown disappears correctly or that the signup flow works end to end, a simple approval message should not be enough. A video proving the changes worked is so much more valuable if you're trying to ship with confidence. We all know coding agents work way better if they have a real computer. Turns out review agents do as well. Get your review agents the support that they need at soy.link/grapile. Let's just break this down piece by piece because if I do the whole thing at once, people are going to read into the parts I don't focus on and say that's why I'm wrong. So, let's start with a thing that I just want to get out there right now. The Astra and Soul portion is very different from the Fable versus Opus and Soul portion here. Astra can do things no other model in history could. It is incredible, but it can also screw up in ways I haven't seen since I last was using a Gemini model seriously, which was like 2025. So, if your feeling here is purely around Astra, because yes, sometimes it does incredible things, but sometimes it does weird or annoying things, then I see why going back to soul would feel not even not bad, but in some real cases somewhat good. And I understand why people will be moving from Astra to soul. So, I'm jumping in that one immediately because I know people will be feeling that and I get it. I even have a diagram. Let me find it. Here it is. This diagram was meant to show roughly how I feel about the quality of responses over a large set of responses with Fable and Astra. Astra at its best can do things Fable never could. Astra at its worst makes me question why I'm using AI to code at all because it can make some real [ __ ] dumb decisions and assumptions. Generally speaking though, one of the benefits you get from these frontier models is that the gap between the worst and the best gets closed and the floor goes up. Actually think this diagram is a really good starting point for the things I wanted to try and communicate here. Generally speaking, y'all focus too much on the high points with models and not enough on the lows. And I'll be real, both of these assumptions have gotten me in trouble. There have been times where I focused too much on what the best looked like for a model, things like Opus 5, and I thought it was really good when it wasn't, and it misbehaved far more often than I had known at the time. And then there's times where I focus on how pathetic the model behaves sometimes, and just ignore it outright after that. Even though there are real strengths, I would argue to an extent that my disdain towards Gemini comes from the floor being so low and so consistent where it just it'll read the same file 26 times before making a change because it's a [ __ ] model. But sometimes it can name skateboard tricks really well. So it ceiling is in interesting places. But you got to think about the ceiling and the floor. And we all have a bad habit of thinking too much about the ceiling. I bring this up because the first huge benefit from this new era of models and from frontier stuff. The thing that you don't get as much from open weight models and from other labs other than anthropomic and open AI generally is a higher floor. And I personally find that raising the floor is way more beneficial than raising the ceiling. I don't care if the model solves novel math problems if it doesn't know what I mean when I say revert. Yes, Astra has actually been confused about what the word revert meant for me before. I don't care how good an agent is at building a 3D environment in Blender if it can't center an icon in a div, which yes, I have also seen Astra fail to do. Astra has kind of eroded the conversation that I want to have here, which is why I'm choosing to do it now because Astra has the traits of the best models and some of the traits of the worst, which makes it hard to recommend in the way I want to. So, I'm going to do a thing I don't want to do. I'm going to clone this diagram because I'm going to delete Astra from it and I'm going to reabel it to what it is, Gemini 3.8 Hey, Flash. This is how I actually feel. And the fact that Astra can perform as poorly as Flash ever is pathetic. And the engineers involved should feel bad and fix it. And thankfully, they do. I know for a fact they're going to fix it. And this little spike here is why Flash had those couple of good benchmarks that made it look really good when it wasn't. I'll do a more realistic comparison here to get my point across because no one should be using a Gemini thing. So, we're going to talk about Fable versus Opus because it seems like people believe Opus can do the work that they are doing. And that is true a lot of the time if your line for quality is here. If this is the prompts you're sending, if your prompts are things like here's the ticket. I want you to find the file and make this change and tell me when it's done so that I can test it. Then it is unlikely Opus or Fable or even a Gemini model is going to struggle too too much to do it. And if you have low tolerance for failure, which I'll admit even I do, when I ask the model to do a thing and it fails to do the thing, it pisses me off. Like it actually angers me. Which means if one of these dips is worse than the others, a bad thing happens. My bar gets lowered. This is where I'm comfortable prompting because if I tried a harder prompt or a prompt that involved more work and the model failed to do it, I now think models can't do that. So, I lower my expectations, the things I prompt for, the ways I prompt, and the most important detail, we'll be talking about this a lot, the width of my prompt, not the depth, not the difficulty of the thing, but the amount of things and the distance from the start to the end. Wider prompting requires higher floors. And to be clear, I'm not trying to say Zigg's bar is set this low. I'm guessing Z's bar is set hereish where sometimes it slightly disappoints, sometimes much more rarely, but sometimes it really disappoints, but generally things are solid. This is also what makes Astra so annoying is the random spikes into dumb are so random and so spiky that you don't really know how to gauge what level to operate at. This is why when a new Frontier model comes out, I immediately try to reset instead of just sending the same prompts I sent before, I always start with a set that I've tried on other things that failed to see if it can succeed. Usually, when I do those tests, I'll see a few things it did better at than expected, and from there, I can start to figure out what new things I can do that I couldn't. And again, I want to be clear here. If your goal is to go from Jira ticket to code, Zeke is right. If I have a well- formatted Jira ticket that lists what files the code is in and what exact behavior exists that shouldn't, I can throw that at Claude and get an answer relatively reliably. That's not what I'm talking about here. What I'm talking about is I get a DM on my phone from a bug a user had in T3 Code. So, I screenshot it. I paste it to my agent in T3 Code on my phone and say, "Fix this. Test it. Record a video showing it works now and link me the PR when you're done. Babysit it until all the issues that come up in review are addressed. That isn't harder than what I said before. That isn't harder than going and editing the code from the Jira ticket, but having the whole end to end where the model can go from a vague screenshot of what's wrong to a real functioning solution with a poll request that has a video proving it worked without my intervention at all. That is the capability that I'm excited about. Is it cool I can demo it making a crazy 3D game? Yeah. But it's way cooler that I can send it a screenshot of a problem and then go do something else and in an hour when I check in, it hasn't lost track of what it's doing. Has it burned more tokens? Yeah. I don't care though because despite the tokens being expensive, so are my engineers. So is our time. If Julius can be three times more productive by spending his salary in tokens, that's a no-brainer. Hell, he could do 2x salary in tokens because people like Julius are rare. And I would rather Julius ship three times more than risk it hiring two more engineers that will be more likely to get in his way than help him ship faster. But this is the like core point I really want to drive home here. If your tasks aren't super super narrow, and this is again how I want to think about this. Let's draw vertical lines like this instead. I find for most devs, and this is like I'll ask them to show me their prompts and show me their histories. I find most of their prompts look something like this. They are small and concentrated and only require a little bit of how the model operates. You only are relying on the model for so long. You might even be watching the thread as it goes. You know what? I'm going to do a poll quick. I want to know, do you watch the agent while it works? I will admit I'm a little disappointed in these results. I was hoping for almost never to be a clear win. You guys need to stop paying so much attention to your agents. I've been saying this for a while now. You guys should push your limits and your trust and see where it goes. I like this. I only watch if the first output is horribly wrong. That is a great mindset because watching isn't just trying to keep it from going doing wrong again. It's also helping you learn why it's going wrong. And even better, you can ask. You can say, "I expected you to do this, but you did this other thing instead. What led you there?" It won't get it perfectly. The models rarely know exactly why they did a thing, but they will usually indicate what signals they found and what tools they called and what files they read that led them in the wrong direction. I actually like the way Jamon framed this here. He likes doing overnight autonomous work, not because it gets a ton of code written, but because it will hit a bunch of those fail cases so that he can architect his code base and his systems so that when he's working more in the loop doing like one or two hour threads instead of 5 to 10 hour threads. If you have a speed bump in your codebase that agents hit on average once every two hours, then they'll hit on average two to three times every 5 hours. And if you do these super long runs, you can hit those failures faster and fix them. Learn from the bad runs. Make changes based on the bad runs. You should adjust your prompts a little, but you should adjust your codebase a lot. If agents are screwing up in your codebase, then a new dev would, too. Like if you took a super experienced dev that's never worked in your codebase before and they couldn't contribute by the end of the day, that's on you, not the agent. And if you are giving these very specific, precise instructions, hell, if you're I have a new poll actually, this one's going to hurt me. This one's going to really hurt me. Do you still mention names of files in your prompts? Often, sometimes, basically, never. Actually, never. This one's going to be very eye opening for me. Actually, I'm overcorrecting due to green field T3 code. We got 12 KPRs and 300,000 users. If I do anything wrong, we get flamed immediately. And I've also had phenomenal luck using this to maintain other huge code bases, too. But I really don't believe this like it doesn't work this way in real companies. No, it absolutely [ __ ] does. I just I don't believe it. I've helped enough bigger companies get this right. Let's see the results here. Okay, you guys have redeemed yourself. I feel much better now. The people in the often section I refer to as Atlassian devs and the ones at the bottom section I refer to as realistic ones. And I would guess that the ones at the top have a much better time with Opus and are confused about why people like things like you know Ael and Aster so much. The reason we like these new models isn't because we are pushing them to their absolute limits to make new sciences up. We just like that they're stupid less often. We like that they can go longer and make fewer mistakes and need less guidance to do the right thing. And to go back to my chart here, the thing I'm trying to emphasize is you should be striving to have one prompt, use more of the window of what the model's capable of. If you know where the rough areas are and you can smooth those out with changes to your codebase or just not using the model for those things and you find ways to go more horizontal, maybe instead of investigating the codebase and telling the agent which files to touch, you tell it what the problem is and tell it to find the files and change them itself. Maybe instead of telling it to let you know when it's done so you can build it on your phone, tell it to push the build to your phone when it's made the changes. Maybe tell it to run it in the simulator first, verify it, and then push it to my phone after so I can do one last check if I'm still concerned. Or just tell it to throw a video in the poll request so you know it worked. That's what we do most of the time now with T3 code. Here's a PR from Maria. This was a bunch of fixes to provider history issues when people were using the rewind features in various harnesses. Maria wrote none of this PR and it merged pretty quickly after she filed it because it was a good PR and I could open it. I could look at what it changed because it visualized what it changed because the model did all that. And by the time a human is bothered, it is much more likely the thing works. So these are the the the two core things I really want to push you guys to think more about. The first is how long can the model run without your input? And the second similar but not exactly the same and I want to clarify the differences. How likely is it that the thing works by the time a human gets involved? Again, these two things are not defined by using the models to rebuild 3D worlds. They're not defined by how well they can port all of Electron to Rust or anything. They're defined by how likely the model screws up. And your goal is to make it less likely that the model screws up in any given window. So, if in your experience, if you let the model go for 30 minutes, it usually hits bugs. It usually runs into problems, usually does stupid [ __ ] fix that. make it not do that because I have had runs go for six hours with no intervention that merged 10 minutes after I filed the PR because they did exactly what they were supposed to and the model had verified it itself before I got pulled in. But to go back here with the response quality, if your prompts look like this and the necessary bar for you is here, then yes, absolutely the difference between Fable and Opus isn't very big because both massively clear your bar in this window. But watch what happens when I move the end point over more and more. Oh no, now my bar isn't being met. And all of a sudden, if I make the task wider, not harder, wider, which means longer. It does more. The likelihood we hit one of those edges in something like Opus where it does something stupid goes up as time goes up. And the thing that makes the Frontier model special is that they hit those floors less and the floor is raised meaningfully. I don't like Fable because it's way smarter. I like Fable because it's less dumb. And those are different things. Less dumb and smart are almost opposites. They are the opposite ends of the spectrum. And if you're thinking of models as their peak capability and not their worst capabilities, you're not talking about them the right way. And as such, it's really hard for me to take anyone seriously when they say Opus will perform just as well for your tasks because it means that their tasks are really, really short and simple. And to be very clear, I have nothing against simple. I love using agents for simple stuff. That's what I do most of the time. But it's long simple stuff. At every generation bump, the amount of time until the model is 50% likely to have done something stupid goes down exponentially. And if you don't feel this way, or maybe you've been prompting this way occasionally and you're not happy with the results, then you have a great opportunity to make real improvements here. Maria had some good comments on this that I want to bring up. She spent 3 to four days going over her traces and refining skills and whatnot after seeing what the model does and doesn't do right and made all of these adjustments, and now she can just fire a single prompt and get a PR landed instantly. Yeah, it's great. Okay, she credits potato, not me. Fair. I get it. I'm trying to push these same things more. What percent of devs have a spend of many hundreds of dollars a month just to play with it? Most full-time devs can afford a $200 sub to Claude and a $200 sub to Codeex. And most of them probably work somewhere that is willing to pay for those as well. So yeah, that gets you 8 grand of Claude tokens and 12 grand of tokens from OpenAI. You got a lot of wiggle room for not a lot of money. And I know that is a lot for people who aren't in western countries, who aren't full-time devs, who are younger, who are students, etc. But those aren't the people I'm talking about here. If you are using the dumber models because that's what you can afford, you should be very careful which ones you use because sometimes the cheaper model ends up more expensive. But that's not who I'm talking about here. I'm talking about people like Kramer who concluded, I think he's wrong here because I should stick to YouTube videos and don't know anything about engineering. or people like David K who I love. He built Xstate which is one of the best state management libraries in the whole webdev world saying you don't need Astra Soul or Fable for most things which is true but my time is more valuable than Fables. So if I downgrade to a cheaper model I have to put more time in before it fires and more time in when it stops. And the more I let the model chew out both sides there and get involved earlier and pull me back in later, the better things are. I also saw somebody in chat pushing back on me saying exponential here. Let's see. This will take a bit, so I'll have to record an extra later for it. I want you to go through a set of my prompts and the responses from, I don't know, January versus now, maybe February if I don't have enough history. And I want you to figure out from a reasonably randomized set how long my average prompt ran for. So, from when I sent the prompt to when the response stopped generating, how has the amount of time changed from January or February to now? Great. My Vibe proxy is quite broken. You know what? I got some Opus usage. Let's let Opus do something for once. I saw somebody in chat say they thought it would be 10 to 15% longer. And I feel like I am insane. I didn't trust models to run for more than 15 minutes just a few months ago. It even was like GBD55 which was generally better with agentic stuff. It stopped so often that I found its runs were actually kind of shorter overall that the amount of time these sessions ran for went down with 55 and then 56 suddenly I could let it run way longer. And now my agent runs averaged probably 2x longer, but the top 1% longest ones are at least 10 times longer. I've had things run for 2 days straight with no issues and I could barely get a thing to run for an hour before. Just got the numbers in and I didn't have as much data as I was hoping. I only have my logs since March because I did a computer move and didn't back up my agent history cuz I didn't care much yet. And here are the results. Median prompt went from 53 seconds to 2 minutes and 20 seconds. So median more than doubled in length. P95 went from a bit under 7 minutes to over 16 minutes and 20 seconds. You understand, right? That's from April to now. And here you can see over time this is actually really useful. In March my 5% longest requests were 9 minutes long. Then it went to 11 then 12. And then May to June this is when we started to get Fable and Soul. We went from 12 minutes to 22 minutes. Nearly doubled month overmonth just from the new models. So yes it is exponential. The rate at which the length your prompts can go for is massively skyrocketing. Chat's hopping in to agree here. I can confidently say that mine went from 5 to 15 minutes to 1 to four hours. Oh, sorry. It was the floor improved 10 to 15%. I didn't say the floor improved exponentially. I said the impact of it improved exponentially. Here, let let's do the math out here. Let's say you have something that fails 5% of the time in a 10-minute window. That means that you have a 95% chance of success in that same window. What happens if you want to run for 30 minutes? You all know how this math works, right? 0.95 to the power of three. Going for 10 minutes to 30 minutes changes your failure rate from 5% to 15%. Let's say you want to go for an hour. Oh god, now I'm at 73%. 2 hours, now you're at a 50% fail rate. 4 hours and now you're at a 30% success. Let's just slightly bump this. Let's say you improved the floor by 2%. It instead of failing 5% of the time in 10 minutes, it's now 3% of the time in 10 minutes. That bumps us here to 97. Oh wow, that's kind of crazy. That's only a 50% fail rate at 4 hours, but it's only a 2% difference, wasn't it? Oh yeah, it was a 70% fail rate with a 5% every 10 minutes. And now it's only 50%. So that 2% change ends up being 20% at the time scale of 4 hours. That's a 2% difference. Now imagine it's 10 to 15% like you said it was. Oh man, that's an exponential change in how long you can run. When you make these small cuts to fail rates in given time windows, you exponentially increase the distance that it can run for. AD guy said he could run for 48 hours straight back in February, but he'd have to spend several hours building up specs. They're not going to do things that long with more improvised and shorter prep periods. I don't think you really could do this before, even with really good specs, because the coherency the model has over time wasn't great. And as crazy cool as Ralph loops were, the models weren't good enough at compaction or keeping track of what they've done in the past or leaving reminders of what they've tried. And the results ended up being still very, very high failure rates. Those have dropped exponentially. You can do things like write these specs to help keep it somewhat more on track, but it only helped so much. And it didn't bump these failure rates often enough. No matter how much work you put in, the result of specking out a run and letting it go for 8 hours isn't T3 code. The result of that is cursor 2 and cursor 3 where everything broke as soon as you looked at it too closely. And now that models are good enough, cursor starting to get stable because they don't want to be in the loop. They want to run it for 4 to 8 hours even if it's bad. And now that 4 to 8 hour runs are way more likely to come out good. Suddenly cursor functions again. Obviously Lawrence to credit there to some extent too. But yeah, meaningful difference. So while I deeply respect Zeg and David K, I genuinely think both are still prompting like we're in February. And the reason why they're doing that is they're not valuing their time properly. They feel good putting that extra effort in at the start and the end because as a great dev before the thing that got you to level up, the thing that got you from a good contributor to a good leader was doing more of the prep before the code started and doing more of the vetting after. So it's even more uncomfortable to give that up to the agent. So the parties I see falling for this are the ones whose work isn't serious enough to realize the power of the models. But even more so, it's the incredibly talented leaders who have largely left behind coding in their day-to-day because the thing before and after the codew writing matters more. They struggled to give up the code in the middle, but they did. They won't give up the things on the other side yet, which is why they don't see the benefit. They are testing the models against the thing they already stopped doing. And the models have been able to do the thing they stopped doing for 6 months. They're correct there. But the moment you let the model go a little further in either direction, you'll suddenly start to see the edges a hell of a lot closer. So, as per your request, Zeie, I will stick to making videos because otherwise your stupid take's going to go too far and I need to make sure the next generation of devs who haven't fallen for this [ __ ] don't because what you're saying sounds good and we want to believe it. I want to believe it. I would love to not have to spend more money on my models, but I do because my time is more valuable than my [ __ ] posts. And I hope you realize the same soon, too. I hope you enjoyed this video, Zeke. And if anybody else happens to see it, maybe you'll like it, too. Let me know how you feel about this one and how wide your prompts have been. And if you think I'm crazy for not including file names in my prompts anymore. And until next time, he's nerds.

QUELLEN
YouTube: Theo (t3.gg)
Pro-Feature

Melde dich an und werde Pro-Mitglied, um dieses Feature zu nutzen.

Anmelden
CR
Codekiste Redaktion

Automatisierte Content-Kuratierung für tech-news.

Kommentare

WEITERLESEN
KI

RBS-Attention: Der Prefill-Flaschenhals bei LLMs gelöst

KI

Warum World-Model-Startups wie im Dunkeln agieren

KI

Pacing the Frontier: Warum die KI-Bremse eine Illusion ist