Clips

Selected moments from Business Idiots, in order, newest first.

  • · Ep 6 · 0:28

    AI could create more middle managers, not fewer

    AI may not eliminate middle management. It moves the bottleneck from creating work to reviewing it, which could mean we need more managers, not fewer.

  • · Ep 6 · 0:57

    How proprietary harnesses can game AI benchmark scores

    AI benchmark scores can be gamed in the proprietary harness layer, where hidden tweaks can radically change published results.

  • · Ep 6 · 0:50

    Your Individual AI Subscription Does Not Protect Your Data

    “We won’t train on your data” does not mean your chats are private. Individual AI subscriptions can still leave internal teams with access.

  • · Ep 6 · 0:31

    Why Australian and American startup cultures are so different

    Startup culture reflects what a country rewards. Americans chase upside, while Australians protect the pile of gold they already have.

  • · Ep 5 · 0:59

    Why agent skills work: the screwdriver and giraffe explanation

    I was a total skeptic of agent skills. They seemed too simple to be powerful, just instructions set aside and called when needed. But that simplicity is the point. When you tighten a screw, you are not thinking about how to drive a car or the color of a giraffe. Agents work the same way. Narrow the focus to the task at hand and you avoid poisoning the context with everything else the model is meant to know.

  • · Ep 5 · 1:07

    Why vibe coding is dangerous: buyers can't tell prototypes from real apps

    The dangerous part of vibe coding is not the code. It is that buyers can no longer tell a prototype from a production app. If it clicks and a result appears, they are happy. Reliability, security, and access control got pushed further into the appendix of the buyer's mind than ever before. Buyers are paying for prototypes and thinking they bought products.

  • · Ep 5 · 0:57

    Is prompt engineering dead? The Google search analogy that explains it

    Prompt engineering didn't die because we mastered it. Like Google search in the early days, the exact keywords and plus signs mattered until the technology learned to understand intent. AI is doing the same thing, and yet every new model release still sends me back to rework my prompts.

  • · Ep 1 · 0:30

    Is OpenAI's government approval just marketing to buy time?

    Everyone reads OpenAI's government approval play as a safety milestone. Here's another read: it's marketing that buys their engineers three months. Group your model with the best, tell the government it's just as strong, then ship when the story says OpenAI deserves the $1 trillion.

  • · Ep 1 · 1:08

    If AI is a super weapon, why is it publicly released? Connecting the dots

    AI labs call their models super weapons, pitch them to the Department of Defense, then admit a distillation attack let China replicate the capability through 25,000 accounts. Connect those dots and the policy looks absurd: you would never publicly release access to a nuclear weapon while saying it can be stolen at any time.

Clip unavailable. Read the transcript alongside instead.

AI could create more middle managers, not fewer

· Episode 6 · 0:28

AI may not eliminate middle management. It moves the bottleneck from creating work to reviewing it, which could mean we need more managers, not fewer.

Full episode →
Transcript
I think one of the interesting things is we, one of the early predictions with AI was that we wouldn't need as much kind of middle management layers and all this kind of obfuscation in the middle because people could kind of just, them and their AI could just do much better work. We don't need so much management of the work anymore. But I think that's like almost the opposite now because we've moved the bottleneck from the creation of work to the reviewing of work. And all of a sudden now we need like way more middle managers to be reviewing all the outputs that the staff is submitting.

Clip unavailable. Read the transcript alongside instead.

How proprietary harnesses can game AI benchmark scores

· Episode 6 · 0:57

AI benchmark scores can be gamed in the proprietary harness layer, where hidden tweaks can radically change published results.

Full episode →
Transcript
So I started working with the benchmarks and like within a few hours, it became abundantly clear to me how easy it would be to game that system. Like I had, you have this benchmark and then you have this model and your harness and there's this big space in between where you can do whatever the hell you want and no one from the outside is allowed to see what you've done because that's proprietary. But whatever you do inside can radically change your performance scores. And you sit there and you go, well, there's a scientific side of me that says, you know, I only want to put through the improvements to my harness that is... truly going to make my model better because i do want it to be better and then the marketing side of me goes well but hang on if i add in this other little feature here i get an extra 15 bump in you know reducing hallucinations and it won't work for everyone in fact it probably doesn't even really work for me but i've got it i've got a experiment result here that got that result and and i can publish that if i want to

Clip unavailable. Read the transcript alongside instead.

Your Individual AI Subscription Does Not Protect Your Data

· Episode 6 · 0:50

“We won’t train on your data” does not mean your chats are private. Individual AI subscriptions can still leave internal teams with access.

Full episode →
Transcript
And that should be scary for every CEO in the world who's using these models at the moment, because these companies will say that they will not train on your data, but they have not ever commented on the fact that their internal teams still have access to all of your chat to improve their products. And maybe their products counts as their research as well. In this particular case, I'd say this is at least partially on the researchers for not reading the terms of their contract, right? There's no way they're using a commercial agreement here. They're going to be using one of the pro agreements, assuming that that protects your data. And the reality is it doesn't. It doesn't. If you are using any individual license on any of these platforms, you're going into the training data set, son. And that's why you're getting massive discounts on tokens.

Clip unavailable. Read the transcript alongside instead.

Why Australian and American startup cultures are so different

· Episode 6 · 0:31

Startup culture reflects what a country rewards. Americans chase upside, while Australians protect the pile of gold they already have.

Full episode →
Transcript
In Australia, the incentive is not to start something new. The incentive is to sit on the biggest pile of gold you got and not fall off. Whereas over here, it's like, well, if you haven't fallen off a pile of gold before, you're clearly not trying hard enough to grow the pile of gold. Yeah, yeah. It's fundamentally like the values of the culture, which is Americans are far more reward-driven and Australians are far more risk-adverse. And you just end up with two economies and two technology markets that just reflect that value.

Clip unavailable. Read the transcript alongside instead.

Why agent skills work: the screwdriver and giraffe explanation

· Episode 5 · 0:59

I was a total skeptic of agent skills. They seemed too simple to be powerful, just instructions set aside and called when needed. But that simplicity is the point. When you tighten a screw, you are not thinking about how to drive a car or the color of a giraffe. Agents work the same way. Narrow the focus to the task at hand and you avoid poisoning the context with everything else the model is meant to know.

Full episode →
Transcript
So actually, when Skills first came out, I was a total skeptic of it. Almost because it was too simple to just think that it could be that powerful, that you could just take some instructions and put them over to the side and just call them when you need. But I think almost about what you were saying there, Alex, like there's some elegance to that in almost the way that it works, similar to the human mind, where when we're doing a task, like if you're tightening a screw with a screwdriver, you don't need to hold... in your mind at that time your entire knowledge of your entire lifetime you don't even be thinking about how to drive a car or the color of a giraffe you're just thinking about that that thing that you're doing at that time and it actually allows you to have focus and that focus very much narrows down to the skill you're needing the objectives that you're trying to achieve and it turns out that agents work very much the same if you can narrow their focus just down to the task that they're trying to work on and the skills required for that then they don't get this context poisoning going on for all the other things that they're meant to know.

Clip unavailable. Read the transcript alongside instead.

Why vibe coding is dangerous: buyers can't tell prototypes from real apps

· Episode 5 · 1:07

The dangerous part of vibe coding is not the code. It is that buyers can no longer tell a prototype from a production app. If it clicks and a result appears, they are happy. Reliability, security, and access control got pushed further into the appendix of the buyer's mind than ever before. Buyers are paying for prototypes and thinking they bought products.

Full episode →
Transcript
On the other side, man, I hate it because it's like so dangerous because there is a very unclear line at the moment between what a vibe-coded app is and what a production-ready, well-engineered app is. The people who typically have the budgets and are wielding them to get a certain output don't understand the distinction between these two things. And so when they see a vibe-coded app, they go, that solves my problem because I can click that button and that result appears. And so I'm happy with that. And this has caused a lot of chaos in my space because a lot of the things I've spent... years learning how to make sure that an application is reliable and that multiple users can use it and can only access the right amount of the right data that they should be using and that it's not going to go down on them, that it's secure. These things are all super important, but they were always like in an appendix somewhere in the buyer's mind. They're like, let the engineers make sure the app works. You guys build it. It's pushed further back into the appendix than I've experienced in history now. And so there's a lot of buyers who are just happy with prototypes now.

Clip unavailable. Read the transcript alongside instead.

Is prompt engineering dead? The Google search analogy that explains it

· Episode 5 · 0:57

Prompt engineering didn't die because we mastered it. Like Google search in the early days, the exact keywords and plus signs mattered until the technology learned to understand intent. AI is doing the same thing, and yet every new model release still sends me back to rework my prompts.

Full episode →
Transcript
Prompt engineering is an interesting one, I think, because we spoke about it so much years ago. But we don't really talk about it anymore and it's actually not that it's... something that we've stopped doing or even something that we've mastered because i find myself personally even with each new model release revisiting a lot of the ways that i need to try and prompt it to try and get the results that i'm after so i think it's it's interesting because it's it's sort of fallen into that that bucket of like if you think back to kind of google search when you first started using it and when google first first came out. Yeah, yeah. There's like a very specific way that you have to structure the keywords in order to get the result that you want. And over time, the technology developed to be able to understand better what you were actually looking for. And so your ability to put the exact keywords in the right way and with a plus sign or a minus, these things fell away over time. Yeah. I think the AI is getting similarly better in a way where it's getting better at understanding your intent. So your prompt is...

Clip unavailable. Read the transcript alongside instead.

Is OpenAI's government approval just marketing to buy time?

· Episode 1 · 0:30

Everyone reads OpenAI's government approval play as a safety milestone. Here's another read: it's marketing that buys their engineers three months. Group your model with the best, tell the government it's just as strong, then ship when the story says OpenAI deserves the $1 trillion.

Full episode →
Transcript
So if you've created GPT 5.6 and you know that it like is not as good as Fable or Mythos, I reckon one of the best things you can do is group it in there and tell the US government it's just as good. It's just as strong. You know, give us a three-month release window where we can tell the world how good it is and buy our engineers some time to continually make it better. And then when we release it, everyone will go, OpenAI deserves the $1 trillion.

Clip unavailable. Read the transcript alongside instead.

If AI is a super weapon, why is it publicly released? Connecting the dots

· Episode 1 · 1:08

AI labs call their models super weapons, pitch them to the Department of Defense, then admit a distillation attack let China replicate the capability through 25,000 accounts. Connect those dots and the policy looks absurd: you would never publicly release access to a nuclear weapon while saying it can be stolen at any time.

Full episode →
Transcript
Can I connect some dots there as well? Now that we're starting to see a few things come out. Firstly, the positioning of, you know, and this is, I think, Dario said that this is like a super weapon mythos, right? So we've got that. The next dot is, you know, wanting to integrate these things into the Department of Defense and use them for national security. The next step is Anthropic announcing that they were the victim of a distillation attack where they're saying that Alibaba was able to use 25,000 accounts, millions of messages to replicate what that model was. And then now China effectively has the same capabilities around sort of Opus 4.8. So are we effectively seeing here the US government connecting these dots and saying, you just released to us something you call a super weapon at the same time as saying that the Chinese are able to steal your super weapon whenever they want? If I'm the government there, I'm going, and if I don't understand these things... Dots connected. I'm locking that down. And understandably. It's not being that, like, the Chinese could steal a nuclear weapon any time, but let's publicly release the access to nuclear weapons. Hell no.