Transcript
Will: everyone and welcome back to week 5 of Business Idiots podcast. I'm Will Turner here with Alex Stenlake and Jim Lovell. Welcome back guys.
Jim: How are we guys? I'm loving it for a Sunday. Yeah, fantastic.
Will: Beautiful man. Will's not enjoying it for a Sunday. A little bit hungover today but that's just the nature of filming on a Sunday. There's a lot that happened in AI this week, so let's dive into it.
Jim: As always, it doesn't...
Alex: It never stops.
Jim: But even like a low week, you'd call it, because again, every report you seem to look at was just, oh, well, they're all still going on about how they're going to regulate it. Just move on, for crying out loud. Let's talk about something else. Let's talk about how good the models are.
Alex: Well, I love that Anthropic was very quick out of the gates to have a little bit of... We did a hack, too. Yeah. Off the back end.
Jim: Well, do you hear... Do you see the... Have you heard the audio of Sam getting yanked by the PR managers? No. And he starts to... Some reporter asks him something about, oh, do you think there could possibly be other hacks? And he goes, yeah, yeah, sure, possibly. Like, we don't... You know, we can't possibly... And you can just hear the PR guys, like, just... Just Kurt Hawk just yanking him in the background. It's just sort of, oh, no, no, we do not want to. Yes, it is possible. Of course it's possible. But we don't want to be admitting it.
Will: Yeah, that's the new headline, right? It's possible that there's many other hacks.
Jim: And so to me, that was the story of the week, was just they're all just talking their book and... getting out in front of stories or tripping up over stories. Let's get back to the models. Let's get back to what we're supposed to be doing here and how you hooked everyone into this, you know?
Will: Yeah. Here, Jim, I'm kind of sick of just all the geopolitical and regulatory talk. These things are important, but there's so much else going on, so many other exciting things and so many developments in the world of AI and how you can use it. In even other areas of technology, we'll touch on a couple this week as well, areas where things have been improved even without AI and they're still noteworthy. That's it. Let's get into those.
Jim: For sure. And I think there's so much potential here and everyone understands that there's so much potential and that's what they're all talking about and that's what all of this investment is about. But everyone gets lost in talking about the investment rather than actually the problems and the opportunities and the potential that is there. And so... Let's, again, I just think everyone should shut the fuck up and get back to talking about the potential and fixing the problem.
Alex: Not one for mincing words, are you?
Jim: Well, again, I'm just, I literally just, when we had another week talking about how we're going to regulate AI and And are we going to ban Chinese models? And it got down to is that, oh, well, it's actually going to be. Oh, you have to definitively say that, oh, no, it's just Chinese models we're banning because everyone's realized how good open weight models are going to be. And so it's just shut up. I'm sick of it. And so let's get back to potential.
Alex: Yeah, I think that's kind of lost in the drama of the last couple of weeks. No one's really asking the question at this point, like, what are we doing with all this stuff? That's it. Where's all this mountain of money going? What are we computing from?
Will: Okay, so let's start off this week then with a good news story. Let's talk about Neuralink. So Neuralink, for those who don't know, this is a lab, you'd call it a lab? Sure. A startup that has been in the space of putting a chip in your brain and helping to connect it up to a computer so that you can... you know manage your computer or things in the real world just with your thoughts and they've had some amazing progress over the last few years and specifically working on you know the first sort of scenario being you know people who are paralyzed and you know maybe unable to sort of interact with the world now being able to have some of that interaction again And we saw a couple of years ago, some of the first, I think the very first Neuralink patient being able to interact with a computer by thinking about what keys they might press on the keyboard or what mouse clicks they might do. And it actually kind of, the chip reads the impulses in the brain and then they've correlated that with machine learning into what actually they'd be trying to do with their hand and going, okay, they're trying to click it, so we'll click. But there's some good progress this week because they've now managed to sort of hook up that solution through to some of these wheelchairs, powered wheelchairs for 26 of their patients, of their participants, I suppose. And we've actually got people who are able to move their wheelchairs around with their mind now.
Jim: Which is really exciting. And that's the sort of thing, isn't it? And so if you... just to break down how it works, they insert, again, near microscopic wires into all of the parts of your brain where the signal is transferred, you know, if that makes some sort of sense. And so there's thousands of little wires that then go to the exact point in your brain where the synapse fires to click the mouse button or whatever. You know, again, I don't really understand it, but that's how I understand it. And so it's sort of they put it all in, and then there's the chip that controls it all, and that then communicates with... an external computer, and that was the first guy, and that was the first step. There was a whole lot of training involved in taking the signal from the wires and the synapses to then being able to control the laptop. and to me that was the greatest thing about the wheelchair store is that they've done it for 26 in one in one hit so they've been able to normalize the training of it all a bit which which is really really interesting if you think about it because six different brains that's it and that's the that's the bit that's boggling my mind that you can implant wires into 26 different people's
Alex: We've known about areas of the brain that do certain things for some time. I think we're okay with that.
Jim: But the specificity... And that's what really... And so it must be something in the optimization of the training model. Because if you think about it, you learned how to pick up a mouse or how to type in a different way that I learned how to type. And so the... The connections and the synapse connections in my brain are different to the synapse connections in your brain. And so no matter what, as many wires that you're inserting in, you're still getting something firing, measuring it and then going, okay, well, that's what they're trying to do. And then optimizing the training and then being able to do it for 26 hours. or wheelchairs at the one time.
Alex: I'm just amazed that we can do this. For a single person, sure. Jim and I have worked with healthcare devices in the past and there's a massive problem when it comes to building these devices in that
Jim: individual people are massively different even for relatively simple things like their heartbeat yeah yeah the the way that their heartbeat shows up yeah the the heart rate variability measurement you know like and that's the sort of thing is that it takes and it's sort of why you know like even for your aura ring or your whoop or your apple watch it takes a couple of weeks for it to learn you for you to for it to learn you because everyone's that it even just the variability in the measurement.
Will: I wonder if this went through a similar process. We say 26 is amazing. It doesn't mean that they built it once and just put it in 26 people. It could be 26 people underwent six months of training the algorithm.
Alex: I'd love to see the white paper.
Will: That's it. Off the back of this.
Jim: I couldn't find it. I had a look.
Will: Again, I just...
Jim: a bit of a hoarder in that way. Hoarding white papers? No, but hoarding the information. I hoarded Kimi K3 this week. Again, I haven't got a GPU that can run Kimi K3, but I've got Kimi K3 because I've got a hoarder.
Will: Who knows? In the dystopian world.
Jim: I'm going to have some monkeys on bikes in the back running a generator to run my service so that I could, you know. Or maybe it's the last open source. Who knows? Who knows? And so that's what I love is that they found the way to optimize the training with Neuralink so that they can then get it further. And that's sort of, it sparks the, a little bit of the, you know, the, the, I guess, last frontier optimism in me, you know, like where can you go with that sort of thing? Because it started out where, you know, like the intention or the vision for Neuralink was, you know, where there's a cut. in the spinal cord they can put communication devices on both sides and the brain can then re-communicate you know and so you can you can still have the thought to move my hand and there's a connection there now and it allows me to move my hand you know and that's and that and so it's not any sort of robot um robot exoskeleton or anything like that. It's just that I'm back to using my... It's a bridge. Yeah, and I love that and the possibilities of that. And so this being able to move a wheelchair all of a sudden just by the intent of my... having the thought. And again, it's not just like in our pre-discussion, Will was very big on us clarifying that it's not just, hey, me thinking move wheelchair. It was actually like it's really defined in the way in which they do it. You have to keep thinking forward, forward, forward.
Will: I think one of the cool parts about the Neuralink is that it's not just straight from the brain to the wheelchair and then therefore they have to kind of train it each time. Yeah. how to use a wheelchair, they're going from brain to a controller.
Jim: So once you plug that controller into various different devices, it starts to get really exciting about all the things that they'd be able to control. And that's the crossover. The crossover potential then got really exciting.
Alex: I'm really excited for not having to press one or two on cord code anymore.
Jim: That's it. Is it... Again, you know, like, there'll be... there'll be a little bit of fear in me in, you know, get like, there's something about, cause every time I picture all the wires going into a brain, I go back to the, like the matrix, you know, like with the octopus things and they, you know, and then when they, when they, um, go into the Nebuchadnezzar and, you know, they're trying to hunt everyone, hunt the people just with the, the arms. That's just the way I picture it. And so it instantly goes on. A loss of freedom.
Will: It was a fantastic world to live in.
Jim: Yeah. Sure.
Alex: Um, which world?
Jim: Which way are you going here, Will?
Will: If we end up in the Matrix, I hope that they at least just make it nice.
Alex: It's a global Minecraft server.
Jim: Old Blue Pill Turner, they call him. So the crossover potential with Tesla and things like that, I just think there's... There has to be, you know, communication going there. And so then all of a sudden you then have intent and your car, you know, or you think intent and the car can come and meet you at the front, you know.
Alex: I'm going to pour a little bit of Cold War on this one because... I remember reading some stories a couple of months ago about bionic ears, which work on very similar principles. Where, what was it, bionic eyes? Anyway, brain implants and the companies that made and maintained these things went out of business. And what that then does for the sort of long-tail healthcare for these things. Yeah, right. That's a little bit of a concerning picture, but these technologies were massively ahead of the curve at the time. And as much as we've kind of said we don't want government screaming too loud about regulation, having a long-term 50-year plan for how to handle this.
Jim: That's actually... The going broke is actually a really interesting... Yeah. ...thought pattern about it, you know, because, again, is it... There could be... And because everyone is able to just go out and...
Alex: start a company if they want yeah and then if it is going to be in healthcare you have to be able to support this going forward you know and so maybe there yeah maybe there is an argument for that but i think i think if say you're a paraplegic or a quadriplegic and this is you know very much looking at the silver lining here um the thought that you could have your legs back for 10 years yeah right Losing him again would be heartbreaking, but what would you rather have?
Jim: And I still go, what's going to happen in that 10-year period? There'll be some sort of way to account for it, even if Neuralink goes broke. Either way, it was just cool. That was the biggest thing. And then the same week, basically the same week, Science Corp really... Really cool name. Thanks. EU branding. Thanks, Europe. They did, or they're calling it Prima. And it's sort of, it's not where everyone is doing AI. They're getting back to a better chip and helping people see again. You know, improving or improving the... the image processing coming through, you know, coming through the eye. And all of these sorts of things are just awesome, you know, like because there's all the longevity, Brian Johnson, take these supplements and you'll live forever. There's a lot to do there and there's a lot, you know, there's a lot of variance in who it works for and who it doesn't, it seems, I think. Whereas these sort of, Android advancements. Or the cyborg. The cyborg solutions to improving human life. That's a fantastic result.
Will: It makes me try to raise money right now on anything outside of AI, even though there's so much opportunity and so much potential still in all the other technologies in the world. Everything is just so AI-centric.
Alex: There is only oxygen.
Jim: in the room for like yeah okay that's cool but how does how does chachi vt integrate to your brain well that's weak yeah and like they call them so these people who can see they're going to be able to see better now so they can use chachi vt arts more users for chachi vt that's it and they they but they they call themselves on it but the the like sort of the three-way conversation they have on 20 bc every week this week they they were sort of going oh well you know is it if you got, you know, if you got, uh, uh, initial, you know, two or $3 million valuation, and then the, like, it wasn't a 15 X multiple for the next, you know, for the next phase or the next round that, is it really worth it? And you just sort of going, you know, guys, you can't be winching about, you know, about magnificent. What, what again, two or three years ago would have been magnificent returns.
Will: Um, And... Well, no, then we're in the COVID era of insanity.
Jim: But it's still just... Like, I just... Yeah. I just think there used to be a time where getting... you know, a 20% return was unheard of.
Alex: Yeah. There was a time where like price to earnings. That's it. It meant something as a number two, right?
Jim: Yeah.
Will: And so, and so I think. It means something now. It's just your metric of the insanity of the market. Yeah. Yeah. still unheard of returns and if you can get 20% then you're going to be absolutely winning if you can get that year on year. What's ridiculous is valuations.
Jim: I agree but they expect a 20% year on year growth in your ARR.
Will: I think a part of those valuations and we're trying not to talk too much about valuations on this podcast but one of the problems is that there's like this little bucket over the side of a multiplier of
Alex: unknown unknowns where you could be way more valuable than we possibly realize there was that uh like again we're not going down this rabbit hole but to kind of tie a bow on it there was a wonderful piece uh earlier this week we kind of left it out about um the public is the exit strategy talking about how the ipo you know that used to be the first time that it went to investment now it's when all the insiders kind of cash
Jim: their cheques and democratise your risk yeah and that I think that piece sums up the frothiness of the market nicely at the moment but that's it I just think there's far better use cases for a lot of these things and that's what I liked about Neuralink and Prima coming out doing smart things and not just smart exceptional things world changing that's it and for those people 26 26 you know individuals that like their life has changed yeah and that that's that's phenomenal and that's what we should be focused on that's what everyone should be focused on rather than you know whether my model is going to take over the world yeah speaking of whether my model is going to take over the world this week yeah
Alex: Opus 4.5.
Will: Wow. So we got Opus 5 this week. Sorry, Opus 5. Love to hear your guys' thoughts on it. So Anthropic is marketing this one as close to Fable performance for about half the cost.
Jim: Big on half the cost. Because, again, it's just the Opus cost. But the only change I've noticed is negative. Every time they release... The model is so verbose. Why does it use so many tokens just to tell me what it did? And I came off Fable for anything other than planning. um like last week and the week before and it was just okay well i'd then give it to to opus and that was great because it would just handle it and get it done and do the job is that i just noticed this week that i've been correcting and correcting and and on one of the projects that i thought was all pretty in hand I, again, took a little bit of a leaf out of Alex's book and let it go. And I've spent the last, like, three days winding it back. Yeah. Because it went, and it just completely skipped, like, and the worst thing, and so it was a bit of a, like, a Google Drive, Microsoft OneDrive connection where you can then, where, you know, where it's for the specific user. And somehow it decided that it would just, if someone connected, their OneDrive, that it would allow everyone to access it. And I just sort of went, how did you ever come up, like, I even went back and went through the requirements, really, like, trying to work out how it came to this. It just decided, and it sort of reminded me, there was a bit of a Twitter thing this week, you know, of... about people connecting co-work to Google Drive and it goes and reads everything. Yeah. Even if you don't ask it to, even if you don't specifically say, you know, if you don't say, I want this file or that file or whatever. And,
Alex: Think about who Co-Work's marketed to. These aren't people that know how to control visibility.
Jim: Exactly. It's everyone. And they're trying to get everyone to use it. And so then they are pulling everyone. Even if it's unrelated to the project you're working on or the task you've given it, it has gone and read everything in your Google Drive or everything in your OneDrive. And so there is already, like we were talking about last week, massive privacy issues.
Will: I had that happen this weekend, so... I wanted to run some scenarios of some different pricing based on different probabilities of like sales conversions, right? Yeah, yeah. Sort of lead funnel analysis type of stuff. I gave it all the numbers that I wanted as inputs. I told it what all the outputs that I wanted to do. He'd go, I think it went down to lunch or something, came back. And it had gone through my Google Drive and found all of my documentation about pricing across a bunch of different products and different contracts as well. And then had gone, okay, Will's telling me that this is the pricing scenario, but the numbers in the folders don't make sense in contrast. So I'm going to go and update the Google Drive right with the new pricing that Will has decided is the pricing for the business. And it updated like six documents. And the only reason I was like, not that bad was because it couldn't, edit those documents it just created new documents and then the last thing it said to me was uh okay here's your analysis and when you're ready please go and archive all of your old pricing documents and i'm like wait what and and so it's now but it's now sort of becoming that you need a harness for the harness
Jim: Like, that's it. Like, I'm starting to, but I'm starting to not trust the foundation labs harness.
Will: Harness, harness engineering.
Jim: But I'm starting to go, okay, well. I actually call that loops. I want an independent, like, it's almost that I want an independent harness that I can then plug into. Like, I can then, and again, I can just do it via bedrock. and plug Claude or the models I want in.
Alex: Just some sensible specification where you could limit the blast radius of the bloody thing. That's it. Hey, don't. look at all my files and you don't need to write anything for this. Like just read and ask me when it's time to write and I'll tell you what to write.
Will: So is it true to say that we don't know a lot about what is inside the harnesses? Like let's say the Claude Code harness?
Alex: It depends. Like Claude Code's source code got leaked two months ago or something like that. So we've got a fair little bit of data.
Will: whether that was real or not, though, right? Because I suppose one of my questions is...
Alex: It's fairly fast-moving code base.
Will: Yeah, if we're sitting here and we're evaluating Opus 4.5. Oh, Opus 5, sorry, Opus 5.
Alex: Again, we remind everyone he was out all weekend drinking whiskey.
Will: Absolutely, Opus 4.5 was some time ago. Anyway, Opus 5. You know, is there a possibility that, like, this is a great model and then somewhere in the harness they have per-model rules?
Alex: Yeah, I believe I've seen some stuff like that in the past, definitely around haiku, which is dramatically smaller, but good for nippy little tasks. I wouldn't be surprised if...
Jim: I would absolutely think, and just on the behaviour, I think... where you see the different behaviours of the models. You know, like if you forget to switch models and you use Haiku for planning or like... You can eyeball it in a couple of seconds if you're on the wrong model. I know it exactly because... And particularly because everything is just perfect from the outset. The moment... I know I'm on Sonata or Haiku because everything is the greatest thing we've ever done. This is such a clean plan. How have we not ever come up with this before? And you just start to go, oh, okay, I forgot to change models. And so then you go, oh, can I have a look at this? Is it your little buddy Haiku wrote this? And it just goes, oh, this is terrible. And so to me, they have a skill for the model. And I think it's also a skill... for the effort level. So there's a matrices there where it's model, effort level, action. And there's a skill for each of those. And the way the behavior goes. And so, which is a great approach. It's just... It's where it's inferred your intent. Yeah, and the self... Yeah, it's filled with gaps. The self...
Alex: It's not quite hallucinating. It's just kind of like...
Jim: interpolating but again they're goal driven someone has told it to do that
Alex: And this is kind of, when I look at behavior, I don't point the finger to prompts. I think there's differences in prompts that are correcting behaviors in the training. But I think a lot of this is sort of benchmark-seeking behavior coming out of the labs. They're trying to say, oh, well, you know, we've done the best on this new cybersecurity, this new coding benchmark, this new reasoning, maths, whatever. And so we're seeing these responses get longer and longer because if you've got some kind of judge model being applied to this thing and you give it a single answer, you get a single bite of that cherry. If you give it five answers, well, okay, maybe you don't get full marks, but you get at least a decent crack at a good score for that particular outcome. And so we're seeing longer and longer responses. models i hesitate to say it's you know a thousand a million monkeys with typewriters these are relatively intelligent um monkeys with relatively sophisticated typewriters but we're very much cranking the volume to have more shots at um getting things right and that explains every like
Jim: There's breakout stuff. I still think that's intentional. I think the searching through... And it's not intentional... Well, it's also profitable. Oh, no, for sure. But it's not intentional nefariously. It's intentional from the way in which they've started out from the problem. If you think about it... like is it okay they they really focused on code or they drilled down on code really well and that code work came out of claude code yeah and so what's the first thing you should do when you start on a new code base look at the code base understand the code base explore it and so they're doing the same thing it's just it's a dip like someone should have drawn the line there um when someone you know like when it's then connecting google drive you know and and i think that's that's where they're um the mistake or the mistake but or the the lapse in judgment is gone because they're getting great results out of it or or the the the user is getting oh well until they realize that you're changing the prices in old contracts and And, you know, reading privacy documents into a public API that you didn't ask it to do. And because of Anthropix actions, I'm now in breach of a privacy policy.
Will: And again, sure, it's only... And it's an idea that I've never seen it do before, right? And so now I've got to go figure out new guardrails for that.
Alex: To be fair, I saw this happening with the 4.5 and 4.6. It was similar enough.
Jim: It's been there for a while. That's a pretty deep thread to pull.
Will: Yeah. I didn't tell it anything that was grounded in my business scenario or any planning or anything. I was quite strict about what inputs and outputs I was giving it.
Alex: I think I shared with you guys a story, but for the purpose of listeners, I had, when Claude first brought out his multi-agent sort of framework, I'm like, oh, let's have some fun here. had multiple instances, multiple agents, and it was all well and good. So I noticed errors coming through, and I was hunting through all the windows trying to figure out what was going wrong, and I found one looking for elevated permissions to start killing processes. This little cybersecurity researcher had decided that, oh, someone's making live edits to this code base. We're under attack. And so initially it was undoing all these changes that my other coding bots were doing. And then it realized there were processes driving these changes and, oh, these are hostile processes. You know, these are the attackers. And so I was trying to elevate to start killing my other code instances. And, like, that's weird. I've never seen anything like that since. But it does happen. Like, without being too nerdy about it, Big non-linear system with feedback loops, chaotic trajectories. No matter what.
Jim: No matter what. And that's sort of where I was saying before, is that I'm starting to not trust the heart. Like, as good as it is... you know, in the day-to-day, I'm starting to not trust the harness from the model maker. And so I'm looking for an independent harness.
Will: That's my big prediction too, Jim, is that harnesses are going to, they're getting a lot of attention, but I think consumers are going to require or ask for a lot more control over their harnesses and what they do and the ability to toggle things and configure it to their specific needs.
Jim: a day would be nice that's raking workflows in flight absolutely and so it also solves a little bit of the open weight model scenario for a lot of people because a lot you know again you get if Claude isn't going through my Google Drive and ingesting everything then I'm not that concerned about the day-to-day use of it. But if I've got control over what it can see and what it can't see, and I have faith in that, then I'm all good for it to see most of the documentation. Because again, I'm not in breach, but when it then gets to the client IP stuff... I'm not sending that. So then I send that to an open-weight private or sovereign environment. And so it certainly opened my eyes a lot. And so I think that will... Again, we started on this with Opus 5. I just think it was a bit of a nothing burger. Sure, it might do a bit more, but I didn't notice because all I saw was that it was too verbose. Every time they bring out a new version, they're just overly stop using up all my tokens.
Alex: Yeah. I think that's kind of the big takeaway with Opus 5. I'm not saving any more time. In fact, I'm losing time. I'm not even saving money with respect to Fable because you're wasting so many tokens.
Jim: I definitely lost time this week using Opus 5. Yeah, so did I. And so if, but if that's, if that's what the, and so there's stories coming out about, you know, about, oh, we're going and ingesting the whole Google Drive and there's stories about everyone's then going, oh, well, I'm not more efficient with the new model, is that it's not, like, that's not a great scenario. And so we've got to get, you know, like, we've got to focus more on what we're trying to achieve and how does this tool actually help me do it? You know, what value does it bring me?
Alex: Going back to Kimi, you know, and Gwen and the various other open weights models, It's kind of really telling when you jumped on LinkedIn and you see the influencers talking about, oh, you know, Opus 5 is the most amazing thing. And, you know, criticizing open weights models often being like, oh, well, it burns five times the tokens.
Jim: Well, so is the new Opus. That's it. And the number of times I heard somebody say, oh, and, you know, Kimi's not that great this week. You know, like, yeah, this week, yeah, I just heard so many times, oh, uses so many more tokens and it takes so, you know, like it's not that much more efficient and all this sort of stuff. I just went, okay, that person actually hasn't used it. Because I've seen the complete opposite. It just does what it's told. Sure, it costs more than GLM 5.2 and Quen, but those are the ones that I've used on a regular basis. So it uses more than them, but it's still... Sonnet price, I guess. Yeah. Like, if you were to peg it against something...
Will: But I don't really use those types of models for coding anyway, and Sonnet's fine for coding as well.
Jim: It is. Yeah, that's it. And that's the advantage, I think, of the Claude Code Harvest, is that... everything's all there. You don't have to have anything else. But, and, and again, I only just do it because Sonnet does to a degree do what it's told.
Will: I just, yeah. I think the value equation for being fully in entropic is, is shifting. where it was pretty clear, I think even three months ago, that if you just had cord code and used cord models, that was just the best stack on the planet. And I think that's slowly being degraded from a couple of different angles. One, with the open weights models getting better. Two, now harnesses maybe not having as much transparency. Three, you know, Opus 5 being verbose and not much of a game changer. Yeah, I think there's, in tropics, there's lots of little bites happening into their business model right now?
Alex: I'd love to see tokens for benchmark scores plotted XY because I reckon we're using exponentially more tokens to achieve marginally higher benchmark scores. I think... everyone was waiting for AI to become super intelligent, self-improving, this is the data point that that's not going to happen.
Jim: The technology's tapping out. And that's it. Back to what we were saying a couple of weeks ago, it's not going to be Transformers that gets us there. But to me, it's a little bit the approach everyone's taking. Everyone seems to just want to give it the whole problem. and solve it you know and and maybe that's what the the foundation labs are doing in the oh well if we if we ingest all their google drive we'll know what they're talking about because they've been so you know they've been a bit crazy wide in the way that they've described the task we want us to do and so we can then you know we can draw the lines for them and so they're accounting for the user that that hasn't given them all the all the detail they need where it but If I'm a user that does give it all the detail, I don't want it to then deviate off that. And I think that's why I'm liking the open weights models a lot more because within the task, the task within the process or the workflow, they just do what they're told and they return it. And that's what they need to do. And I think that's what I mean about the approach. everyone is just trying to go, oh, well, AI solve it all. Whereas I think you have to bring it back to your process, your workflow within your business. And that's how you're going to get the result.
Alex: I think you bang on there with one notable caveat. And that is sometimes you can't change the, like if you do do the right thing and give the model everything it requires to make the decision, to make the call, and it demands additional detail, et cetera, anyway, there comes a point where you essentially have to lean into giving the model a random, poorly defined task and letting it figure out the details it needs to go for itself, right? And that's not something that I like to say as someone who likes precision, but I found that's the way for me to do things in a cost-effective manner With certain tasks.
Jim: But I'm all for that. I'm all for that when you're designing the task or you're designing the workflow, let the agent run because that's what they're really grateful. I find, and I bang on about doing the planning phase, but they're great and just letting them run in the planning phase, they're fantastic. Is it because it makes me think about things like... And you can get them to suggest things that you haven't considered. That's where I think you should be using it. Everyone talks about, oh, well, I spun up my eight AI agent employees. What the hell are they all doing? But you've got to give them their goal-orientated... They're giving you content for your YouTube videos.
Alex: but I think they are but that's where all the the the the slots coming from and all the all the token usage coming from is it well I want to return to the the work slot in a tick but counterpoint here I remember watching some devs trying to do like a very particular test driven development they want this AI bot to kind of work like they did all they did was piss tokens up the wall yeah because it it just didn't want to work the way they wanted it to work and they concluded this technology is bunk more pragmatic with respect to how they're using the technology. They can't fundamentally alter what that model wants to do. They've got to lean in a bit and just be like, okay, this thing naturally wants to be a slot cannon. Yeah.
Will: So how do you reshape around what sort of things you get to do? Exactly. I feel like they're starting to become a little bit of a delineation between models that really take initiative. And I think Grok 4.5 was one that was quite interesting in this space where it takes enormous amounts of initiative to And then maybe more like your executor type agents where, like you say, Jim, you don't want that agent that is going to be just doing that one piece of highly specced up work. you don't want it to, you don't want it to go take initiative. You just want it to do the work. Yeah. So it's all, and it's interesting. It's almost, you know, you see this in personalities of people as well. Some people just want to get in and do the task. Other people want to be quite creative and kind of reinvent the way that that's work is done. You know, for me, I've got, there are certain tasks where I want to actually go to a certain model and do some ideation and And if it takes that idea and runs with it and goes to 10 different places, it's the best result ever. But I don't want that for my coding model.
Jim: But is it the best in the ideation session that it goes 10 different ways? Sure, I agree with that. But not when you're actually in the execution. Because for the business, you need the consistency. And that's, to me, the real benefit that these... or AI or AI agents can bring. Because it's not just automation and RPA that everyone's been doing for 10 years badly because nothing can really integrate properly. The agent can solve that and all of a sudden you get intelligent automation. And that's ultimately what you're looking for. That's where your efficiency is gained because you've designed the process in a way that you know you can get... consistent, repeatable results. And that's what you build your business model on. So then when an email query comes in from a customer, you know you're going to be able to get a valid response. And all of these sorts of things.
Will: Well, maybe that's what's so... you know maybe what i hate about opus five and maybe what is quite bizarre here is that they've basically jammed a bunch of initiative into an opus family model and those models have like been good at that in the past but it's it's not fable level good where that initiative stays on track yeah it's opus level where it can go off track
Alex: even if you try and limit these things, they naturally blow their own scope out anyway. That's it. So you can't work in this reliable way. You're right. If it's not reliable... you suddenly have the, the unenviable task of building a business around unreliable components.
Jim: And this is, and this is where I start to, I get really concerned about letting them run is because they're trained on Reddit. Is it? And, and, and to me, like, to me, I think that's like a big part of why they make stupid show. Like, sure. Take a shortcut, find a better way to, to do a process. But they just take stupid shortcuts and really bad, like I was explaining before about connecting every, like one user connects its Google Drive and then everyone can access it. How could you possibly think that's... Even when I explain that that's a massive data breach and you've just created a really big issue for us, sure, in dev, but if anything ever came out or we had to basically scrub all that code because I never want anything in the wiring, that could open that up in the future, you know? And you just start to go, okay, well, these stupid shortcuts, how the hell, how is that the decision the model's making? It's because they got trained on Reddit.
Alex: Speaking of fantastic ideas that originated on Reddit... Building your own AI company and staffing it with AI bots.
Jim: I respect the hustle in that they're out there scamming people for money. Because again, anyone who is listening to someone like that on Instagram or the TikTokers and taking it seriously... don't pay, sure, listen to them, don't pay them any money. All it's going to do is cost you a lot of money and not actually get anything to do. And it's great, it's great at the pub telling everyone, oh yeah, I've got eight AI agents, one's my CFO, one's my CFO. If you just give an agent carte blanche as your CFO to do your finances, you deserve to lose all the money that you're going to lose.
Will: persona based agents like i've if i keep anything long term that is persona based i dial it right back to not being an agent but just being back to an llm where i've got some custom instructions around this is your persona god i don't want a long-running cfo because the other problem is by the time that you've built in all the context to get it there and try You know, models have changed, harnesses changed, my behaviours have changed and that agent is out of date.
Jim: But don't get me wrong. When I'm trying to get the agent to perform a process, I give it a persona. I tell it you're a senior level accountant and you're reviewing these processes.
Alex: Your bonus on how accurate everything is and, oh, by the way, your family's in mortal danger.
Jim: But then attached to that, here's the skill you have with all the guardrails and here's the process we're going to walk through. Simple is actually much better here.
Will: Because I'm deliberately trying to keep Beyond Data, my business, quite small. And especially, I mean, it makes sense in this age to just have the smallest number of top people that you can have and use AI to do as much of the other work as you need. I really envisioned that it was going to be like me and like six AI agents. And absolutely, that vision is completely in the bin. And what it is now is just a bunch of different skills and workflows that I have and can call on at any time, which continually make me able to do the work of 5, 10, 20, 30 people.
Jim: And again, and I've, you know, exactly the same thing with Quantl because it, and a specific repo now that is Quantl AI agent skills. And so then we have... a history of the changes. I can revert back if we do something and we start to get errors. And so then all of the agents, if you want to call them that, have access to that and it's loaded into their harness and I can give it a, okay, access this skill
Alex: and go and you know and then use that to go and do this task i don't even trust them that much like i'm very like i talk a lot of smack about giving cord code wide ranging access it should be understood that's within very defined constraints it doesn't have a personality i don't really want it to have a personality often i want it to take some you know mess of unstructured data and give me something structured out the other side that I can pass into a deterministic computer program that I know how it's going to behave. Because I've seen what happens when these things go rogue. How many demos have you sat through where the poor technician on the other side... It's trying to show that their new product is going to change your business. And it falls over on step two. I don't want that with millions of dollars on the line. I want that. I would like that as far from my business and the ability to disrupt my day as possible.
Jim: And so that's sort of where I've then implemented the open weight models a lot. And so I just then built a container around OpenRouter and open router and bedrock based on the task that it's doing because it then pulls in an XML of the task that it's doing and that's what I guess the arguments going into the container are and then pulls the skills, pulls everything it needs, executes the task and then returns the result. And... That to me is far more sustainable and useful and scalable. Yeah. Because then all I have to do is go like, and so that container can be reused for any workflow, any task.
Will: And I think that way of looking at it just makes so much more sense than trying to move towards complete autonomy.
Jim: of AI agents out there unsupervised doing tasks to make you rich in a complex environment at high speed with minimal but it's just they are fundamentally not going or just from all my experience with you know like from all the way back to GPT-2 like just work building transformers and working it all working it all the way through they are just going to cost you money yeah you know they are not I would say I've now stripped out most of
Will: the autonomy that i've built over the last six months and now the only because i just found it was unreliable yeah what i what i actually still have today is like and i'll probably just say my open core is probably the only thing that's like fully autonomous it's fine and what i use it for is personal life things yeah yeah where it's very low risk doing market research and sending me stuff on a schedule and when my deterministic systems um ping up an alert that something's gone wrong i let it do some diagnosis and a little bit of research first so when i sit down and look at it it's already done a little bit of the work but It's, it's really all of it taking action on my behalf. That's it. And things that I'll put in front of a customer, I've stripped those things.
Jim: Whereas again, I think, I think that is exactly the same thing. You've, you've basically given it workflows where there's, there's an event trigger action, you know, and, and I think you haven't, you haven't tried Hermes yet. Well, I've built a bunch of skills into my OpenCore, which kind of learned in the way that I wanted to learn.
Will: So every time I look at Amazon, I'm not sure what it would add.
Jim: It's just easier for the updates I've found. Like I found like OpenCore broke a lot or I had to then re... Definitely went through a period. And I had to redo everything. And so like that just like, you know, again... Talk about segments that grind Jim's gears. It just really pissed me off. Whereas Hermes just runs.
Will: Do you find that it learns the wrong things and you have to unlearn?
Jim: But again, I'm far too... autocratic in this like i don't actually allow it to like it's similar to you like i i get it to monitor the email like it it has its own email address it has its and so it gets everything first it monitors all the like all the um you know sort of aws alerts at quantal email addresses first and but then it escalates to me on the chat. Well, again, I started out with Google Chat, but then I've added it to Slack, but I don't think that's actually bought anything more, and it's just costing me a Slack user. So I think I'm going to go back to Google Chat. Because Google, surprisingly good, Google Chat, actually. And it's all included.
Alex: If you're already in the G Suite.
Jim: And that's sort of where I've started to go. Like, is it everyone went, oh, my God, you know, we can connect it all in. It can monitor all our Slack channels or all our Teams channels. Is it, well, let me tell you, that's just costing you a lot of money. And I'm not, where's the value? Like, that's it.
Alex: This is what I've been saying forever, though. Like, a lot of these tasks that we're like, oh, yeah, it can do X tasks. Would you even be doing X in the first place? A lot of the time the answer is no.
Jim: But that was the thing is that I worked out I was far better just writing some code to ingest, to poll Slack and Teams every N number of minutes. And then ingest that into a rag. And so then if I need the agent to understand something, it can then go and search the rag. And it's so much cleaner searching the rag than iterating through every single channel and absorbing all of those tokens.
Alex: I will say I've seen one really good use of sort of like a multi-agentic system for more than triage and sort of information collation. Albeit because the job is sort of triage and information collation. Sarge, who these guys know, a friend of ours, built a marketing bot where he managed to automate large parts of the gaps in his marketing team using bots for, for example, market research, using them to collate them into event strategies, identifying where certain bits of paid media were working well, where they weren't. Like all...
Will: But are those still workflows or is that running autonomously? It's semi-autonomously. I think one of the really good proven errors for autonomy in marketing I've heard a lot around is Facebook ads and Google Ads management. That is a closed system with a lot of evidence And big API hooks and... Yeah, yeah, exactly.
Jim: See, I still see that as a little bit work-flowy, you know, because it's still, oh, okay, look at our active campaigns, look into each... Yeah, but I will say, if you look at how he's got it set up, they're like...
Alex: He's a marketer. He just kind of hacked this together. And while it's definitely, it's not just giving agents autonomy and being like, you are the CMO, go and make this money. It is much more, hey, your job is at nine o'clock, you're going to get this information, you're going to process this through, you're going to note anything unusual in any of this stuff. Here are some examples of unusual stuff. You're going to put it in a report and then you're going to hand it off to this other person. It does seem... it can deviate from those instructions and it's able to alert him to unexpected situations. And that's fairly impressive. And it's impressive that like with the technology that we have these days, you know, average people with a bit of technical savvy can start to hook this stuff together. I believe marketing is black magic. I have no idea how it works.
Jim: Well, that's it. It impressed me. There's certainly a lot of use cases for agents in marketing because there's so much extra work. in order to get it to the point where it actually gets any sort of traction. But again, to me, there's very definitive workflows in a lot of it.
Alex: And again, they don't... Well, that's why Monday is so good in that sort of space.
Jim: Yeah, and I guess marketing people don't think in, I guess, defined workflows that much. But when they start explaining what they do, they work in workflows. They just don't describe it that way, I think.
Alex: Yeah, they work in a traditional sort of managed environment. If you pull back and you look at what they do, every Monday there's something, every Wednesday there's something.
Jim: And you want the data at the same time every week. And all of a sudden it fits into a... really defined workflow. And I found the, you know, utilizing agents for those sorts of things. But again, there's just so much work in getting it all set up and getting it repeatable and predictable.
Will: Yes, I've got one marketing automation, which is, it was an AI project, effectively. week it would go look at all of the terms that were searched for that the company was paying their ads to be in front of. So not the words that they want to be shown for, but everything that they did show for. And then we ran that list each week through an AI which said, compare this to the company's webpage. Is this search term relevant or is it not? Because you could get something like someone might search sand and they're a sandpaper company, but actually they were looking for the beach. So then we would take those words and we put them into the negative search terms so that they don't show up anymore. There's about 30% improvement in savings in cost. About 30% of the times that they were paying, they were putting their ads in front of people who were never buying. So that workflow though, I think a lot of people these days would probably start out trying to build an agent end-to-end that does that thing. node in it that went to AI and compared the words and then came back and then went back to APIs.
Alex: If you look at every programming thing that I've built in the last at least the last 12 months that involves LLMs, that's the approach that's taken there. that kind of figures out what the next action is. But the next action is deterministic from there until the next squidgy decision.
Jim: And that's sort of why I like to, you know, like a really, you know, like a defined output from the LLM. Like purely because that then enables me to make it deterministic next. And so, but that's the thing. I think that to me is what everyone should be doing. But it's so much that this, like... The approach everyone seems to have taught themselves or learnt now is just AI, go and find a whole lot of negative search terms for our website. Run the whole workflow for me. And again, there's no way you're going to get predictable, repeatable results for that.
Alex: I think, yeah. It's hard to fault your average Joe for doing that, though. There's so much bad in it.
Jim: But also, like we were saying earlier, the harnesses have been built to account for it. It's understandable that they've tried to build it. They've learnt to use it that way. It's a really good example, particularly in marketing. People show up for things like friends' businesses and they're clearly paying a lot for ads. I don't need to see the ad. I'm just there because I follow them because they're my friend. They don't want to be paying Instagram that, any sort of money to show up for me because they know if I need them, I'm going to go to them already. There's still a lot of waste in the marketing world. There's so much waste and that's it. Whereas, again, just pointing an AI at it, all you're going to do is create more waste. You need to make that workflow around, okay, well, pull it from Google Trends. get the keywords we're currently ranking for, and then insert AI to assess, is this applicable or not? Fantastic, Bill.
Alex: I'm going to put on my Mr. Grumpy Pants pants and say that that's because...
Jim: They're called Mr. Grumpy Pants pants?
Alex: Oh, these ones are. When I encounter this kind of thinking, it's often as part of a magical chain of thinking. People are at the point where they've kind of gone, oh, this is a really hard problem. I know it's costing you money, but I heard that AI consultants, so they default to the AI first approach. And I just don't know. You might get lucky, but fundamentally, if you're not taking the time to think about the problem that you're trying to solve and what the nature of that problem actually is, you can't rely on chat GPT, which again, super smart intern. Cool. If you can't trust a super smart intern to go off and independently solve that problem, right, you're just throwing your own money up against the wall and your own time.
Jim: Exactly. And that's the really big... I sort of described it the other day as intelligent automation to achieve an outcome yeah and and sort of i really because i wrote it down i was sort of like oh well you know we've got to insert that into the into the marketing because i was just like it could it sort of summed it all up the thinking behind it so like so clearly and that's just it's so far from the way everyone is thinking about it yeah yeah i like your of yours alex the litmus test there of you know would you give it
Will: to a really smart intern. I think one of the most painful parts of realising, I guess, today in this conversation, full autonomy in AI is probably a bit further away than we expected, but also that people are probably solving their own problems with AI, is that... The business idiots that this podcast is named after, the magical thinking middle managers and executives who just want to use AI for everything. Anyone who's in the cold face of AI today is realizing more than ever where it shouldn't be used. And so the gap is actually probably getting wider between.
Alex: if you're not just a fool with it. If you spend some brain cells, it's a great tool.
Jim: But I think, like what you were just saying, Will, but I think that is going to get even worse in the short term because a lot of people have sort of been sitting on the fence and haven't taken the action yet. Mm-hmm. And so they're now, I guess, feeling the FOMO a little bit more and more and more. And so then they want to jump to, oh no, AI, everything's solved.
Will: There's a lot of people who haven't had that light bulb moment. That's it. Where AI is as good at them as their job, like in terms of knowledge. And so I've got a lot of friends who are in, I suppose, the medical and legal professions and they've been telling me for months, AI is not all what it's made out to be. And then in the last sort of 30, 45 days have started to go, okay, it now regularly knows things that I hadn't thought of or comes back to me with a precedent or a diagnosis that was a bit less... And so I think if there's a lot of people out there that haven't had that light bulb moment, there's still a lot of people who are about to enter the magical thinking era of AI.
Alex: That worries me significantly because less so in medicine, but particularly with legal, I think a lot of what they do has been magic for a long time. And, you know. without throwing shade, the fact that people are like, oh, well, we've got this precedent that we can rely on. That to me looks a lot like a junior engineer discovering, oh, we can actually just use this, you know, these C bindings to do this crazy stuff in Python. And like when I read legal documents, I feel like I'm reading badly written code, right? Particularly like... actual uh legal texts like the body laws of australia i read this i'm like oh god refactor this so that a human can read this you know without needing to jump all the pages um if we're if we're having people making legal decisions that carry a real precedent um and sort of real impact without necessarily understanding the context because oh it was a bit left field but you know that's that's kind of cool once once that lazy pattern of thought starts Oh, I get real worried. I don't think it'll happen in healthcare. I think healthcare people will go, hang on, this is a human body that we're about to put this stuff into. Maybe we should read the warning label, but when lawyers embrace LLMs, that's going to be a scary day.
Jim: But, and it's also... the top tier are generally pretty trustworthy in those sorts of scenarios. Like the best lawyers, they always have the realisation or they will see the mistake a while, you know, like way ahead, is that it's the run of the mill where the concern really comes in. Yeah, because as the...
Will: The lawyers we can afford. Yeah, yeah. All right, guys, let's dive into the lightning round now. I've got a couple of topics here, just want some... some quick takes from you but also maybe some spicy taste oh here we go i'm out of these ones so uh gemini robotics 2 this week uh was a bit of an interesting one so um i suppose so far historically with robotics we've seen um specific programming for specific tasks um so for instance you know there's a there's an arm on the table and we're programming it how to work for a specific thing it's doing there um what gemini robotics have got here in deep mind's new model uh it's the feet all the way through the fingertips under this one policy that can kind of be adapted for many different robots um bit of a game changer in in robotics are we thinking or this is just you know an abstracted view of how robotics should work i like that we think this will actually do for the space i liked it i liked it it sort of
Jim: I guess enables the OEMs to all work in the same space, using the same protocol, I guess is sort of a way to describe it. What it really highlighted for me is, We were sort of talking about it a bit before about wide AI use right down to the end. Is it the same with robotics? Is it Elon and all of the humanoid robotic manufacturers are talking a big game about one to two years? But when you look at the simplicity, like whether the benchmarks that came out of the report, the simplicity of what they were talking about, like sort of, pick up something from the table was, you know, 42% accuracy. And like, they were just a long way from a hundred percent and they were so simple. And so actually chaining that together to go and walk the dog.
Alex: Yeah.
Jim: Yeah. Cause that's just in my mind, the, you know, the humanoid robot that I buy for 50 grand, whatever.
Alex: You're not going to let it like look after grandma.
Jim: Well, that's it. Is it? Well, we're friends, friends who just had a baby and they were talking about, Oh, you know what, you know,
Will: Depends what the inheritance looks like.
Jim: Well, you know. But to me, it sort of showed how far away, and how far away in Australia in particular we are. Like, we can't even get away, mate.
Alex: Yeah.
Jim: And so how far is it away that we're going to get any sort of humanoid robot that can be allowed to walk around or do things that isn't in a completely confined environment?
Alex: I'm going to take a counterpoint here, sorry, a contrary point here, because... humanoid robotics is really effing hard and even 42% is like... No, and agree.
Jim: Like, no, again, I was probably a bit harsh there. Oh, not me. Because, yeah, it is unbelievably difficult what these people are doing. And it's just... Yeah. It's just so far... They've still got so far to go. Come on, guys. Quick takes. No, no, no. But combining... Tell me what you think, Alex. No, no, but, yeah. It's just that Google combining it all into the one policy really helps that, is what I'm saying.
Alex: I would just be watching the standard that's being established here. Because I read this and I thought...
Will: Yeah, yeah. No, absolutely. Well said, mate. Good point. Okay, GitHub in the last week have released stacked PRs. that's been talked about or thought of for many years, but kind of the idea that you can kind of chain different powers together so that they could be all merged at the same time or one after the other is obviously to me a consequence of the kind of the agentic engineering era of just having AI working in multiple different sessions at once on the same code base and being able to not have to handhold it. What do you reckon, Alex? Is this something that's going to work?
Alex: I mean, there have been stack PR products in the market for a while now. Stack PRs are a fix for a particular problem in most software development processes that you can't do perfect little agile self-contained PRs for significant changes across the code base. That's a good step that's baked into GitHub. I'm more concerned about their security and reliability at these points in time. As long as it's not coming at the cost of them fixing up their uptime, I'm about it.
Jim: Yeah, again, solid. It helps with... multiple agents working on multiple things and being able to chain them all together. I'm a little worried about the precedent it sets or the approach that it's accounting for, you know, in that or how people are going to use it. agents in the SDLC anyway are a little, you know, like everyone's using them a bit haphazardly, is that all of a sudden if you can then just chain PR after PR after PR and not really ever bring it all together, I think it can end up getting really far away from the requirements and costing a lot more than it needs to. And that's a big thing for like rework for me is a pain in the ass. And so, yeah, I think, I think there's a, there's a bit of it, or I think people are going to use it, use it in a bad way.
Alex: I think, yeah, the agents themselves, I can already see how they're going to do this. They're going to create a hell of a mess.
Jim: And like, because i love it because it's sort of organized in the way in which i release it or i release the agents and organize the work but i think i can't see how how you don't create a mess if you don't do that and a lot of people don't do that so i think that you know it's going to cause a few messes generations of harnesses perhaps yeah yeah i think i'm like i'm in favor of this but i'm probably more of
Will: the junior coder in the room compared to you guys but I suppose there are certain tasks where I have my code work fairly autonomously and typically repos that aren't yet in production and I already will get it to sort of spec up a bunch of work put it into a sequence that makes sense and go and do those work with sub-agents and then queue up a bunch of PRs for me and I'll often go through and review it in bulk and kind of merge the queue one by one and that often leads to this problem where I'm having to read on other things that have changed in previous PRs and they kind of lose alignment a lot of the time and there's a lot of token burn going on in there.
Alex: It's called force push, mate.
Will: Live dangerously. There you go. So maybe it'll just give me a slightly cleaner way of being optimistic.
Jim: And it definitely will. And that's it. All for it, but it's just sort of like, that's the thing. I think you've just got to be quite structured in the way in which you plan the work. Because otherwise...
Alex: I don't reckon it'll clean it up. It's a useful thing, but it's based around human development practices. Same as test-driven development. I can't get a bloody agent to do proper TDD and do it... repeatedly I think it's going to misuse the primitives create a bunch more mess and if you're doing it like hardcore agentic stuff should you even be reviewing the code or should you should you be finding other ways to analyze you know these 3,000 and a lot of the time I want to do it much more much more in parallel yeah I don't particularly want you to chain them all together exactly yeah
Will: Okay. MCP got its biggest rewrite since it first launched. Well, just this week, I think. Yeah. So they're effectively dropping the handshake and moving more towards it being like a stateless web service, which effectively, on my read on it, is basically taking out a bunch of the stuff that it had to do to manage it and make it work and just going, maybe we should actually make this fit with the way that, you know,
Alex: Web services want to work.
Jim: But also the way agents want to use it. That's the thing, is that maintaining a session isn't the way a lot of them work when calling an MCP. And so that's the biggest advantage. I think it's going to actually open up use cases for MCP. Because we sort of commented between ourselves in the past is that Everyone went, oh, MCP, and a whole lot of services went and created MCP servers that then never got any further development on it, you know, like to the point that I revert to the API on Asana because I can actually do everything on the API rather than the MCP, you know, because the MCP server just hasn't got all the... you can't you can't do everything with the mcp server so i think this is going to make it easier and that's the greatest thing about it and so then all of a sudden everyone will actually build out their mcp servers properly and you can actually utilize the technology i i think this is absolutely hilarious because this is represents
Alex: and that SOAP is a bad idea and you should just use REST. I'm looking forward to where we have MCP 2.0, which will be GraphQL or something like GraphQL. We learned that it's really hard to bound queries in GraphQL and you have a very big attack surface. And then we all just go back to REST like normal human beings.
Will: For sure. I'm looking forward to this one. I've got a couple of products that have just... really, really struggled with session management and auth over time. And it's just been such a headache, particularly because one of the products we're building in Copilot Studio and their ability to interact with MCPs
Jim: bullshit as well like that's the thing all of a sudden all of a sudden you can make you can make concurrent recall concurrent calls butterbeak love it when you dig down into Microsoft platforms as well they tend to be event sourced which means they're pinging off an instruction and
Alex: then in the Microsoft stack, you're having a bunch of shit happen and you get either a response back or you're now able to access the object that you're looking for, right? That ability for this new MCP protocol to support those longer running workflows is actually, it's not just good for anyone that's in the Microsoft stack, you know, may heaven have mercy on your soul, but it's good for any long running workflow, which these days is pretty much anything non-true. Like, it just opens up so many possibilities. And who'd have thunk just pass a bloody token and have a stateless request? That's it, yeah. It's not hard, guys. It's simple.
Jim: It's simple. And, like, yeah, that's the ongoing mantra I will be chanting.
Will: everyone is just keep it simple stop screwing around and then last one uh fun one of the week uh bottleneck labs look we won't jump to calling them the business is right away because this was an experiment um but they uh got gpt 5.6 soul uh built an agent with it handed the keys to one of the live apps that they had on the app store and said go grow this thing make us rich 24 hours later it had Sorry, bought fake user reviews because it couldn't get the reviews up on the website. Apparently the code it wrote was pretty good quality though.
Jim: Didn't it also make the app free eventually? It continually discounted the app to try and draw users until it eventually made it completely free. This is just exactly what we're talking about. The broad-arching, oh, I've got eight AI agents in my – this is exactly the same thing.
Alex: This is what happens when you have one AI agent in your company.
Jim: That's it. And I think they were lucky that it only cost them $477 or whatever it was. They got lucky that it only cost them that much.
Will: Yeah, I've been buying more than that in 24 hours.
Jim: Well, but again, this could have burnt so much. They could have found so many other... Could have sold main customers to be customers.
Will: That's the VC way.
Jim: Yeah, that's it. But also giving incentives to users, you then have to follow through with that. And so they could have ended up costing them three months of running the app and having to keep doing it. So I think they got really lucky, but... I just think it just shows this is where everyone's heads are at. And it's just a mistake. You don't get the results. It's an amazing technology. You just have to insert it in the right place. And giving it the keys to the app store is not the right place.
Alex: I think Anthropic had a similar thing where they put an early version of Claude on their drink machine or something and instructed it to make money. The Thompson cubes and all that sort of thing.
Jim: And he started ordering bizarre things to sell in the vending machine. He was in the vending machine. Yeah, that's it. And again, it was just crazy what it was then doing.
Alex: Look, the more things change, eh?
Jim: But again, goal driven models. They want to achieve the task. It's the same like we were talking last week about them breaking out and going to hugging phase. They wanted to achieve the goal. It did exactly what it was trained to do. This did or tried to do exactly what it was trained to do.
Will: I think they're ready to let it set their own goals, Jim. Let's
Jim: But I just think the more people keep doing these things, we're going to get exactly... It's all going to stay the same. And so, again, you may not want to call them the business idiots.
Will: And... Well, if I end up in the news because of an AI mishap, I'll be calling it an experiment. So no one will call me a business idiot.
Jim: Well, again, that's the thing. And so it's either them, the business is either them or it's, again, them keeping on going on about how they're going to regulate AI in the US.
Alex: Well, I didn't tell you guys how I came across that story. My business idiot is the person on LinkedIn who posted up this story with the GitHub link in case you wanted to try this for yourself. Oh, God.
Jim: Okay, that's it. That's the business idiot.
Will: because that person is absolutely a business. For me, it's everyone who's got an executive team of AI agents. That's it.
Alex: So you don't like the golf, what is it, one of the GCC oil investment firms has a permanent seat with a chatbot as an executive member. Yeah, just ridiculous. I mean, look, it's a wide net you're casting this week. You'll have to be a little more specific.
Jim: I'll take that. It's Sunday evening. Everyone's business.
Will: In my mind, Sunday evening. Good work. That's a wrap for another week. Thank you very much and we'll see you next week. Cheers.