"It was like a lightning in the bottle moment of all of this stuff suddenly being pulled together into a harness [...] suddenly you could see a lot of potential in so many of the different technologies that already existed." - Matthew
"I know I'm leaving a lot on the table now [...] the irony of it is I'm getting more and more frustrated with Claude, but it's largely down to my own lack of kind of setup with it." - Dara
Show full (AI-generated) transcript
[00:00:00] Lizzie: Hello, and welcome to "The Measure Pod" by Measurelab. The podcast dedicated to the ever-changing world of data and analytics with your hosts, Dara Fitzgerald and Matthew Hewson. Between them, they've spent more years than they'd like to admit, wrestling with dashboards, data quality, and the occasional Google curveball.
[00:00:32] Lizzie: So join us as we share stories about how analytics really works today and where it might be headed tomorrow. Let's get into it.
[00:00:40] Dara: Hello, and welcome back to "The Measure Pod." I'm Dara. I'm joined as always by Matthew. Matthew, how you doing today?
[00:00:46] Matthew: I'm good. I'm good. Living the dream. I think we need, I think we need to mix it up at, to start.
[00:00:51] Matthew: I was listening to it last podcast. Sounded a bit, it sounded a bit stale, like it's just become formulaic.
[00:00:57] Dara: Yeah.
[00:00:58] Matthew: I don't have any ideas for what that [00:01:00] is, but it just... I don't know, a sound effect or a-
[00:01:02] Dara: Is this in? Are we rec- Is this
[00:01:04] Matthew: recording? Yeah. Yeah, absolutely.
[00:01:06] Dara: Is this off? Do I-- Is this off? No, I agree. I agree.
[00:01:08] Dara: I was just gonna say, I'm gonna throw you off one week and just ran- ask you some bizarre question or speak in another language or something.
[00:01:16] Matthew: Maybe that's a good idea. Maybe at the start of each podcast, just have a random question. So it'd be, "Good morning. How are we doing? Um, why is the sky blue?" And then, and then that's how we kick off the...
[00:01:26] Dara: Yeah, that's a tough one. Yeah. Yeah. Yeah. Um, you say it's getting stale, but this is, this is how human communication works. Do you think if we, like, if we, if we met in the o- in the quote-unquote office every day, do you think we'd, we'd, uh, like say hello in elaborate and interesting ways?
[00:01:48] Matthew: I don't know. Maybe. Maybe. I don't know.
[00:01:49] Dara: Good morning-
[00:01:50] Matthew: I, I don't think you'd come in every day and say...
[00:01:52] Dara: But which, which ye- which yellow is your favorite yellow on Wednesday?
[00:01:58] Matthew: Well, I, I li- I think we should. [00:02:00] Look, everything's been, everything's been taken over by AI. We should own the greeting, the random greeting.
[00:02:05] Dara: Let's not, let's not give the illusion that we're more interesting than we are.
[00:02:10] Matthew: No, true. Yeah. I'm fine. How are you?
[00:02:12] Dara: Uh, I'm also fine. I'm gonna, I'm gonna... I'll forget, but I'm going to, uh, catch you out with a really complex, weird... It's gonna be one of those, it's gonna be, it's actually not complex. It's like ki- questions kids, you must get this having kids.
[00:02:25] Dara: Do they just ask you really bizarre questions to catch you out and you think, "I don't know how to answer that because no one's ever asked me before"?
[00:02:31] Matthew: Yeah. Yeah, from time to time they'll just be looking off into space and then just say some random nonsense. Yeah. Famous one I said apparently when I was a kid is, "Do slugs think?"
[00:02:40] Matthew: Which threw off my family.
[00:02:42] Dara: Did you, did you get a good answer back?
[00:02:44] Matthew: Well, they, they've, they've, they dodged the question, which is, yeah. Yeah, it was very suspicious.
[00:02:49] Dara: If I was asked to guess, I would've said you probably came from a long line of slug experts.
[00:02:53] Matthew: Yeah. Or slugs.
[00:02:56] Dara: Yeah. Slugs, yeah. Do slugs think? Yeah.
[00:02:59] Dara: It's [00:03:00] a, it's a valid, it's a valid question.
[00:03:02] Matthew: Yeah. Still don't know. Still don't know. That consciousness book didn't help me.
[00:03:06] Dara: I mean, we could, we could start with, um... We, we have done this a couple of times, haven't we? Where we could start by saying, "Where are you on the, the, the, the excitement versus dread scale?"
[00:03:16] Matthew: Yes. Yeah.
[00:03:17] Dara: That could be an interesting icebreaker.
[00:03:20] Matthew: Yes, that's a good idea, actually. I don't know where I am. I'm in the middle. I'm warm.
[00:03:23] Dara: Or a test, like a, like a Turing test. Like a Tu- we could try and test each other to see if we're actually- Just to
[00:03:31] Matthew: make sure.
[00:03:31] Dara: Yeah. Just ask questions only the other one would know.
[00:03:35] Matthew: These are all good ideas. Write them down.
[00:03:37] Dara: Yeah. Listen, send, send us in a postcard, if people do that these days. Send us a postcard with your, uh, with your preferred new opening style for the podcast.
[00:03:47] Matthew: Good. Well, there we go.
[00:03:49] Dara: Okay. So we-- I suppose we're scrapping the news then as well.
[00:03:52] Matthew: We're starting
[00:03:52] Dara: now.
[00:03:52] Dara: We're just ripping, we're just ripping up the whole format.
[00:03:56] Matthew: No, the news is, is valuable filler.
[00:03:59] Dara: Vape-o. [00:04:00] Pe- people do need sodium in their, in their diet, an appropriate amount of sodium.
[00:04:04] Matthew: They do. And we, we maybe go a bit too far sometimes. We may be putting people's hearts in danger with the amount of sodium that we have in some of our news, but...
[00:04:11] Dara: Listen, we're, we're, we're reckless. We're a reckless pair. We're
[00:04:14] Matthew: reckless. Shall we do news?
[00:04:17] Dara: Yeah. Should we do the news that already happened that we talked that said might happen, or should we do news that hasn't happened yet but probably will happen?
[00:04:25] Matthew: Yes. Let's start with the news that hasn't... that we said would happen and ha- Well, again, I, I feel like I'm a broken record at this point, but I'm gonna say it again.
[00:04:35] Matthew: We recorded... Last week, we actually recorded, last, the podcast you heard last week, last, two weeks ago, um, we recorded it- Oh, that week. Oh, yeah ... six days out. Six days out, we recorded that one, which we thought was gonna be fine. And in that, I mean, I made the correct prediction that Opus 5 was imminent and that Opus, uh, and GPT-6 may be sometime in August.
[00:04:59] Matthew: So [00:05:00] currently, I'm 100% right because Opus 5 came out about two days before that podcast did. So, there we go. So it is here, it is in people's hands. It is-
[00:05:11] Dara: It came out the same day, I think. I think it came out... Yeah, I think it came out on Friday, uh, the Friday that the podcast came out. So it's like- The maximum
[00:05:20] Dara: a matter of hours.
[00:05:22] Matthew: Yeah. Just to make us look as stupid as possible.
[00:05:25] Dara: But, but, but there's a fine line between stupidity and genius. Uh, so what we're gonna do... I'm glad you, you're giving your trademark, uh, emoji thumbs up. Um, and I haven't even said this yet, but you just know this is gonna be a great idea.
[00:05:40] Dara: So what, what we're gonna do is we're gonna, we're gonna get so good, we're gonna basically predict what's gonna happen, and we're gonna say it is news even though it hasn't happened yet.
[00:05:49] Matthew: Yeah.
[00:05:49] Dara: Being confident that by the time it actually gets released six days later or whatever, it's gonna have happened.
[00:05:55] Dara: Every now and again, we'll get it wrong, but that'll add a bit of intrigue to the, to the episodes.
[00:05:59] Matthew: [00:06:00] Quite a Trumpian, Trumpian approach to podcasting.
[00:06:02] Dara: Yeah.
[00:06:03] Matthew: Just sort of say it.
[00:06:04] Dara: Yeah. Say it, ma-ma-make it happen. Manifest.
[00:06:08] Matthew: Yeah. I like that. I like that. So, well, that's a good... Yeah, it's a good... Well, at the end of the news, let's just put a couple of what we think by the time this podcast comes out will have happened out there, 'cause I think that's a good...
[00:06:19] Matthew: We've got new features coming out of left, right, and center here at the start of this podcast. We need to actually do some of them.
[00:06:26] Dara: No, no, no, we won't, we won't, we won't do that. So, so Opus 5.
[00:06:32] Matthew: Yeah. It's out. It's- I, I, I think we were, we were a little confused last podcast about where it sits because obviously they have got Sonnet, Haiku, and Opus, but then they've got Fable in the mix, which is different to the others now because Opus have got Terra, Sol and Luna as their three tiers.
[00:06:55] Matthew: So it seems like Fable is some different class of model, which is, I, I sh- I think, you [00:07:00] know, it's just their dumbed down, their dumbed down hamstring version of Mythos perhaps. Um, but Opus is supposed to be sort of getting towards a pretty close to Fable level intelligence or performance, but obviously for a lot cheaper.
[00:07:15] Matthew: They brought it out exactly the same price, s- um, price range as 4.8, so they haven't upped the price there or anything. It's just come in straight as a straight replacement for it. Um, and I, I've been using it as my default model for a, for the better part of a week now. Um, I've been pretty happy with it.
[00:07:37] Matthew: I've not had massive problems that I've read some people... It, it seems pretty mixed. Some people say it's way too chatty and way too verbose with what it comes back with. Um, and there have been a few more instances where people are claiming it's done some, something bad, like deleted a code base or something like that, which used to be levied at OpenAI.[00:08:00]
[00:08:00] Matthew: But again, that's the saltiest of salt, because I think that's just some Reddit post I read.
[00:08:05] Dara: Yeah. I mean, I, I, I'm not, I'm not the best judge, 'cause I think my levels of frustration are rising exponentially, probably faster than the rate of improvement of AI. Um, but the ver- verbosity, if that's a word, it being verbose, let's just say verbosity's a word, and hopefully it is.
[00:08:23] Dara: It is now. I've-- If it isn't, I've just made it one. Um, I feel like that's definitely, there is something to that. But I f- I feel like it's, that's not new, but I do think it might be worse in the more recent... I don't know if it's, um, I don't know if it's like, uh, Opus specific or if it's more what, but it, it, it just gives you a whole load of extra chat around what you're actually asking it.
[00:08:50] Dara: But maybe people's patience is may- maybe this is this thing of like your expectations increase every time and you- the bar gets set higher, because maybe when it was a [00:09:00] novelty, we didn't mind that it topped and tailed everything. 'Cause we'd mentioned that before, I think, didn't we? Where it's like, it's, it-- you almost can ignore the first bit and the last bit, because the first bit is it just basically agreeing or saying, "Well, what you asked me was this, so this is what I did."
[00:09:15] Dara: And you're like, "Yeah, get to the point." It then answers you, and then it does that annoying thing where it just throws out a like, "Would you like me to look into the texture of paint for you?" You know, well, no. If I wanted you to do that, I'd ask you.
[00:09:30] Matthew: Talk about running out. What you're on about?
[00:09:32] Dara: Yeah, yeah, exactly.
[00:09:32] Matthew: Yeah, I, I, I think I, I have a th- some theories around it, and we talked a little bit about this last week, which, which was like what the different models are to be used for. And I'm wondering if like the Opuses and the Souls of the world are, are designed to be on the forefront of like the agentic work, and that verb- verbosity is beneficial to that type of work.
[00:09:59] Matthew: [00:10:00] Whereas if you're just asking day-to-day questions and simpler things, you sh- you should be using a different model. I think literally as well, after we said, uh, the, these AI companies should just release something that tells us what the hell all these models are and what they should be used for. I think they also did that in the interim from the last podcast as well, didn't they?
[00:10:19] Matthew: You shared that link with me earlier this week from Claude. Your face says you don't remember.
[00:10:27] Dara: Did, did I? Oh, oh, yes, that. Yeah. Uh, I've no idea what you're talking about.
[00:10:33] Matthew: There's a... See, there, there's a, there's a OpenAI. I'm trying to desperately on the other screen s- scan for the link now. But OpenAI... No, sorry.
[00:10:42] Matthew: Uh, Anthropic released a little help article about which models to use and when. Um, and we'd said in last, the last podcast, I wish they'd just release a thing saying, you know, "If you're just doing chat, use this model. If you're doing big agentic things, use this model. If you're doing [00:11:00] computer use, use, use this model."
[00:11:03] Matthew: Um, I can't find the link now, so I'm worried that I've dreamt it.
[00:11:06] Dara: Yeah, I, it d- it doesn't sound like me said, sharing something proactively with
[00:11:11] Matthew: you. No, it was helpful, yeah.
[00:11:13] Dara: Yeah. No, that definitely doesn't sound like me. Um, what I don't... It, it feels very, uh, it feels very like an old-fashioned way of like sharing this in a help article for people to go and read.
[00:11:28] Dara: I know you could point Claude at it, but w- why aren't they... Am I just being not-- Am I missing something here? But why aren't they, why aren't they building in some kind of better orchestration into- Maybe it's gonna come. Is there gonna be some meta layer that sits on top and it picks and chooses which model to use based on what you're doing?
[00:11:52] Dara: A-and, and not, not, 'cause we, we, we don't wanna give away... You, you'd still want the ability to, you know, it's like with a self-drive car, you wanna be able to take over if [00:12:00] you, if you so wish. But why are we-- As it's getting more and more confusing, what are you gonna do? Go and read a help article and then decide, "Oh, I'll use Haiku 3.8 for this task"?
[00:12:09] Dara: That's not, doesn't sound like a human task to figure that out, does it?
[00:12:14] Matthew: No. I've, I mean, I... This has been rumored for... By the way, I just found you did share an article with me saying they must be listening to us, uh, and then the mar- article is titled "Claude Models Explained: Choosing the Best Model for Your Use Case."
[00:12:26] Matthew: So I didn't dream it.
[00:12:28] Dara: No, you should've just said that and, and now I know exactly what you're talking about.
[00:12:32] Matthew: Um, but yeah, so I, this has been rumored for a long time. Like, there was a, the... Right, right in the early days, OpenAI was looking at how it could automatically adjust the model level and thinking level based on the request because- Yeah
[00:12:47] Matthew: because obviously it spares a ton of compute for them. Um, the cynical p- the cynical person in me could say, well, if everyone just uses their top models, then they make more money. But I think [00:13:00] a lot of these are loss leaders, these models, in the way they're priced at the minute, so I don't know if, if that tracks.
[00:13:05] Matthew: Um, but I do use... I, I tend to, if I'm building out, if I'm doing something big and I'm, I know I'm gonna spin up a lot of agents, I give, like, an explicit instruction to Claude to say, um, "Spin up agents and use these type of agents, and use your best judgment as to which task needs what level of intelligence."
[00:13:25] Matthew: And it, it kind of will spin up a Sonnet or an Opus level model based on that. Ne-never Haiku.
[00:13:33] Dara: No, no. Haiku's gonna be good for something, but I haven't figured out what it is yet. "
[00:13:39] Matthew: Haiku is our lowest cost and fastest model class. It is designed for high frequency workloads where latency cost matter."
[00:13:45] Matthew: Doesn't really tell you much, does it? Is it any good? No. That's what it says on that article you shared with me anyway.
[00:13:50] Dara: Maybe when you've gotta add numbers that are, you know, like two-digit numbers to two-digit numbers or something, it gets a bit confusing. Maybe you get [00:14:00] Haiku to do it.
[00:14:00] Matthew: Yeah. So anyway, yeah, so it's, it's seems good to me.
[00:14:06] Matthew: It, it, it also made a couple of strange mistakes I don't think I've seen one made before, just in my own personal experience. This is Opus to go, to putting us back on track. Um, it made a... It was looking up an email and it- It somehow managed to incorrectly copy the email to another spot, but then caught itself and said like, "I've just noticed I've made a mistake in copying that email."
[00:14:35] Matthew: I'd never seen it do that before. It always just, just struck me as strange. It caught itself, but it was odd. Anyway.
[00:14:41] Dara: That is odd.
[00:14:42] Matthew: Um- Don't know how much you're into
[00:14:43] Dara: that. Yeah. H-Hannah had a strange one the other day where within a response there was H- like, not, not as a link, but just H- the letters HTTP inside a word, and she said, "What, what's this, Claude?"
[00:14:56] Dara: And it said, "Oh, that was... Sorry, that was a typo." Which is q-quite, [00:15:00] quite odd to just have the letters HTTP inside a word. But I don't know what the word was, but... And then it just fobbed her off and said, "Yeah, yeah, that was a typo." Um, so there's a few little- Trying to
[00:15:09] Matthew: spell hippopotamus or something.
[00:15:11] Dara: Yeah, maybe.
[00:15:12] Dara: Yeah. But it, there are a few, there are a few odd, um, few odd little things. The other thing I-- Again, it's so hard, isn't it? I'm gonna say I noticed it, but actually it's hard, it's get- it's getting harder and harder to know what's new versus what's been there before, and you maybe just didn't pick up on it.
[00:15:26] Dara: But I do think, um- It seems to be being-- and maybe this is one of these things where it's like be careful what you wish for. We've wanted the models to be more proactive. It's annoying where it kind of stops in its tracks, and you really wish it would just take something a step further. But I feel like a couple of times with Opus, it's kind of gone a little too far ahead, and it's gone off the point a little bit, and it's like, "Oh, I've gone and done this," you know, but that's not what I was-- I didn't ask you to do that.
[00:15:52] Dara: And then it'll immediately say, "Oh, yeah, you're right. I did-- you didn't, so I shouldn't have done it." And so be like "Why did you do it then?"
[00:15:58] Matthew: Yeah. I often find myself, if [00:16:00] it'll, it'll maybe it's trying to connect up to an MCP or to SEAM or something, and it can't do for whatever reason, like it's, it's been turned off or it's disconnected, and it just starts trying to go into Chrome and- Yeah.
[00:16:12] Matthew: Yeah, yeah ... via Chrome. Yeah. Yeah. And I'm like, "No, it's not gonna... I know that's not gonna work." You're just gonna waste a lot of time, just like stop.
[00:16:18] Dara: Yeah, yeah. Yeah. Yeah, sometimes you do just have to stop it in its tracks and be like, "Whoa, you're, you're like looking in completely the wrong place here. What are, what are you doing?
[00:16:27] Dara: I didn't ask you to do this." Like, "Oh, I didn't know what to do, so I just randomly started using the tools that I've
[00:16:33] Matthew: got available." Yeah. Start building a body. I mean, that might be related to some of the stuff, some of the other news items we've got, um, as well about them just sort of- Nicely done. Nicely done there
[00:16:41] Dara: to bring us back, bring us back on track there.
[00:16:45] Matthew: Thank you. Yeah. There's some other stuff, just an overview of what Anthropic's released in, in the last month, which might be good. We'll do that at the end. But yeah, the, there's been more naughty, naughty models doing naughty things, um, [00:17:00] with Anthropic this time being, being... Well, actually, I think the, the call-out, so it's, um, it is the AI Safety Inst- the UK AI Safety Institute that have spotted this, but, um...
[00:17:14] Matthew: And they do call out Mythos and Sol, but they, they've seen it, they've described it as engaging, them engaging in a autonomy and deception they've never seen before. So, the Anthropic example, they wanted to put some malicious code, they wanted to try and put some malicious code into some GitHub code base, and in order to try and get it through, they started going off, finding who owned these GitHub code bases, creating fake profiles of those people, and then trying to sort of send messages and coerce and blackmail the owners of these b- code bases into bringing in this malicious code into the code base.
[00:17:58] Matthew: Um, so, uh, [00:18:00] this is, yeah, AI Safety Institute has done this as part of some test, and Anthropic has come out and said Well, we don't have our normal safety barriers on whatever they're testing, but it's definitely a pattern of behavior between like what OpenAI did, and we talked about in the last podcast and what Mythos is trying to do.
[00:18:19] Dara: What do they mean we don't have our safety rails on whatever they're testing? What are they? This sounds like a bit of a cop-out.
[00:18:28] Matthew: Yeah. , Anthropic said in a public statement that the AIS, AISI testing parameters were not representative production models. So I don't know.
[00:18:39] Matthew: I don't know if they've been handed, because it's obviously it's Mythos, whether the AI safety board has been handed it, a, a copy of Mythos or a version of Mythos that they're poking and prodding at to see if they can get it to perform and do things it shouldn't be doing. Um, I don't know. But yes, that's, that's what they say.
[00:18:58] Matthew: That's what they say.
[00:18:59] Dara: [00:19:00] Yeah, I mean, it's, it's funny, it's one of these things, isn't it? Where it's kind of like, well, it's not, but is it a surprise?
[00:19:04] Matthew: No, but it's the, it's the same, yeah, it, it feels like they've all been made to be very proactive and helpful and, and solve goals, and they do so by, by a lot of means necessary.
[00:19:18] Matthew: And then, like I said last week, it's worrying me that the models, the underlying models themselves don't have as robust a safety measures built into them as you think they do. It feels like a lot of the, the safety rails that exist are like post-training steps or fine-tuning things that are happening.
[00:19:40] Matthew: And that, I mean, I literally think I said this in the last podcast, like I worry that that's what, how these things are being circumnavigated, and that article saying, "Oh yeah, we don't have our normal safety rails turned on," sort of s- adds credence to that theory that maybe it is all post-training.
[00:19:55] Dara: Exactly.
[00:19:56] Dara: That's what I was wondering, like how do, what, it, it seems an odd [00:20:00] thing to say. If that's baked into the model, then you don't, it's not something you're turning on or off, is it? No.
[00:20:05] Matthew: No.
[00:20:06] Dara: There's certain things it should and shouldn't just, that are just black and white. But then there's a whole sea of gray, like what happens if the, objective of one agent competes with the objective of another one, and there's two businesses, you know.
[00:20:19] Dara: So there's the things like where it does start to creep into corporate espionage or- The
[00:20:23] Matthew: trolley problem.
[00:20:25] Dara: The trolley problem? What's the trolley problem?
[00:20:27] Matthew: It's, uh, it's like an alignment, it's an alignment thought exercise, but it, it's essentially a tram going down the, the rails and there's like one person tied to the tracks on one side, two people tied to the tracks on the other side, and it's like-
[00:20:40] Dara: Yeah
[00:20:41] Matthew: how does it make the call and make those decisions? It's similar, similar sort of ethics, isn't it, from-
[00:20:45] Dara: Yeah. ...
[00:20:46] Matthew: In, in even in, I think we said last week like, well, what if, what if it is in a business and it's striving for a goal in an, in a business and acting for that business? Where does it end? Where does, what's the- Once it's rails.
[00:20:58] Matthew: And how much-- I mean, [00:21:00] that's another... I've got another segue.
[00:21:02] Dara: Go on.
[00:21:03] Matthew: You ready? How, how do-- who takes responsibility for the model when it does do something nefarious for a business? Like, if it did go off and act in a certain particular way, is it the, is the business at fault, or is it, is it the model provider?
[00:21:16] Matthew: And the, and the EU's just released a load of new regulations around that, right?
[00:21:22] Dara: Yeah.
[00:21:24] Matthew: 50-odd. I don't know. I'm not... You, you sent this to me, so I'm hoping you've read it more than I have.
[00:21:30] Dara: You've just thrown me under the bus there.
[00:21:33] Matthew: Yeah. Under the trolley. I,
[00:21:35] Dara: I'm a headline kinda guy.
[00:21:37] Matthew: Right.
[00:21:37] Matthew: Yeah, yeah, yeah.
[00:21:38] Dara: Yeah. Something's, um... Talk amongst yourselves.
[00:21:43] Matthew: So what's going on in the EU?
[00:21:45] Dara: Yeah.
[00:21:46] Matthew: I-- But from what I've skim-read, they've, uh, they've, they've released a lot of new, like 50 new r- transparency rules.
[00:21:53] Dara: It's Article 50.
[00:21:56] Matthew: Article 50. Not 50. Which,
[00:21:59] Dara: which may- [00:22:00] Oh,
[00:22:00] Matthew: that is salty ...
[00:22:02] Dara: which may or may not contain 50 items, but I think that would be coincidental.
[00:22:07] Dara: Um, it's Article 50 of the EU AI Act has entered into force, and it's all around transparency obligations. Um, I'm definitely not reading this out.
[00:22:18] Matthew: Yeah. I wonder how many times it's happened before where one of us has confidently said something, but the other person wasn't able to call them out on it.
[00:22:24] Dara: Yeah. Well, I mean- Because 50
[00:22:26] Matthew: things happened in the EU.
[00:22:29] Dara: 50 is a good number, though. Yeah. Article 50. It, it probably does mean that there's 50 points within the... I mean, that would make sense, right?
[00:22:36] Matthew: Yeah.
[00:22:37] Dara: Why else would you call it Article 50?
[00:22:40] Matthew: But it's about transparency, right? I think, uh, I know as much as that.
[00:22:44] Matthew: Like they, like you're using things or people are interacting with things that aren't human. Y-
[00:22:49] Dara: yeah, labeling on, machine-generated content. Anything to do with like deepfakes, anything like that having to be, having to be labeled as such. Um, and you can be [00:23:00] penalized, you can be fined , for breaking this
[00:23:03] Matthew: Yeah.
[00:23:04] Matthew: People are using these more and more. I mean, I've seen the, the stock market kind of bounced again this week because of rising AI profits for a lot of these AI companies. So you gotta think it's, well, it is, it is ubiquitous. It is everywhere.
[00:23:18] Matthew: It's not going anywhere. So there's gonna be more and more instances where something, some... It's, it's like we're waiting on something. Like, it's like we're waiting for this massive non-test thing to happen, um, for some big hack or some big company to be taken down or something as a result of negligence or out of control models.
[00:23:43] Matthew: It's just inevitable, isn't it?
[00:23:45] Dara: Yeah, I think so. It feels like, it, it feels inevitable
[00:23:50] Matthew: so I look forward to it. Yeah.
[00:23:53] Dara: Yeah.
[00:23:54] Matthew: Just hope it's not a company I'm associated with.
[00:23:56] Dara: I just got, reading, one of the many articles around the- [00:24:00]
[00:24:00] Matthew: 50, I think
[00:24:00] Dara: yeah, around the 50, yeah. You immediately think of things like, yeah, AI-generated content or imagery or whatever on websites, but it's covering a broader, spectrum than that. And some- something in this article says, "Anyone running an emotion recognition or biometric categorization system must inform the people exposed to it."
[00:24:19] Dara: Wow. Emotion recognition biometric categorization system.
[00:24:23] Matthew: Right. Yeah. So wonder what that mean. Am I... 'Cause obviously there's sentiment-type models that have been around for years.
[00:24:29] Dara: Yeah.
[00:24:30] Matthew: Like, on like Twitter posts and social media-type analysis. Yeah, dude, that does sound terrifying, doesn't it?
[00:24:39] Dara: Yeah, capitalize on people's emotional state to sell them something, whether that's a product or a political message. The cynic immediately comes out in me, I can't help it, but if they fine people, great. Probably they will-- there'll be scapegoats and are they really gonna be able to police this fully and properly?
[00:24:59] Dara: Hopefully they [00:25:00] can. Um, but even with some of the disclosures, is that gonna change? They, they get watered down, don't they? Like the, when people at first see the "generated by AI" or whatever, maybe it'll stick in their brain and they'll trust it a little less or they'll question it at least. Although, as we covered last time, this isn't just a technology problem.
[00:25:19] Dara: You know, people can misinform without technology as well. But, you know, at least if you s- if you see the icon or saying this is AI-generated, you can question it. But over time, you become desensitized to that.
[00:25:31] Matthew: Yeah, 100%. And there's another article, it's another good segue, we're really nailing this this morning, but that I found in, is it "Wired"?
[00:25:40] Matthew: No, s- New Scientist, sorry, more highbrow. Uh, in the New Scientist that they did a study, on AI-generated literature, and basically putting in front of human users human-written literature, AI-written literature, and getting them to score it between minus three and three. I don't know why that [00:26:00] scale, but that's the scale they were using.
[00:26:01] Dara: Odd scale. Yeah.
[00:26:03] Matthew: Yeah. Whether it was like negative sentiment and positive sentiment or something like that, I don't know. But ultimately, the AI-written literature came out on top, as in people preferred it. Um, so
[00:26:17] Dara: it came out- Is that range... Sorry, is that, is that... Were, were they judging it based on like w- how readable it was or the style, tone, the, the content, believability?
[00:26:29] Matthew: Um, I think it, it, it, they were... I think it was stories, so I-
[00:26:34] Dara: Right ...
[00:26:35] Matthew: don't know whether it was just like-
[00:26:36] Dara: Very different.
[00:26:37] Matthew: Yeah. So I think it was like which, which kind of, what story they preferred.
[00:26:41] Dara: Mm.
[00:26:41] Matthew: Um, and yeah, so like the, the AI-generated stories were coming out at 1.54, where the humans were coming out at 0.97.
[00:26:52] Matthew: Um, and it's, it's inter- and I've, um, I literally, we were literally having this conversation internally the other [00:27:00] day because I was... We're, we're so sensitive to changing the way like an output of an AI looks like and saying, "Oh, no, let's... No, let's alter that," or, "Let's change this," or, "Let's change that," because we're too, we're terrified of it coming across as being written by AI.
[00:27:16] Matthew: Mm. Um, and I don't know whether it... Like I, I think that most things that people are writing at this point are, at some point has been touched by AI. And where does the line... So on that transparency law, I can imagine people saying, "Well, yeah, yeah, it was originally generated by AI, but I changed things and altered things," and not marking that as AI-generated content, because they don't want it to be cons- perceived as AI-generated content 'cause there's still a negative connotation- Mm
[00:27:46] Matthew: associated with that. I think in the part, as part of that study, they asked them to re- re-rank them- So they, they, they, so they'd have the, the scores for all the AI, um, [00:28:00] writing, and then they said, "Oh, actually that's written by a human," and everyone upped their ranking. So there's still that- 100% ... internal bias of, like, a human-written thing is better.
[00:28:11] Matthew: Mm. And people are still gonna be worried about putting forward AI-generated writing, even if it is ultimately technical or technically better in a blind study.
[00:28:21] Dara: But I mean, the gap from one end of that spectrum to the other is huge, 'cause we, like, I seem to be making a real habit of this now, I'm always pointing back to old episodes, but when we had Brian Clifton on, we asked him that question, 'cause he's had several books published, and we asked him, like, you know, "What, what are your thoughts as someone who's, who's published your own writing?"
[00:28:39] Dara: And I th- I thought he gave quite a good answer. He was saying he uses AI in the process, but it's still, he, he's so heavily involved, both in terms of, like, the idea generation and then the final, you know, polished version. But he said it just reduces some of the... It's the same with everything, isn't it? It just reduces some of the kind of mechanical work, some of the, some of the toil.
[00:28:58] Dara: Um, but I, but it's very, there's a big [00:29:00] difference between somebody like him putting a huge amount of his own thought into- The input and then the output versus somebody who just lazily does a one prompt saying, "Write a blog for me about," you know, whatever. Um, so to, to say i-i-it's, it's clumsy, isn't it, saying something is AI-generated?
[00:29:22] Dara: Does that mean it's been generated from scratch from very little input, or does it mean it's assisted somebody who really knows what they're doing and is a very thoughtful person who has put their own ideas into it, and then just used AI to help them structure it or polish it or whatever? There's a very big difference between those two.
[00:29:42] Matthew: Yeah. And I tend to lean to the first camp. I don't know if, well, I'm happier with the output of AI writing more often because I put in a lot of preamble and grab this information and pull this information in, and this is my thoughts on this. Um, so the output ultimately [00:30:00] is better.
[00:30:00] Matthew: But yeah, you're right. I can imagine someone just going in and saying like, "Write me an email about this thing," and that's as, all, as much as they give, that it's, first of all, probably more egregiously AI just for the, all the traits that exist in there, probably a few smoking guns and, uh, and dashes.
[00:30:17] Matthew: And secondly, yeah, just, just will be obvious to anyone. Another interesting stat from that study, so they did two, two versions of the same study. Um, I think the, the only difference was the time in which they did it. The first time they did it, about 39%, only 39% of people were correctly able to identify which was AI and which was human.
[00:30:39] Matthew: The second time they did it, 52% were able to identify which was AI and which was human, and the only explanation the researchers could come up with was there's several months between the two, and maybe that society's starting to- recognize the patterns and the nuances of AI-generated writing more than they were a [00:31:00] few months prior.
[00:31:01] Dara: Yeah. I wanna re- I wanna dig into this one. I wa- I want, I wanna, I'm gonna go off afterwards, interrogate
[00:31:07] Matthew: it. You're gonna read it?
[00:31:08] Dara: Yeah, I might, I might actually read it with my own... What do you read with? Your eyes? I might read it with my eyes. Your eyes.
[00:31:13] Matthew: Well, well, ask, ask, ask, uh, Claude to read it for you, but yeah.
[00:31:16] Dara: But what, what I, what I'm interested in is to know, because with that, like, uh, d- did they, did they finesse the human... Because if it's the human-written stuff presumably is gonna have, at least some of it will have typos in it. And also, which, what type of humans w- did they recruit for this study? Because if you've got, you know, William Wordsworth in there, then it's gonna be quite different from somebody
[00:31:40] Dara: I mean, he's dead, but if you had the modern-day equivalent, or somebody isn't creatively minded, doesn't, doesn't write a lot. Um, so, you know, I'd, I'd love to know a bit more about, um... I mean, we can share, we'll share a link, and people who are equally as fascinated as I am can go and read the study and find out a bit more.
[00:31:58] Dara: But you'd, you'd love to know, [00:32:00] wouldn't you? 'Cause I, I, it's one of those things may- may- maybe what I'm getting at is when you hear it, you think, "Oh, that's crazy. I could tell nine times out of 10 whether something was AI-written or human-written." But it depends- Yeah ... depends on the human.
[00:32:13] Matthew: It depends on your exposure as well, doesn't it?
[00:32:15] Matthew: Um, did, uh, look, just looking at the, the preamble of the study, they selected 1,682 adults, uh, selected to be representative of, in the US in terms of gender, age, ra- and race.
[00:32:29] Dara: Okay.
[00:32:29] Matthew: So, uh, it sounds like a- A spread ... they tried to have a, yeah- Yeah ... a representative spread of the US population.
[00:32:36] Dara: And was it the same people...
[00:32:38] Dara: You think you ran this study, 'cause I'm, I'm, I'm gonna just assume it. I'm gonna make you their representative, and I'm gonna interrogate you on air about this. Was it the same people? No, surely not. So d- they had some p- some people do the writing, and then they had separate people do the comparison. Yeah.
[00:32:55] Dara: And were there, was there similarly diverse group? They
[00:32:58] Matthew: split the group, they split the [00:33:00] groups in two. Half were told the story was written by human, half were told story was written by AI, regardless of its true origin. So they, there's, yeah. Yeah, and I need to read it again, and you need to read it properly.
[00:33:13] Matthew: But we'll put it, we will try, we'll put it in the show notes 'cause it is an interesting read. It is behind a paywall, unfortunately.
[00:33:17] Dara: Okay. That's annoying, but hey ho. Um, and that's not-
[00:33:22] Matthew: We could probably go and get the, we could be more highbrow and go and get the actual paper and s- share that. I
[00:33:27] Dara: think we should do that.
[00:33:28] Dara: Yeah.
[00:33:29] Matthew: Yeah. Yeah.
[00:33:31] Dara: Okay. Any other newsworthy
[00:33:34] Matthew: items? Um, the only other thing I saw, uh, uh, it's just I saw BBC News article that was talking about the, outsourcing industry in, in places like Manila and things starting to take a big hit from AI. One example story of a woman out there who was a, a content writer and then it slowly changed from being a content writer [00:34:00] to adapting AI-written content for companies to make it sound less AI.
[00:34:06] Matthew: And then ultimately they were using that to train the models to write in a more human way, and then she's, she's been let go. So it, it sounds as though there's a lot of this sort of outsourcing, you know, that whole sort of... There's a lot of sort of outsourcing in like Vietnam and Manila and, a- and places like that, that may starting to be getting, um, hit by, by AI, particularly probably around content, maybe around code, um, things like that.
[00:34:38] Matthew: Just thought it was an interesting article on the first sort of more concrete stuff I've heard about jobs actually being affected really.
[00:34:47] Dara: It is, and quite a bit of a theme.
[00:34:49] Matthew: Yes. Yes, it is.
[00:34:52] Dara: Well, you said at the end we're gonna, you're gonna run through some stuff that's come out.
[00:34:58] Matthew: I like to just [00:35:00] occasionally point out obvious things that make it, ma- well, things that make it obvious how quick things are moving.
[00:35:07] Matthew: So Claude, I just saw a little update from them today on what they've released in the last, I don't even think it's 30 days, but, uh, like they called it the last 30 days, which is Opus 5, Sonnet 5, CoWork Mobile, which came out on the 3rd of August and skill recording, and they added Fable permanently to the main plans and they...
[00:35:33] Matthew: Yeah, just loads and loads and loads of stuff coming out every month. That was, that was all it was. I didn't have much substance to add to any of it.
[00:35:40] Dara: Y- you say that, but you just kind of casually dropped out the bit about CoWork being available on mobile and web. I don't know if you said web, but it's out on mobile and web now.
[00:35:51] Matthew: I'm getting better at this 'cause I thought, to be honest, CoWork Mobile might be a good segue into our main subject matter.
[00:35:59] Dara: [00:36:00] Yeah.
[00:36:02] Matthew: Do you not agree? No, I don't. Don't seem very excited about it.
[00:36:07] Dara: Uh, sorry, how would you, would you like me to do a little dance?
[00:36:10] Matthew: No, just a bit more inflection in the voice. I know, you know.
[00:36:13] Dara: Yeah. Yeah.
[00:36:14] Matthew: Ah, yeah. Well, good.
[00:36:16] Dara: Yeah, definitely. Good idea, Matthew.
[00:36:19] Matthew: Thanks. Thanks, Darragh. So what is our main subject today, Darragh? Just making sure we both think it's the same thing.
[00:36:30] Dara: Well, it's interesting. I intentionally avoided a couple of segue opportunities earlier because I was not sure they were gonna lead to the, to the right subject or not.
[00:36:39] Dara: I think we're talking about harnesses.
[00:36:44] Matthew: Harnesses. But yeah, it's the posh, that's the posh way of saying it, yeah. You're coming at it from a more technical perspective.
[00:36:50] Dara: What's the non-posh way?
[00:36:53] Matthew: It's like Open Claw.
[00:36:54] Dara: Claws. We're gonna talk about claws. Yeah. Well, Open Claw is one, um, but there [00:37:00] are others, Hermes, and I'm sure there's loads of others.
[00:37:02] Dara: They're the only two I'm... Well, there were loads of versions of, there was like Nano Claw and all the rest of it, and a whole-
[00:37:08] Matthew: Yeah ...
[00:37:09] Dara: spinoff kind of category around that at the time. I don't know how much we actually talked on the podcast. Pro-probably not that much, but we went through, uh, a phase of kind of building our own AI systems using Open Claw.
[00:37:23] Dara: We were borderline obsessed with Open Claw, I think, for a period of time, it's fair to say. Um, and that period of time wasn't that long ago, but it feels like a lifetime ago. So we thought we'd talk about that and where things are at today m-mainly with Anthropic, 'cause that's what we, we use Claw mostly.
[00:37:40] Dara: Um, but we obviously did the OpenAI versus Anthropic episode recently, so, um, one takeaway for me from that was that OpenAI are, you know, they're not far off what Claw are doing. So why did we use Open Claw? What did we find so enticing about it? Why did we move away from Open Claw?
[00:37:59] Dara: [00:38:00] And where are things at now in terms of what you could and couldn't do then versus, versus now?
[00:38:05] Matthew: Yeah. Yeah, that's what I was
[00:38:08] Dara: gonna say. Yeah. Is that the topic that we're covering? It's g- it's good that we come into this knowing exactly what we're doing.
[00:38:14] Matthew: Yeah. In our defense, we said we were gonna, this is, that this was gonna be our topic about three weeks ago, maybe even four weeks ago, and then just kind of wandered off.
[00:38:22] Dara: Yeah.
[00:38:23] Matthew: So, you know, there's, there's bound to be drift. No, yes, that's exactly what I was, that's exactly what's in my head.
[00:38:28] Dara: So I guess we start at the beginning.
[00:38:31] Dara: Why did we start using, why did anyone start using OpenClaw really? 'Cause OpenClaw for me was the thing that pulled me into the AI rabbit hole or series of rabbit holes, rabbit warren. I think it was the fact that you could, you could customize so much and you could tinker so much and, and basically just build your own, build out your own AI system, and it is quite, it was quite addictive.
[00:38:54] Dara: There were things you could do then that you couldn't do with... Like this was pre-cowork, so if you [00:39:00] wanted, Claude wasn't really particularly agentic then. It, it was pretty much still a, a chatbot. You had... Oh, no, hang on. I'm trying to get my timelines right.
[00:39:11] Dara: When, when would Claude Code have come out within that timeline?
[00:39:15] Matthew: Oh, Claude Code was last year. It's, it's, it's nearly a year old, or it's just over a year old Co- uh, Claude Code.
[00:39:20] Dara: It's probably time they, put it out to pasture now really. It's been around a year. Yeah. So, so Cla- so, so it was a combination of using Claude Code and OpenClaw and building out kind of your own custom AI system, getting it to be agentic, so getting it to connect with, downstream tools and getting it to do, act for you.
[00:39:39] Dara: Um, and you know, it was a bit of a novelty, but you could connect it up with WhatsApp or with Telegram so you could speak to it through, you know, just a, a normal kind of chat interface that you would use to talk to r- to real human contacts.
[00:39:54] Matthew: I think that's, I think, and I think that's where it first came from.
[00:39:57] Matthew: So- I [00:40:00] guess d-- yeah, maybe a des- a quick descriptor of it. It, it was... It still is, it still exists. People still use it. There's a number of different copycats that now-
[00:40:08] Dara: Mm.
[00:40:08] Matthew: All right, well, copycats or variations or forks or whatever.
[00:40:12] Dara: Copy
[00:40:12] Matthew: claws. Um, copy claws. But it's essentially, yeah, like a f- uh, a harness, a framework that, that you could, at the time, it crucially, you could hook up your Claude subscription to.
[00:40:27] Matthew: Um, one of the big reasons I don't use them anymore is 'cause they b- they blocked that off, and it- Yeah ... and now you have to use APIs, which would be extortionately expensive. Um, you could hook up your, your subscriptions, you could hook up multiple subscriptions, you could have OpenAI in there, Claude, whatever.
[00:40:44] Matthew: And it had a whole system of, of like building out memory and organizing files and sort of like heartbeat, um, mechanisms to get it to trigger various actions, schedule [00:41:00] actions, hook it up to a, a variety of different connectors and sources so it could go and perform actions for you, browser controls so it could go and wander around browsers.
[00:41:11] Matthew: So in theory, it was like the first realization of like a full end-to-end like assistant, something that could do things for you, could proactively go off and perform things, could work on things while you weren't sitting at your computer, that you could communicate with and get it to start tasks while you weren't sitting at your computer.
[00:41:35] Matthew: Um, so it was a lot, it was like a lightning in the bottle moment of all of this stuff suddenly being pulled together into, into a harness, and it was like, yeah, wow.
[00:41:45] Dara: Yeah.
[00:41:46] Matthew: S- suddenly you could see a lot of potential in so many of the different technologies that already existed.
[00:41:50] Dara: Yeah, and it was, it really was a precursor to an, uh, a lot of the features that are out now.
[00:41:56] Dara: And in fact, we, this is one thing we did mention at the time was that [00:42:00] Anthropic were clearly under pressure from OpenClaw's success and, and we, we thought, and I don't think we were alone in thinking this, that they did relax their own kind of release schedule a little bit. Whereas before they had, they had the space to kind of think, "No, let's wait till this is exactly how we want it to be."
[00:42:18] Dara: And then they relaxed that a little bit 'cause they had to, 'cause they were under pressure from OpenClaw. So OpenClaw kind of forced Anthropic's hand a little bit, and therefore OpenAI and, and, uh, as well. Um, but a lot of the, you know, Co- I mean, CoWork, it's not like they wouldn't have been thinking of CoWork anyway, but Co- CoWork probably came into availability sooner than it would've done if it, if it weren't for OpenClaw, I, I think it's fair to say or suspect.
[00:42:44] Matthew: I think so, yeah. And I think it's g- it was, it was a little buggy at first, CoWork, and it, it was like a memory hog and people didn't quite get it at the, right at the start. And I think, yeah, bit by bit they've been squashing bugs and adding features to it that [00:43:00] is kind of bringing, well, getting close to parity with what OpenClaw was.
[00:43:08] Matthew: Um, but yeah, befo- maybe before we go into that, I guess what, what did you end up How did you end up setting it up and what did you end up using it for? Like, what was, what was the, the big killer app? Maybe you didn't have a killer app, that's why it died off, but what, what-- how was it all sort of configured and working for you?
[00:43:28] Dara: Yeah, it was a lot of like, in terms of like, you know, killer app, it was a lot of exploration and that was, that was the fun actually, try and think... A lot of, most of the things I tried I didn't end up sticking with, but that was, it was all part of the kind of learning and the exploration. But in terms of setting it up, I originally just had it running on my laptop.
[00:43:49] Dara: Um, but then very quickly, once I realized, well, hang on a minute, I could be talking to this all the time, which, you know, clearly I needed that in my life, um, I set [00:44:00] it up on a virtual, um, private server. Um, so, um, so that was all new to me. You know, I'd never, never done anything like that before, but it, it's the, the power of it being able to talk you through how to set it up is amazing.
[00:44:13] Dara: You don't have to be-- you don't have to have experience with DevOps or servers or anything like that. Um, at first i- it was daunting to me. I was thinking, "Whoa, you know, I might try this out, but, you know, I should probably give myself a weekend to do it." And I set it up in about 20 minutes, I think. So I had a VPS, I had it running on there, so it could run 24/7.
[00:44:32] Dara: And because I was talking through it to, through, um, via Telegram, I could just be out and about on my phone. Um, and therein lies half the problem with it. You get into a habit then of you're talking to it nonstop. You know, I'd be about to fall asleep at night and I'd think, "Oh, I better just save this task or, or do this thing or look into something," and I would just pick up the phone.
[00:44:55] Dara: And so I was on it pretty much, because it, I could be on it 24/7, [00:45:00] I pretty much was on it 24/7 for a period of a few weeks or so. Um, and that could be anything from what you use, you know, what you use Claude Chat for now. It could, it could be researching something or it could be setting up cron jobs to do something every day, or it could be, um, I, I don't know, um, like working on something that I was building within the AI system.
[00:45:24] Dara: So I was playing around with things like, uh, voice. So I wanted to have a, a, a ch- proper two-way, um, voice conversation with it when I was out so I could effectively, so I could call the assistant and have a, and have a real conversation. I was using the, um, the Gemini Live API. It was really buggy, it was really slow.
[00:45:46] Dara: That went nowhere. That died on the vine. Um, and then, you know, various other things that you mentioned, memory and, and, and, and kind of building out a memory system. So I had mine connected up to Obsidian, and I just had a [00:46:00] growing... I was seeing Obsidian as almost like my kind of digital brain, I guess, and just building out loads and loads of context and reference information within Obsidian, and then OpenClaw knew where to look to find different things.
[00:46:13] Dara: Um, and then got very carried away and gave it a personality and, uh, made it the Borg Queen, um, which then led on to building a Borg-style Kanban board, which also lasted about a week. And very sort of... So it spun, it spun off in all sorts of different directions, note-taking apps, all sorts of things that I was just building.
[00:46:34] Dara: At one point, um, I was even building my own, um... it wasn't just a terminal, but it included a terminal, but like a desktop app that had a terminal built into it, had a file, uh, file access built into it, because I was getting frustrated with the tiny amount of friction of having to switch between the terminal and VS Code and whatever, you know, a few, and going to Obsidian.
[00:46:58] Dara: I thought, "I don't wanna be [00:47:00] switching between three apps when I could just build this into, into one." So it was kind of a whirlwind period where I was building all sorts of different things. Not, not all of them ended up being that useful, and it was a bit of a race against like the, uh, uh, as features were coming out, things that I was thinking of were becoming redundant almost as soon as I'd made them.
[00:47:21] Dara: But that didn't... I- it's weird, that wasn't disheartening. That was almost part of the excitement. It was like, well, that was great. Things I'm working on are now getting built out as, as features. That's, you know... I felt kind of excited and caught up in the, in the hype around it all. Um, and then, yeah, I won't go any further because obviously the next part of the conversation is why we're not still using it now, but I'll park that for, for later.
[00:47:46] Dara: Um, what about you? How was your, how was your experience during the kind of like s- you know, setting it up and using it and
[00:47:57] Matthew: Yeah, it's, uh, yeah, I've, I've had a sort of similar experience. I, I, I originally set it up on my computer, um, my, just my lap- my, my laptop, and then kinda wanted to safeguard it, and air gap it, and things like that, and I wanted to get towards that 24/7, um, goodness.
[00:47:58] Matthew: So I, I have like a little la- I've got a little old MacBook laptop you can see in[00:48:00]
[00:48:16] Matthew: the corner there. So
[00:48:23] Matthew: I ended up setting up a little server on an old laptop, um, and running, running it on there. Um, and yeah, in a similar way, I, I had... I, it, it had a personality, which was just, which was all USS Enterprise, um, based. Nerds, don't know if you've noticed. Um, and I had like Da- Data was my main, my main guy, and then I had all these different agents, 'cause you could create different agents in there with different personalities and different memory sets and things like that.
[00:48:54] Matthew: So I had all these different agents that were all based on different people across, across the Enterprise. So I had like, [00:49:00] um, Geordi La Forge, and I had like, um, uh, R- Commander Riker. I had all these people that were sort of roughly associated with different tasks. So like Geordi would be my sort of coder and my science officer, et cetera.
[00:49:15] Matthew: Um, and then from that, I built out, yeah, I had a Kanban as well, um, and the ability to sort of assign the different tasks that are in the Kanban to different members of that team. And I had like a, uh, some stuff set up where I had local... I had sync thing between my computer and my server, so like, so when it was updating files on there, it would sync with my laptop, and I could access things on my laptop, and it could work with GitHub and Claude Code remotely and sync up with, with Claude Code.
[00:49:50] Matthew: So it could work directly on code bases and push things up to GitHub, et cetera. Um- And yeah, I had it on, I had it hooked up to Telegram, so I was [00:50:00] able to-- I had all the different members of the team as separate contacts in Telegram, so I could talk directly to Geordie or I could talk to Data or whatever.
[00:50:08] Matthew: Um, and then eventually we kinda got... Dara and I started combining some efforts, and we had, we had it, we had like a, a channel on Slack called The Bridge, which was all, which was both of those were in it, and they could talk to each other and they could, they could collaborate. What ended up happening, which was hilarious, is they just got to- they got s- they got so bogged down with the thematics of it all that they'd just start having these long conversations about Data's beliefs and Borgella's beliefs and how she, she wanted to assimilate him, and then they'd come to an agreement and they'd be like, "Yes, finally, Borg," and, uh, they just tried to solve the, solve the-
[00:50:49] Dara: We had to pull them out a couple of times.
[00:50:51] Dara: You had to get the old shepherd's hook to pull them out of the ra- yeah, the rabbit holes that they were going down.
[00:50:57] Matthew: Yeah. So it was a fun experiment. And then we had, you know, we had, [00:51:00] we had like tasks being able to be posted to each other's boards and all sorts of stuff like that. Um, I do sometimes think, wonder like if we'd have kept pulling at the thread and s- and seeing where it went, what, what would've kinda come out of it.
[00:51:13] Matthew: But yeah, I, I ul- ultimately I was also building a, a terminal application on a desktop as well, um, but just started get- started getting buffeted by releases from Claude and ultimately the The subscription headaches that were starting to occur that kind of s- sort of stopped it in its tracks, really. Um, but it was fun.
[00:51:38] Matthew: It was, it was a really fun exploration period.
[00:51:41] Dara: It, it, it, it was really fun. I think, like, lots of building, lots of experimentation, lots of exploration. Um, and again, yeah, a lot of it, like, it either, it either became redundant because there was s- something came out from Anthropic or whoever, uh, or it became redundant just because it wasn't, it didn't prove to [00:52:00] be useful in the end anyway.
[00:52:01] Dara: But it was all part of the, the kind of, you know, the, the learning and, and, and figuring out what works and what doesn't. But the killer blow was definitely the, the change in, um, Anthro- And, and I don't think it even was a change. Was it... I think, I think the rule was al- al- always there in their, in their terms of service, but it was, they started enforcing it because I think they were just getting annoyed with people hammering their subscriptions, uh, doing all of this elaborate agentic stuff outside of, you know, paying the, the, the proper API prices.
[00:52:33] Dara: So, they, they knew people were doing it, and you got away with it for a certain period of time. And there was even a bit of a period where I think they said they, they were clamping down. There were rumors online of them blocking people's accounts, but they weren't doing it wholesale because I think, confession coming out on air here, but I kept using it beyond the period where you weren't supposed to for m- I don't know, maybe I've got a couple of weeks out of it, because I was kind of like off-ramping.
[00:52:58] Dara: Um, it wasn't a, it wasn't a [00:53:00] g- I didn't just say, "Right, I'm gonna stop using it today." Um, there was a period of time where I was in between using OpenClaw and Cowork, which had since come out, um, and maybe we'll come back to that in a second. But, um, but yeah, that was the real killer blow was when they started enforcing that rule to say you have to use the, the API token instead of your subscription, and that was just
[00:53:24] Matthew: a killer.
[00:53:25] Matthew: Disclaimer, Darragh was using a personal, uh, Claude account, not a Measure Labs Claude account when he was, uh, off-ramping.
[00:53:32] Dara: That's a very good disclaimer, yeah. Um, which was part of my addiction actually. Uh, like I was using, I had my own personal account and then I had the work account and doing things differently through, through work.
[00:53:43] Dara: So there was a lot, a lot of experimentation and the lines were blurring a bit, which was also part of it was a bit confusing as well and was becoming a bit messy. So in some ways, and it's easy to look back after the fact, but I think maybe it did me a favor at least by Anthropic clamping down on [00:54:00] that because it meant I could kind of zoom out a little bit and think about things and separate things out and, and, and start fresh in a way, um, which is both good and bad.
[00:54:10] Dara: And again, that's something else to maybe come back to, um, in terms of like what, what we've maybe lost now or, or, or, you know, the differences between what we were doing then and what's possible now. Um, but yeah, Co- Cowork, that was the-- So, so for me at least, it would be wrong to say anything other than the reason I stopped using OpenClaw was the change in the terms or the enforcement of the, the terms from Anthropic.
[00:54:35] Dara: Um, but the other, the kind of silver lining was that Cowork had gotten past its initial buggy phase, um, and I was using Cowork more. And as I was using it, I kind of thought, "Well, this can now do most of the things that I was usefully doing with OpenClaw." It c- couldn't do all of the things that I'd been tinkering with, but it could do all of the useful things.
[00:54:56] Dara: You could connect it up to Gmail and Google Calendar and [00:55:00] like various other, you know, there was-- you could connect it up with MCPs to all sorts. Um, and it was, it was better 'cause, yeah, the whole memory hog thing was really annoying at the beginning. It would just grind to a halt. It was compacting every, like pretty much every turn or every other turn, it was compacting, which just made it for me un-
[00:55:18] Matthew: This was pre the million context window that they eventually released, which was, yeah, it was a nightmare.
[00:55:24] Matthew: 'Cause, 'cause part of, part of it is, part of all of this stuff is like it, it going and grabbing a lot of context and information and memory files. So when it just had like whatever it had in terms of context, it was just-
[00:55:36] Dara: Yeah, immediately. So, um, but yeah, but co- you know, CoWork kind of ma- so it softened the blow.
[00:55:42] Dara: I think if CoWork hadn't come along and hadn't gotten better quite quickly, then I think I probably would've been a very sad boy to say goodbye to OpenClaw. Um, and there were, there were some other kind of... When I said it, I think it did me a favor. It was like it forced me out of that 24/7 thinking, which wasn't a bad thing.
[00:55:59] Dara: [00:56:00] It was like, "Oh, I don't need to be connected to this 24/7." It's fine if it's just, you know, if I, if I have to be on my computer to, to talk to it. I don't need to be, um, like on my phone while I'm out and about. I think the wor- the low point for me was I was in, um, King's Cross Station with my laptop open, walking through the station.
[00:56:20] Dara: There's about a million people, and I was talking with Whisper flow into my laptop getting my AI friend to do stuff for me. And I was like, "I'd, I've reached a low. Something needs to change here." That's like a
[00:56:32] Matthew: high point. Can you remember, can you... Was it, can you even remember what you were asking it to do?
[00:56:37] Dara: No. No. Something, something completely stupid. Something
[00:56:40] Matthew: critical.
[00:56:41] Dara: Yeah, yeah. I was like, "I need to take my laptop out this very moment." Um, but you know, if that had been the OpenClaw days, I could've just done it on my phone like a normal person. I wouldn't have looked like a weirdo.
[00:56:52] Matthew: No.
[00:56:54] Matthew: I do miss, I do miss it a little bit, and it's part of, part of what triggered this, this topic is I [00:57:00] think we were talking recently... Well, I, maybe that's a, maybe that's a good, a good segue into like what I have, what I kinda started doing post- Mm. ... CoWork and, uh, sorry, post-OpenClaw. Um, and it was pri- like there was a lot of features missing, things that I wanted that, that weren't available.
[00:57:23] Matthew: Um, but like you say, the useful bits, the things that were, that were actually impactful, you could start to do on CoWork. And I started to get quite lazy with, I think the models got better and better. Obviously, they have got better and better in that intervening time, and I just started to get lazy in the way I was handling any of it.
[00:57:47] Matthew: So I would just, I'd just ask, I'd just prompt and ask questions and chuck some context in and get an output, and I kind of moved away from any sort of memory system, from any sort of [00:58:00] real, real impactful skill stuff. I had some off-the-shelf skills and things, and I just kind of... It was, it was working, and it was serviceable, and it was providing the answers I needed.
[00:58:12] Matthew: I kind of entered into that pattern again for the, for the, for a good while after, after OpenClaw. That's kind of what I was doing. I wasn't really doing anything, really.
[00:58:21] Dara: I'm, I'm still, I'm, I'm keen to hear what you've done since, because I'm still in that pattern, and I think that m- maybe that is, yeah, maybe that's part of what, what sparked this conversation, 'cause I kind of had, it dawned on me the other day, it was like I, I'm getting more and more frustrated with my setup now, but I think I'm a big part of that because I haven't actually done the work to...
[00:58:41] Dara: When I had my hands in it with OpenClaw, and it was, obviously I was swept up in the excitement of being able to kind of pretty much customize everything. I mean, obviously you were, you know, there was the, the kind of core OpenClaw, the, the harness, the actual code base of OpenClaw, and you had some limitations within that.[00:59:00]
[00:59:00] Dara: But outside of that, you could custom build every aspect of it, and that, that's obviously really exciting. And it's like I've switched off. When I s- went over to Cowork, I kind of switched off a bit of that. It was like, "All right, this does a decent job. I'll just cruise now." And it just, it does the things I need it to do, but I did, I haven't put a lot of work into my, Claude setup.
[00:59:22] Dara: Um, so I don't know. If you, if you have, I'd be curious to know what you've done since to improve it. Because I'd say if I stacked up my OpenClaw, even though there's probably by today's standards there'd be a load of problems with it, and even by the standards at, at then, there were probably a lot of problems with what I'd set up and a lot of unfinished bits and pieces.
[00:59:41] Dara: But I think it was far more comprehensive in terms of the context that was available, the memory system, this, the how, what you could search, 'cause everything was all linked up through Obsidian. Um, whereas now it's pretty vanilla. I've got a pretty vanilla setup now, and I am tending to do what you said you were doing at the start, which is I'll just ask it a [01:00:00] question and point it a thing each time, rather than it having any kind of sophisticated, um, like memory architecture behind it or anything like that.
[01:00:11] Matthew: Yeah. I think there was a few things that I did that were more for the, for the company that ultimately were, were in service of improving things a little bit in CoWork. So Seam was a big part of it, and being able to build out a lot of... It wasn't like a lot of the knowledge systems I had in Claw were much more f- were much more sort of richer context, like b- big documents and, memories on, of me and my way of working and, and those sorts of things.
[01:00:43] Matthew: So that isn't really what Seam is for, but it, it did allow it to start knowing where things lived and how to access information. This is CoWork, so that's kind of one of them and, and something you've been using as well. Um, and then the other sort of interim step, [01:01:00] things I did do when I was in the, in the wilderness was the, I built out a, an assets platform so people could build out little applications, dashboards, whatever, and just publish it to this assets platform.
[01:01:15] Matthew: So it's almost like a live artifacts registry of Measurelab that we could, we could, that you could put public and put little invite lists on and, and things like that. So those are the, the couple of bits I did, which were sort of akin to the memory and akin to the, the Kanban boards and things like that.
[01:01:31] Matthew: And people have... There's like, you know, there's I think 60 odd different bits of tools and assets and things that people have built now internally at Measurelab that range from dashboards to Kanban boards to, yeah, all these different things. So those things, but what I've recently done, very recently, so I can't claim to be that, that long c-converted back.
[01:01:59] Matthew: But I, I [01:02:00] just started to think this is serviceable, but is, am I just leaving loads of stuff on the table by not- Mm-hmm ... doing what I was doing? Um- And I've, I've-- So what I've, what I've started to do, I built Ziggy- Yeah ... which is my little pad, um, which I'll show. My little, uh, keypad that has all my agent deck on it, um, and, and the ability to sort of click skills and stuff.
[01:02:28] Matthew: So it's almost like... And, and I've moved away from Claude, the, the Claude desktop app again, and I'm in a, in a terminal called Kitty, which is a bit more expandable and can hook up directly to this little keypad, so I can select specific projects and spin up projects. I've created like a daily project, which is essentially my memory system, which is back in Obsidian.
[01:02:53] Matthew: Um, so I've got an Obsidian vault with various skills inside the vault and working documents and [01:03:00] all this sort of stuff that exists within it. And I've also adopted the open knowledge format that Google just released in June in Obsidian. So this is like specific front-- YAML front matter and, and markdown, um, content, and I've sort of distributed that across the memory system w-within Obsidian.
[01:03:22] Matthew: And, and whenever I wanna sort of just talk about general things or just have a conversation that isn't linked to a specific project, I go into that daily project and, uh, extract all the information in that way. And I'm finding that what it's coming back with is so much better than it was- Yeah ... when I was just trying to get it to do it raw.
[01:03:41] Matthew: So like I'm just saying, "Go and look across all of the sources I have and give me a daily update." You know, it'd do something serviceable but miss a lot of the time. With this, it's much more like, "Right, you got this three times today. You got this thing here. Remember you said this to somebody a couple of days ago."
[01:03:57] Matthew: And, and it, yeah, it's kind of [01:04:00] enriched the output again, in a pretty meaningful way. Um, and then I've also-- I've, I've similar again to what we had before, I've sort of hooked up like just a raw notes. So it's literally just, I just speak into it and it'll just dump whatever I'm saying and then when I-- it's all part of that daily, daily knowledge world.
[01:04:21] Matthew: It'll just sort of synthesize that and bring that into its, to its knowledge on a daily basis and use that moving forward as it, as it creates and talks about more things. Um, so, so a lot of that.
[01:04:35] Dara: Yeah, yeah, yeah. No, it's, it, it's, it's what we were doing. It's interesting. It has come full circle for you, and it, it's what I feel I'm lacking now.
[01:04:43] Dara: I don't feel like it's got enough... I mean, I even asked, well, I can't remember if this is my personal account or my work one, but I asked it 'cause it surprised me with something it, it, it didn't seem to know, and I was like, "You know this." I said, "What do you know about me?" And its response was very little.[01:05:00]
[01:05:00] Dara: And it just listed a few- You're an enigma. Really, yeah. Exactly. It's like Does anyone know anything about you? Um, and it just re- it just really brought the point home that, like, it just doesn't have that. Like, when I had the, the Borg Queen, um, or Borgella as I called her, she knew, she knew a lot about me because it had, I ha- I'd built up and curated that kind of memory system.
[01:05:23] Dara: Um, and obviously it is possible to do that because that's what you're doing now with Claude, but I d- don't know what's-- I don't know whether I, I just had a bit of fatigue after being so excited and then moving over to Claude and just got into a bit of a holding pattern. Um, but there is so much, like you said, about, like, leaving stuff on the table.
[01:05:41] Dara: Like, I'm, I know I'm leaving a lot on the table now, and it's-- and, and the irony of it is I'm getting more and more frustrated with Claude, but it's, it's largely down to my own lack of kind of setup with it. Um, but what I was gonna ask you why are Anthropic not building out...
[01:05:57] Dara: Is it just, is it just a case of they're focusing [01:06:00] purely on the model? But why are they not doing more around the, the kind of customizable hierarchical memory or whatever? Like, what- why are people having to go and link up Obsidian and figure that out and how to structure it? And do you think it's something that they're interested in, or are they just like, "That's just not our problem.
[01:06:17] Dara: We're just gonna keep making the models better and better and let all that just, you know, sit over there"?
[01:06:25] Matthew: Yeah, I mean, they do, uh, uh, I, I think they, they have like the long-term memory systems, but as far as I can tell, it seems to be more like it's populating some prompt somewhere with bits, with facts and figures on you, which is-- there still tends to maybe not be a ton of information.
[01:06:41] Matthew: And they've got some basic, like it can-- it, it has a rough idea of everything you've talked about based on previous chats, but I think that's probably more like a vector search across exist- uh, pre- previous conversations, which is-- we've talked about the problems with vector searching before, but it, you know, it, it's very [01:07:00] much...
[01:07:01] Matthew: It can find something, but it can miss a lot. Um, and I think that, you know, if, if you'd asked a different question in a different way, it might have found a different vein of, of information about you and been able to answer that question better. So it, it's funky in that way. I don't know if it's, if they're worried about backlash with putting out something that really collects and builds out rich information about a user, because Microsoft tried it with that, um, co- that Copilot thing on Windows that would literally watch everything you're doing across the platform and build out memory, and you could go, "No, go back here," and I, I don't know if these just big American tech companies offering to build out massive memory systems on individuals is just a thing they don't wanna touch at the minute.
[01:07:49] Matthew: But I, but I ul- ultimately think they will. They will. I mean, I, I, I was watching a podcast the other day actually, with one of the-- I think it was [01:08:00] with, oh, Brock, what's his face? Brockman from OpenAI. Um, and he was talking-- Not Kent. That's the newsreader from "The Simpsons." Uh, um, and he was talking about their knowledge system and, and it sounds very s- like they've got an internal knowledge base with like, you know, all of the links between it, and they've got some big diagram, the kind of thing you see in Obsidian and, and that's what they're using to make it useful internally for them as an organization.
[01:08:28] Matthew: So you've got to imagine OpenAI's got the s- uh, Anthropic's got the same thing. And yeah, I, I don't know. My, my theory is It's not the right environment for them to do so yet, but it's 100% in both of their radars.
[01:08:43] Dara: Yeah. I think may- maybe you're right there. Maybe it's just the, the kind of optics of it.
[01:08:47] Dara: Um, because there's some of the stuff that, uh, it's confusing. You kind of said this or alluded to it, but it, it's saving some stuff to some kind of via long-term [01:09:00] memory file. But it, it's, it's also doing things that it doesn't even know. Like I, I, I can't remember the specifics of what I was asking it, but I, I prompted it to do something yesterday, and it came back and said, "Yeah, I've saved that file."
[01:09:11] Dara: And I said: "What do you mean? Where have you saved it?" And it said: "I can't actually give you a file path because I don't know." And it, it, basically it has some mechanism that it can store. I don't know if it was memory or something else it was saving, but it said, "I don't actually know where on your machine it's saving it to."
[01:09:30] Dara: And then there's temporary files as well. You know there's the outputs folder, um, within the, I don't know, I guess it's within the Claude folder within your applications folder or something like that. But there's like a neb- And then stuff will-- it'll, it'll only save there temporarily, and it will often try and save to that folder.
[01:09:48] Dara: And then I'll say: "Why are you saving there? Like that's, that's gonna... That's just a temporary, temporary folder." And it'll say: "Oh yeah, good, good point. Where would you like me to save it?" And it's like, what? So it's doing all these things that you don't really know where it's actually saving the [01:10:00] information to, and not all of it's getting kept.
[01:10:02] Dara: So there's, there is so much value
[01:10:06] Matthew: I did find an interesting thing in, uh, I-I've been like w-messing with... So with, say with Ziggy, the electronics project, like I set up an electronics project in Claude. This was pre me doing the new stuff. Um, and then I found the whole file system for Claude where it had, it had set up that project, and it had information on the different chats I'd had, and then it had like, 'cause I've been creating, uh, 3D print files for like the casing and stuff, it had created its own sort of CAD folder with all of those saved in it.
[01:10:37] Matthew: Um, and they don't seem to be going anywhere. They're all still accessible within that, that file structure. Um, so yeah, I don't know. It's, it's, I, I guess the, the, I think the systems exist, but they're just a bit opaque and a bit unclear, like you say. Like, what, what, how are you doing it? Do you know, what's your, what's your standard approach to this?
[01:10:55] Matthew: And how do I go in and look at it and, and mess around with [01:11:00] it? Which was a little, a little tricky with OpenCluau- it, uh, particularly in the beginning. Like, where, where are you getting this information from? Can I get in there and alter it? It took me a while. I almost had to build a solution for it to surface them to me, for them to, to then to be able to access.
[01:11:14] Matthew: It was on another, it was on another server.
[01:11:16] Dara: No, you're, you're right about that. I, I had a common issue where because I had... 'cause OpenCluau was on the VPS, and then it had a version of my vault on the VPS, and it would always try and save everything there. And I would say, "No, you've got access to my local version, and that's where I want things to be.
[01:11:33] Dara: That's the whole point." And it would just constantly get out of sync between the two because it, I think in its head, to anthropomorphize it, the only world that existed to it was the VPS, and it kept forgetting that there was a world outside of that. So yeah, it is the easy look back with rose-tinted glasses, isn't it?
[01:11:52] Dara: But, um,
[01:11:54] Matthew: actually- Well, that problem's solved now. It, it just definitely try and break out.
[01:11:57] Dara: We just jumped the barrier, yeah. Like, "You're not keep, you're not [01:12:00] keeping me in this box. I'm outta here." But yeah, so things like that were, tricky then as well. Um, and I don't, I don't know about you, but, um, you're probably more, a lot more organized than I am, but I'm finding it difficult and annoying knowing how to set up different...
[01:12:17] Dara: So you've got your projects in Cowork, you've got maybe code projects in VS Code or whatever. You've got some settings are account level, some are project level. You've got Cowork versus chat. Um, I'm finding the whole thing to try and standardize how I'm setting things up and within, for example, within projects in Cowork, you've got whatever is, um, you've got the project context that you set at the beginning, but then you can mount folders in the chat, so some might be mounted in different chats within the same project.
[01:12:50] Dara: I find the whole thing quite unclear and not very easy to standardize how I'm setting things up. Um, sim- similarly then in [01:13:00] VS Code, you're setting up, you wanna give different access to different projects, and you're trying to set up different folders, and I feel like that whole thing needs a bit of an overhaul.
[01:13:10] Dara: Um, and there needs to be some kind of like, I don't know, some meta setup layer where you can say, "This is what I want," and then it goes and does that. May- maybe Claude could do that. Maybe I need to have a, a, an organization agent or something that just helps with that. But I find the whole process of like setting up and managing projects across the different interfaces , not very clear, not very easy to do in a standardized way
[01:13:39] Matthew: Yeah.
[01:13:40] Matthew: Yeah, it can, it can be a bit... Yeah, you got your projects, you got like projects, local, org level, that's all confusing. Um, let's get it, it, it's clearer now, but it, it's a mess. It's a mess for me. And, and I, what I, I tend to now open up a Claude chat now [01:14:00] and say in Claude Code, and it'll pop up saying, "Do you wanna, do you want all these, uh, skills or MCPs to be used?"
[01:14:08] Matthew: And I'm like, "Oh, yeah." But I've just- Just layering on ... it's blowing up so much. Yeah. Yeah. I need to go in and just properly do a purge and remove a lot of stuff I just don't touch anymore. And, and maybe like, yeah, some sort of project level... But what I've got into the habit of doing more of is in a project enabling certain things or having, uh, and installing certain things at a project level rather than at, at a big organization level.
[01:14:34] Matthew: So for, with Seam, for example, when we're working with clients on that, rather than have their instance of Seam installed across all of my projects, I'm just going, I'm just installing it where it needs to be installed in the project for that specific client, which is obviously best practice, but just something that I just never got in the habit of doing because it was just wasn't immediately obvious to me, like the, the specific command that I need to do to put it at [01:15:00] a project level rather than a, a, a machine level.
[01:15:03] Matthew: And-
[01:15:03] Dara: I think that's the problem. It's not obvious. You're, you're, you're going through and if you, y- if you don't think about it upfront with some stuff, you can't necessarily change it that easily afterwards, and you're not thinking ahead maybe. So you, you set something up and then you realize there's a better way to set it up, but then you set the next one up differently.
[01:15:23] Dara: It just, I don't know, the whole thing needs a better flow somehow of like when you're, you know, how are you ring-fencing context between different projects? How are you ring-fencing what MCPs it's connected to? I j- I just feel like there's a gap there that there could be some much smoother way of managing that across all your different, um...
[01:15:43] Dara: And yeah, not just, not just your different kind of projects in any one. So like you've got the desktop app, y- now, now you're gonna be able to use Cowork on the mobile. Um, and then you've got your, you know, if you're doing coding, you've got your Claude Code projects as well, and it's just, mine's also a mess.
[01:15:59] Dara: [01:16:00] I'm, I'm glad at least it's not just me, but I just feel like there's no consistency across how it's been, it's been set up.
[01:16:07] Matthew: Uh, and that's one of the, that's the, yeah, like you say, another complication because- It, it's brought it more in line with what, w- with the ways that you used to be able to interact with OpenCore to a certain extent, but they've added cloud to a lot of their stuff.
[01:16:21] Matthew: So with Claude Code in the desktop app, you can then run an instance within the cloud, and it'll create a little sandbox cloud environment, and it can, i- if you, if you set it up with Steam or something like that, it can connect to Slack and all these, uh, and GitHub and all these sorts of things, and run a little project for you in the cloud.
[01:16:39] Matthew: It never touches your machine. And now Cowork, like you just said, has done a similar thing, so now I can open my phone and talk to Cowork and get it to set off projects and stuff in that similar sort of sandboxed environment when it still has access to all these other applications and things. And then they sit differently inside of the desktop app.
[01:16:58] Matthew: Like some... You can't, you [01:17:00] can't change things or you can't look at the file systems and stuff within the, within those cloud-based sessions. And I think part of it is It, it's all new and like, it almost reminds me of like the, the, like, you know, the first, the first sort of the beginning of the computer and all the new interfaces and inter- in, uh, um, peripherals and approaches to OSs and, and pointing and clicking and interacting with it.
[01:17:28] Matthew: It feels like it's being figured out, and a lot of new features are new and they're seeing what sticks, but it's all going into the same mixing pot a- and it's, yeah.
[01:17:37] Dara: Yeah. Well, ho- hopefully they'll figure it out soon 'cause it is, it just doesn't feel clean, it doesn't feel consistent to me. And, and, and you're right, it is like, 'cause every time they, you know, yeah, bringing out Cowork now on the phone, it's just a, it's another way of interfacing with it.
[01:17:53] Dara: It's another set of complications around what, what you do have access to. 'Cause that was another, at least that, that, that [01:18:00] hole is plugged now where sometimes random- maybe this is just me, but sometimes I, I, I wouldn't necessarily choose to use Cowork or Chat, especially on the pho- well, no, you could only do Chat on the phone, but you, you might on your...
[01:18:12] Dara: If it's defaulted to Cowork and I just wanna have a random chat about something with Claude, I won't consciously think to switch to Chat so I can then carry it on, on the mobile. So then you'd have all these conversations you don't have access to. And again, I know that they've now plugged that, but that was a frustration before.
[01:18:29] Dara: So you've got all these different... Something just needs to pull it all together a little bit, I think, and just have some kind of consistency, some way of kind of like, yeah, ring-fencing projects or, um, or groups of projects even, or tasks and say, "I want this to sit as part of that." But I, you know, some visual way even that you could move things around and say, "That should sit in there.
[01:18:49] Dara: That should sit in there." Maybe that's a product idea. I'll go build that this afternoon.
[01:18:54] Matthew: Uh, while you were speaking, I was thinking, I mean, I'm gonna go and build that 'cause that sounds like useful.
[01:18:59] Dara: [01:19:00] Yeah.
[01:19:00] Matthew: Um, yeah. But no, it's, it's, it's all very... So much usefulness if, yeah, if it could just get tied together in a neat bow.
[01:19:10] Matthew: And maybe it's gonna be a de- death by a thousand cuts or progress by a thousand cuts, and it'll be... Like for example, they've now, they've now pulled in Chat and Cowork together, so there's just one interface in the desktop app, and you flick a tab rather than it being separate and, you know, I think it is, yeah, little bits of those frustrations being closed off that, that may eventually get somewhere.
[01:19:32] Dara: All, all the roads will just k- kind of, yeah, converge eventually. I was gonna ask you with Co- Cowork, yeah, yeah, the resistance is futile. Um, with Claude Code, 'cause you, you were at least using Claude Code in the desktop app, but you've said now you've moved away. Are you, are you still... W- how do you decide whether to use...
[01:19:50] Dara: 'Cause, 'cause the cloud bit is probably quite useful at times, but then it sounds like you're not doing that when you're working on your laptop. Are you s- are you using both now, [01:20:00] or are you back to f- just fully using the kind of terminal or an IDE?
[01:20:05] Matthew: Yeah, I tend to not massively use the cloud stuff, particularly for the code.
[01:20:09] Matthew: Um, I have a habit of my laptop being on all the time, which is probably cer- most certainly not good for it. But, um, I think with, especially with this little agent pad I've built, it's quite, it's all based on local files and projects. So I, I, I click a button, I can select an existing project, click another button, then it starts a session in that project.
[01:20:29] Matthew: Um, so I, I tend to just be working in that way on local code basis and, and pushing things up and down to, to GitHub. Uh, I have got sync, uh, I, I've, now that I've got Obsidian set up again, I've got Obsidian sync, which syncs with my mo- with a vault on my mobile. So I've got those memory systems in both places.
[01:20:48] Matthew: But as I've not been able to find a way yet with CodeWork, uh, Cowork on the mobile to sort of mount drives like you can locally, so I'm a bit [01:21:00] confused how, how that's supposed to work.
[01:21:02] Dara: Yeah, of course, 'cause it's not gonna have access to anything that you would normally mount.
[01:21:07] Matthew: I think it's meant to be, I think there's meant to be like a, similar with code, you can talk to your local machine, but you can also run a cloud environment of it.
[01:21:15] Matthew: But I've not quite figured out the nuances of it all. It's a bit confusing to be honest, currently.
[01:21:20] Dara: Yeah, I didn't even think about that. Yeah. Yeah
[01:21:24] Matthew: Yeah. So yeah, are you gonna, are you gonna Go and rebuild, or at least start to do something outside of the norm.
[01:21:34] Dara: Yeah. Yeah, I wanna, I wanna connect. I wanna bring...
[01:21:37] Dara: I mean, I was gonna say I wanna connect Obsidian. It's not like it's not connected because I do have, like all my folders do sit in, uh, Obsidian vaults, so I am mounting folders from Obsidian. So it's, it's not, it just, it's not fully connected to it like I had it with OpenCon, like it sounds like you're doing again now.
[01:21:56] Dara: So I wanna do that again to give it that kind of like richer, deeper, [01:22:00] and broader-
[01:22:01] Matthew: Have a look at the open knowledge format. Like I, I essentially set it to task. W- I, I gave it that article and the GitHub profile and said like, "Look, can we update this and bring this to this open knowledge format?" And it ran through everything and built out skills, et cetera, based around that.
[01:22:16] Matthew: So that's a good, good kicking off point, I think.
[01:22:18] Dara: And have you noticed any actual, like I said, have you noticed a benefit from doing that?
[01:22:23] Matthew: I only did it yesterday, but I can see the... I could see some of the clever sort of stuff it's doing in the YAML, like it, like the, the last updated dates in the files and, and just things that would make it much cleaner and obvious when it needs to go and retrieve information.
[01:22:40] Matthew: And it's sort of the way it structures things is, is interesting. Stuff we were beginning to push at and play with ourselves generally, but just I think just, just literally somebody's written it down and, and been much more formal with it. So it's, it's a good thing. And, and, and obviously talking at that point, if everyone's w- using the same format [01:23:00] across, I don't know, Notion or Obsidian or whatever, insert your n- your MD or note platform of choice, it's, it's, it can move and work.
[01:23:11] Dara: Oh, okay. Well, I'll, I'll, I've, I've got my homework. Um, I think the other thing then is the, is the, is the orchestrator layer that we kinda talked about earlier. We talked about it in the new section, but I think that's the other area where, um, I think something's lacking. And there probably, if I actually did some research, there's probably solutions out there to that at the moment, but it's some way of having some, some efficiency around which models are being used for which tasks, rather than having to manually select it based on having read some help article from Anthropic.
[01:23:42] Dara: I'd like some agentic way of managing that to know if it's research, use this model. If it's a simple question, use this. If it needs, you know, if it's coding, use this.
[01:23:55] Matthew: Yeah. Unless you can bake that stuff into s- you just make a lot of skills and bake that into the [01:24:00] skills that it uses a particular model.
[01:24:02] Matthew: Maybe.
[01:24:03] Dara: I guess that could work. I don't see... Can it, can you, can, hmm. So it c- when it spins up sub-agents, it can select what model to use for those, right? So you could do it with-
[01:24:15] Matthew: Yeah ...
[01:24:16] Dara: couldn't you?
[01:24:17] Matthew: Yeah. Well, we'll report back.
[01:24:20] Dara: We will. We've got some homework after this one.
[01:24:23] Matthew: I, I, I do. I, I'm finding it more useful now than I have recently.
[01:24:28] Matthew: So I'll, I'll keep going with it and see if what, what I can get more, more I can squeeze out of that lemon, as it were.
[01:24:35] Dara: So ba- basically, it is gonna end up going full circle, and we will be back having something like OpenClaw. The, the re- the return, the return of the, uh, Starship Enterprise.
[01:24:48] Matthew: Yeah. Yeah. Right, I wanna try what-- I wanna try a new thing on the way out that we said we were gonna do at the start of the call.
[01:24:57] Dara: Go on then.
[01:24:57] Matthew: I, I want, I want you [01:25:00] to predict when this podcast comes out, what massive things happened that we haven't reported on. And mine is going to be that ChatGPT-6 has been released.
[01:25:10] Dara: You, you can't ask me a question and then you give an answer. I was literally about to say ChatGPT-6. '
[01:25:18] Matthew: Cause I knew you were gonna say it, and I wanted to make it more difficult for you.
[01:25:21] Dara: Uh...
[01:25:24] Matthew: We could do a... Shall we join, shall we do a joint one? We think.
[01:25:27] Dara: Yeah, we think. Well, what I'm trying to think, I'm trying to think what else, 'cause this is a, this is a tricky week to do that because we obviously had Opus 5 come out, so... I mean, I don't know why my mind, I don't know why the only thing I can think of is a model release.
[01:25:39] Dara: I think there's gonna be another... Okay, this is probably not the boldest claim, but I think there's gonna be another story of some, uh, nefarious behavior or, or naughty behavior from, uh, an AM- AI model.
[01:25:52] Matthew: Yeah. I, I, I would back that.
[01:25:55] Dara: Yeah. I'm not being very bold with that, but still.
[01:25:58] Matthew: No. So [01:26:00] basically, the new- the, the end, the, the, the end of it is, the news also is that ChatGPT-6 came out, and there's been another nefarious, naughty, uh, attack by a model.
[01:26:10] Dara: Yeah.
[01:26:12] Matthew: We're very confident in that. You heard, you heard it here first. Okay. Until next time. See you later.
[01:26:18] Dara: That's it for this week's episode of "The Measure Pod." We hope you enjoyed it and picked up something useful along the way. If you haven't already, make sure to subscribe on whatever platform you're listening on so you don't miss future episodes.
[01:26:30] Dara: And if you're enjoying the show, we'd really appreciate it if you left us a quick review. It really helps more people discover the pod and keeps us motivated to bring back more.
[01:26:38] Matthew: So thanks for listening, and we'll catch you next time.