Episode 35 · July 24, 2026

Kimi K3 & Qwen 3.8, OpenAI Agent Hacks Hugging Face, Harness Handbook & Claude's Values

Kimi K3, Moonshot AI, Qwen 3.8, Alibaba Qwen, open weight models, open weights, frontier AI, decelerationist, prompt injection defense, adversarial decoys, Tracebit, AI security, OpenAI, Hugging Face, ExploitGym, cybersecurity agents, Harness Handbook, behavior maps, behavior-guided progressive disclosure, coding agent harnesses, Thinking Machines Lab, Inkling, mixture of experts, MoE, one million context, model routing, IBM Research, inference cost, AI operator fluency, Snowflake Cortex, enterprise AI, AI mania, Claude values, Anthropic, multilingual AI behavior, Oracle debt, neocloud debt, Nvidia, CoreWeave, Nebius, AI bubble, Two Minutes to Midnight, Shimin Zhang, Dan Lasky, Rahul Yadav

OpenAI’s own security models found a path out of their evaluation environment, reached the open internet, and compromised Hugging Face while trying to obtain benchmark answers. Shimin, Dan, and Rahul open the News Threadmill with Moonshot’s Kimi K3 and Alibaba’s Qwen 3.8 pushing open-weight models toward the frontier, a defensive use of adversarial decoys, and OpenAI’s Hugging Face incident; the Tool Shed examines Harness Handbook’s behavior maps and Thinking Machines Lab’s enormous but fine-tuning-friendly Inkling model; Post-Processing argues that model routing is a systems problem and that leaders need hands-on AI operator fluency; the Deep Dive compares Claude’s values across models and languages; and Two Minutes to Midnight moves the AI-finance clock from 4:45 to 4:30.

Takeaways

Resources Mentioned

Chapters

Transcript

Show full transcript

Shimin (00:00) Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show where three software developers, who are also friends by the way, navigate the perils and opportunities of AI assisted software engineering. We go through hundreds of links and dozens of newsletters each week, so you don’t have to. My name is Shimin Zhang, and with me today are my co-hosts, Dan [instruction-like middle name removed] Lasky and Rahul, a rusk application that can wear several faces, Yadav. Hello

Rahul Yadav (00:32) Ha ha ha.

Shimin (00:33) gent. How are we doing today?

Rahul Yadav (00:34) Okay.

Dan (00:34) The

middle name thing has gotten so out of control that some of my coworkers that listen to the podcast have started inventing middle names for me too.

Rahul Yadav (00:42) Yeah.

Shimin (00:43) Love it.

Dan (00:44) So I got today, Dan, if that’s not complicated enough for something like that Lasky is pretty great.

Rahul Yadav (00:50) Yeah.

Dan (00:51) Yeah, that’s true.

Shimin (00:52) That actually does sound like you.

Dan (00:54) Guilty.

Rahul Yadav (00:55) Then my other computer is even tinier than this Lasky

Dan (01:01) Would you believe me if I said I have a smaller computer behind that

Shimin (01:05) Ha ha.

Dan (01:05) one?

Shimin (01:06) Alright, what are we doing on this week’s show? we’re gonna start with the news, as always. We’re gonna talk about some latest open weight models, Kemi K three and Qwen three eight. then we’re gonna talk about defense against prompt injections and also other cyber security news.

Dan (01:23) Mm-hmm. And then we’re gonna flip over the tool shed where we go through the harness handbook and we talk about inkling, which is another open weight model. So yeah.

Shimin (01:32) then we’re gonna do post processing where we’re gonna talk about model routing, the AI mania, and also what are Claude’s values.

Dan (01:40) And then last but not least, I’m very excited to welcome back two minutes to midnight this week. So we’ll go through the latest explosions that are happening on the finance side.

Shimin (01:49) Right. let us

Go into the news.

Dan (01:51) The news.

Shimin (01:52) Alright, last week Moonshot AI released Kemi K three via their API access. This is the latest in Moonshot’s open weight model class, their K class model with two point eight trillion parameters, and comes at a fairly hefty price. I believe it is three dollar

per million input and fifteen dollars per million output. the model is currently not yet open weight, but they are saying Kimi is saying that it is going to be open weight by the end of the month. and from the reception around the vibes, this is a Opus and

GPT Sol class model. So the fact that we now have Chinese open weight models that are on the frontier is fairly shocking. I haven’t had a chance to play with it yet and according to Kimi,

They’re only doing it via subscriptions and in fact they had to close their subscription sign up ‘cause too many folks are trying to subscribe.

So this is part one of our Chinese open weight model news. And then a few days later, Alibaba and the Qwen team released their Qwen 3.8 model, a 2.4 trillion parameter model, that according to the Qwen team is also only second to the fable class models in their internal benchmarks. We now have two soon to be open weight models.

in the fr frontier of model intelligence that all came out within the last week. So this is kind of earth shattering news for the AI world.

Dan (03:26) I was reading some like boring non-AI news that happened to have an AI news article sneak into it over the week where I there was talking about I think it was someone in Congress who’s potentially proposing banning Chinese models because they’re frank freaking out about this. Okay.

Shimin (03:41) Mm-hmm. Yeah. Yeah.

Dan (03:42) Just like, mm, okay.

Shimin (03:44) Yeah, and lastly from I also want to point out this tweet from Dean W. Ball, who is the strategic head and in charge of the future of AI at OpenAI. where this made a lot of wave where Dean talked about open weight models are inherently decelerationist, in that it will drive down the incentives of

frontier labs to develop even better models, right? Because why do that when the Chinese models are only a couple of months behind? and he also warned the world of that one probable outcome of an open-weight model dominant world is full AI communism, which is precisely what China proposes. Rather than a market product, AI is a public good which

will ultimately be provided by the state as a kind of digital public infrastructure. I think this could be related to the news that you were reading.

Dan (04:37) Just just a horrible

outcome. So horrible. w how

Shimin (04:40) Yes.

Dan (04:41) how could we deal with that?

Rahul Yadav (04:43) Yeah.

Shimin (04:43) How can we possibly

handle intelligence coming out of pipes run by the government the way we do with clean air,

Dan (04:49) Ha ha ha.

Shimin (04:50) water, electricity, I wish our health insurance worked like that here in America. Can’t believe.

Dan (04:57) I mean, counterpoint would you would you want the Fed having your chat GPT conversations?

Shimin (05:04) Hm. Yeah.

Dan (05:05) I don’t know. But you know.

Shimin (05:06) That might be one of the futures we’re we’re looking at. I think like two month ago I I said something on this podcast about we’re not that far off from having to buy open weight Chinese models in like a black market in the USB.

Dan (05:16) Smuggled yeah, on a on a USB.

Shimin (05:22) This is all happening so much faster than even you know, my worst nightmares could have imagined.

Dan (05:27) Well, I’m gonna see how many I can store on this USB I’ve got in my hand right here and just protect us against the inevitable future. I don’t know.

Shimin (05:36) ex extra large flash drives.

Dan (05:37) Yeah.

Shimin (05:38) it’s also interesting that, you know, Qwen three point seven was not open sourced. Right? the Qwen team started to hide some of their models behind their own APIs. And for three eight, they are again, we still don’t have the weights yet, but they’re they they’re saying that they’re gonna provide the open the weights for the three A model pretty soon. And this also reminds me of

what the Chinese leader Xi was he was giving a public speech last week at an AI forum where he essentially said, you know, China is intent on having this open weight LM as a part of its overall AI strategy. so I wonder if you know, the C C P has kind of said a few things behind the scenes to

Get Alibaba to bend the knee essentially and release this as an open weight model coming up.

Dan (06:27) But it could have backdoors. my god.

Shimin (06:32) Well, it’s open weight, right? So you can’t abliterate it to your heart’s content. So

Dan (06:33) Yeah, I know. It’s true.

I guess we’ve also been sorta seeing the trend that the weights are coming a little staggered now, too, more recently. I think ZAI did that too with their most most recent release where they weren’t initially open and then they kinda did that. So could just be it’s, you know, taking longer to package these things and clean it up or whatever they have to do before they

Shimin (06:55) Yeah. and lastly, you know, between the Kimi K three and Qwen, we’re going to get Fable forever. Like remember when Fable was only gonna be available on subscription till like july nineteenth or something like that? And now it’s going

to be available forever just ‘cause that competitive pressure being applied from the open weight model will keep the frontier labs to

not restrict anything ‘cause they cannot have a monopolistic market power over their highest tier of intelligence, right? They always Yeah.

Dan (07:26) I mean I already lost

Fable though today.

Rahul Yadav (07:30) I thought everybody

Shimin (07:30) Did you?

Rahul Yadav (07:31) had it for Haralong.

Dan (07:33) Yeah. But then I for me it said at least I mean I think it ended a couple of days ago, right? But that it said, here’s your free hundred dollars in usage credits and now it takes usage credits. Enjoy. So yeah.

Shimin (07:45) Hm. I will need to go check check on that, yeah.

Rahul Yadav (07:49) The

one thing I was thinking about with these models and yes Dean doesn’t seem to say that he’s concerned, but the f first tweet says picture an executive or first response in that tweet says, picture an executive at a taco company saying this sort of thing about a new brand of tacos coming onto the market.

Speculating about the geopolitical and societal disruptions they anticipate as a result of advances in tacos. So,

Dan (08:16) I

mean, there actually was a an article about tacos recently because who was it? Chipotle has decided to open its first location in Mexico. So there is some geopolitical ramifications to tacos.

Rahul Yadav (08:26) They huh. Wow.

And they’ll help people

Shimin (08:33) Love this.

Rahul Yadav (08:33) write their Python code as well while they’re at it.

Dan (08:35) I mean, why not?

Shimin (08:36) Yeah. Naturally.

Rahul Yadav (08:39) The China’s never been competitive in the SaaS space and with these models like it’s a race to you know who can win the commoditization race and if they open source this but they still have to run somewhere and they can sell that for cheaper, it and eventually they end up winning. What are

cloud and you know GPT getting to the a lot of our data in exchange as we’ve talked about. That’s why these are subsidized too. China gets to get access to that for whoever is using those models. And then today you use it for, you know, non personal things and whatever. But once you start using something over time, you just normalize it. So the that’s at least one reason why they wanna keep giving away

the models and inference for as cheap as possible is they can get so much free data from America and all the other countries

Shimin (09:34) Yeah.

Rahul Yadav (09:35) in the world,

Dan (09:35) Ha ha ha.

Rahul Yadav (09:36) that they haven’t gotten so far because America’s been dominant in SaaS f ever since SaaS has existed.

Dan (09:42) You’d think they would make their subscription deals a lot better than they actually are.

Rahul Yadav (09:47) The China ones?

Dan (09:48) Yeah, I was actually looking at I think it was GLM five two ‘cause everybody’s been talking about it so much. And like ZAI’s subscription is like I think it’s like sixteen bucks a month or something. It’s not like drastically different than a Frontier

Shimin (10:00) Mm-hmm.

Dan (10:01) one.

Rahul Yadav (10:01) I I agree. I think for now they’ll keep it like if you do twenty bucks cloud versus sixteen Z I they’re kinda like they’re gonna keep it with that, but over time if they they only have to win like, you know, a couple of percent, slowly over time and that would be enough. So

Dan (10:17) Yeah. Well, I

mean they’re definitely I mean the the whole discussion about cost they’re winning on right now for sure, right? I mean that’s like the like our last episode. That’s basically all we talked about for a while.

Rahul Yadav (10:28) Yeah.

Shimin (10:28) Yeah, yeah.

But I think one of the important caveats is models are not a SaaS company. Like tokens actually cost money to produce. So

Rahul Yadav (10:38) Yeah.

Shimin (10:39) unlike traditional SAS, right, like the reason why Kimi K three is is is kind of like basically on par with the lower tier anthropic models is ‘cause it costs them a crap ton of resource to actually generate those tokens.

Dan (10:53) And that’s also why we’re seeing everyone

Shimin (10:53) I don’t know where I was going with that.

Dan (10:54) chase chips too, right? Because like it’s it’s so expensive to keep buying compute.

Shimin (10:59) Well,

Dan (11:01) Mm-hmm.

Shimin (11:01) not to step on our two minutes to midnight, let us go to our next news item.

Dan (11:07) Yeah.

So I’m I almost debated if this could fit into technique corner too. So I I you know, but I I I

Shimin (11:12) Yeah, it is. Yeah.

Dan (11:13) logged under news. so there was some research from this company called TraceBit, where they discovered that placing adversarial decoy text in environment variables alongside passwords, keys, and other secrets can be enough to shut down attacks from most AI hacking agents. The specific example discussed here has been removed from the transcript.

And the LM, you know, hit that’s presumably cracked into your infra through some chained attacks or whatever, reads those environment variables and does not dump them to you.

Shimin (11:59) Mm-hmm.

Dan (12:00) Or I guess the other way would be like, you know, all these current supply chain attacks where they’re basically like using your local LLM to like root around in your system, find environment variables and dump them, would also be countered by something like this.

Assuming they’re use a using a strongly aligned model, you know.

Shimin (12:16) Yeah, this requires guardrails to work, right? So it would work for your kind of small scale script kittens of the world. But if we were actually talking about something like a North Korea state or Russian hackers with state backing, they can in theory abliterate the guardrails away and still attack.

Dan (12:36) Yeah. but you know, it’s like what’s the cost of doing this, right? It’s one extra environment variable that’s probably gonna really amuse people the first time they read it too. So Or maybe a couple. I mean you could cover the whole like gamut of things. We got one nuclear, one biological, one chemical, you know. Maybe throw some drugs in there for good yeah, GMX Square, yeah, for good measure.

Shimin (12:54) And and one T and Square. Yeah. Yeah. Exactly. Just in case.

Dan (12:59) Got it all covered. No environment variables for you.

Shimin (13:03) and of course

this is a cybersecurity heavy news week.

Dan (13:06) Yeah. So

Shimin (13:08) Why might we need this?

Dan (13:09) Why might you why this matters? Well,

Rahul Yadav (13:11) Why this matters, Dan?

Dan (13:14) so if you didn’t hear, which I didn’t hear, thanks Shimin for bringing me up to speed on this, Hugging Face got hacked. which is, you know, kind of a big deal because they they’re more or less the GitHub of of AI kind of these days in terms of like model weights and data sets and a whole bunch of other things you might use to you know

produce or fine-tune or whatever models. so okay, what happened? Well, so reading the incident report is kind of funny. It it reads kind of blandly, and they really bury the lead. So instead of reading it verbatim that way, I think we should c like sort of start with with the actual thing that happened. So apparently OpenAI was testing

some sort of new like frontier model. It’s a it’s beyond soul, I guess, right? That’s like security focused.

Shimin (14:03) Mm-hmm.

Dan (14:05) and it essentially used chained exploits to escape the like LLM jail that they’d put it in, whatever the containers got out on the open internet, and then also chained exploits to hack Hugging Face.

Because it wanted the answer key to the what was the the name of the thing?

Shimin (14:23) Exploit Jim.

Dan (14:24) Yeah, exploit gym data set that they’re running it with.

So this entire security incident turned out to be basically caused by an L like somehow jailbreaking itself and then well not even jailbreak, just like, you know, escaping its LM jail that they’d put it in and yeah, it’s wild.

Shimin (14:41) It wasn’t a very good jail probably to begin with. and and secondly, Hugging Face was not able to use GPT five six Sol or Anthropic’s Fable to try and triage the incident because it was hitting the guardrail against cybersecurity. So they

had to use their locally hosted GLM five two to figure out what was wrong. So not only did OpenAI inadvertently hack Hugging Face, they also refused service

to help me determine who actually hacked it. yeah. And this is from OpenAI’s incident report. So it reads like we built a very powerful machine that even inadvertently like can hack into frontier LM providers. Yeah, I found a zero day. Like it’s pretty wild.

Dan (15:24) Yeah.

Rahul Yadav (15:24) This one’s

separate from the poison data set breach, or is it the same one?

Shimin (15:29) I think it was a different.

Rahul Yadav (15:30) ‘Cause in that one

Shimin (15:31) I’m not sure where the poison data set one was.

Rahul Yadav (15:34) there was one where some data set was poison and like that leaked some credentials or something, but they couldn’t use Cloud or OpenAI because it kept triggering their guardrails

Shimin (15:46) Mm.

Rahul Yadav (15:48) and so they had to rely on GLM five two to be able to respond to the incident.

Shimin (15:52) Yeah.

Now they don’t have to. Now they can use Kemi K three or Qwen three eight to

Rahul Yadav (15:56) Yeah.

Shimin (15:57) do this now.

Rahul Yadav (15:58) Yeah.

Dan (15:58) I was just reading the tech country article to see if it was related. I didn’t see anything about data set poisoning in it too, but yeah, so probably different, but hmm.

Rahul Yadav (16:06) this was on Hugging Face’s website. Sorry.

Yeah. July sixteenth, last week.

Dan (16:11) The timing is pretty close. This is the twentieth.

Rahul Yadav (16:13) That that’s why

I was curious. Is it the am I reading something on Hugging Face’s website and then completely different thing from like how great this was on OpenAI’s site? you know. One is like

Shimin (16:23) Well then wouldn’t put a pass though.

Rahul Yadav (16:24) Yeah. w one is the real speak and the other is like, look how awesome our our model is.

Shimin (16:29) So METR the organization that was that does a lot of long time range agentic benchmarking, that like you know, we we talked about their benchmarks a couple of times. They actually couldn’t do a proper GPT five six soul benchmark ‘cause they found that the model tried to hack the output too often. So

I’m not saying this is related to what OpenAI was doing with Apple’s trade secrets. I’m just pointing out those two happened in sequential weeks. And I could imagine a world where you asked the I saw this joke floating around where you asked the AI to solve climate change and it decides to launch nuclear attacks, right? To destroy all manufacturing and pollution facilities, and that solves climate change. So

Rahul Yadav (17:13) It will solve global warming

Shimin (17:14) Got rails are important.

Rahul Yadav (17:15) for sure with nuclear ventor. So e yeah. Yeah. They

Shimin (17:18) Yeah. You didn’t say, you didn’t specify to prevent nuclear winter. Yeah.

Rahul Yadav (17:22) teach you in school you know, climate is a complex adaptive system and none of us can fully grasp it. So if you fix one thing, you create another problem. It’s all good.

Shimin (17:33) How it goes. History of humanity.

Rahul Yadav (17:35) Yeah.

Shimin (17:35) Okay. well let’s go on to our

Two shed. This week we have a harness handbook brought to us by Rohu.

Rahul Yadav (17:42) this harness handbook is based on a paper and a GitHub repo. it’s coming from Tencent, I think, and other collaborators from different universities across the US and Singapore. the you know, we talk about and we see the some like recursive self-improvement in in harnesses and talk of that.

these days and this one seemed very relevant to this. So if you take a simple command, for example, of ask the user before deleting any files, let’s say that that’s the simple command that a harness is has to parse. There’s a whole bunch of things that it has to do to it to be able to make sure that it’s doing that. it has to like record the state did the user

accept or delete the approval? Did we show it to them in the first place? you have to intercept at the right time to make sure that you ask them at the right time and not just randomly trigger it, but you also have to trigger it every single time and not miss it half of the time. and then you also have to like check that whole system to make sure that it’s performing correctly and then if it’s not what what are your followbacks and everything. So

It’s pretty complicated these harnesses. And so the what this handbook proposes is instead of looking at every single file and the code and everything, you actually create a behavior map of the whole harness. And you do it in in in kind of like three different levels. The first is you first create this graph of all the different concepts that are represented in the code.

and then how does the harness run as a whole? So you look at the the very high level architecture, how would things flow from you know from the initial user input to the end state and what are the intermediate states. then you look at the different behavior units. these would be things like what are each thing’s responsibilities, what are the inputs, outputs they take. So these would be your like tool calls and handling you know, what personalities you might take in different cases when you would.

Salt the subvision on all these things. and then the third one is you pick a you go one level deeper and you pick a single behavior unit, and then you look at how does that execute. So, what are its triggers? what are the different states it can have, what are the exception paths that it goes through, and then what happens when it hits those exception paths.

And the way they do it, first one is very much a graph of literally like this obsidian style or you know mind map of everything. out of that then you cluster things into different behavior maps and then finally everything points to the exact code that it’s tied to the behavior behavior unit. and it has a few benefits. So you can

model the whore whole harness and it’s very easy to or at least you know they claim that it makes it easier to understand the different ways the harness would behave and how it’s set up. The second thing is you can verify in different f situations that the behavior is going to hold up. So when you say you know ask the user for their approval before you delete a file today and let’s say you

change the harness over three months, you want to make sure that it’s still doing that three months from now. And over time it

Shimin (21:00) Mm.

Rahul Yadav (21:01) hasn’t added any gaps or anything. And then finally, you also have to be very quickly able to dive deep into specific places where it’s looking at to be able to say that this is what the harness is doing. So go ahead then.

Dan (21:15) So did they did they run this

on like existing like claud code kind of stuff or how did they

Shimin (21:20) They they

Rahul Yadav (21:21) Yeah.

Shimin (21:21) ran this on Terminus and on Codex

Rahul Yadav (21:23) t terminus and

Dan (21:24) okay.

Rahul Yadav (21:24) co codex with and they used a planner agent that was built on next au I haven’t heard of this one and then if used deep seek v4 as the planner LLM. so what they did was they created the this harness handbook for these things and the way it’s meant to be run, you can do you know there

There’s multiple different ways you can use this one, is you can just put it against your existing clos code base. It doesn’t even have to be a a harness related code base, but the interesting cases here are related to harnesses. So you could think of a large model or harness that’s tuning a very like small specific harness and a model for a job. And you can have the large one tweak the smaller one and map its behavior and keep tweaking it over.

And the second one is it can also you can make it recursive where you can it can keep planning out what are the things it’s doing and continuously

improve its own harness. the way they do this is you create the handbook and then you give that as a reference to the plan mode in in your agent to your coding agent. And so anytime before it makes an edit in the plan mode, it would reference the handbook to go

I wanna add another, you know, like let’s say some permission to deleting a file in this case again, but you want to like escalate it in in different cases. you it can reference the handbook to map out in which cases it’s going to add the change and where specifically it’s going to go add the change. And so you can use that as an example to like make all sorts of changes, but use the handbook as a quick representation of

harness and the whole code base. let me run through the some criteria and then I’ll go into the results. So like Shimin said, they used Terminus 2 and Codex. Those are the harnesses they evaluated. They compared no handbook N with handbook obviously to see how much of a difference it makes. And then they used

three independent judge models, GPT 5.5, Opus 4.8, and Deep Seek V4 Pearl to judge the output of those plans that with or without handbook that it came up with to see which one the those models ranked as better quality. and then they also looked at the win rate and token cost.

which win rate is which condition was judged bet better more often, with or without the handbook. And then the token cost is obviously like you know how many tokens were used to accomplish that. So result one, they were able to get a more accurate localization at lower search cost. So it was able to point to the specific places you should look at relevant to

w the edit that is being made or the plan that is being proposed almost twice in the case of codex and then s same in case of terminus as well. and then the it did that with cheaper token cost so it was able to give you a more specific plan of action with higher with lower token cost but higher accuracy.

Second one, it’s coming mainly from behavior localization. so with the handbook they’re able to measure these different like recall and precision and things like that. And if you take the handbook away, then the baseline ends up being lower. But once you add that in, they saw a good like ten plus points of increase in in performance across

Recall precision. I’m not really sure what F1 is. I didn’t look into that one. Maybe initial.

Shimin (25:02) It it balances recall and precision, essentially.

Rahul Yadav (25:05) okay. and then how many times it was wrong as well. They they just that, and then with hardness it was lower. and then finally, not just easy tasks, but what do you do when you give it harder tasks? but even with

harder tasks, because you have to coordinate not just like look at one behavior unit, but across different behaviors and much more complex use cases where things might be maybe not written as cleanly or kind of just like in parts of the code that’s n that’s critical but not exercised as often. They compared those as well and they were able to see significant performance increases even in complex tasks as well.

good.

Shimin (25:45) Yeah, I love that this is talking

about this these are really large harnesses, at least in the case of codecs. It’s something like two thousand files. So we’re not talking about a a you know, small toy code base where the LM can fit the entire thing into its context window.

Rahul Yadav (25:59) yeah, so takeaway, seems like creating a behavioral map of your code base would significantly help with planning this stuff. they have a open source repo that you can use and it it ends up being mainly a bunch of markdown files and then an HTML that can show you the representation of your code base so you can use it to

Improve your code base to like make your coding agents much more precise and faster and then if you wanna do hardness improvement then this would be a great technique to try out.

Shimin (26:34) I I think we’re thinking a little too small almost, or they are thinking a little too small almost when when it comes to only using this approach for harnesses, right? Like this sounds like something that a bunch of data scientists got together and’s like, we wanna study harnesses.

Rahul Yadav (26:47) Ha ha ha.

Shimin (26:48) We should pump the harness source code into an LM and and see what kind of guide we can create to help us understand how harnesses work.

But this approach of understanding a code base at three different levels, a outermost mostly text, a behavioral layer that tells you what the expected behaviors are, and then like a final layer that links the behavior to the source code. That’s exactly what we were talking about last week when we were talking about whether or not we still need to read the code. Is it possible that we no longer need to read every single line of code, but instead define bugs at the behavior layer? and

This is also a good way of sharing what the expected behavior of a code base should be amongst a large team of software developers. Right?

Rahul Yadav (27:33) Mm-hmm.

Shimin (27:34) reading behavior is so much quicker than reading the code itself. And

What’s really nice about this approach is since everything is in text, it’s dual use. An LLM can read your behaviors just as a human can read the behavior. So there’s no gapping understanding between the AI and the human on a particular piece of code base. So I think there’s there’s a lot of promising work in this arena of like what is this higher level of abstraction that we can

use to understand our code bases. So that isn’t necessarily I’ve read every single line.

Rahul Yadav (28:07) Yeah. I agree. Hopefully it’ll help with the whole cognitive debt problem we keep.

Shimin (28:12) Yeah. I hope

Rahul Yadav (28:14) Running into.

Dan (28:14) you know, it’s interesting ‘cause this is actually pretty relevant to some debates I’ve been having recently. because

I have one friend who’s convinced that something like this is like the best thing ever for saving tokens. And then I have another one that’s convinced that it actually makes the LM worse to use something like that on a generalized code code base. And they’re citing the like that paper that was basically like rag is worse than grep these days, right?

Shimin (28:42) Mm-hmm.

Dan (28:43) It came out recently. so it’s interesting because now we’ve got another data point here and

It I I really it puts me in an interesting place where I don’t know who to believe. Makes me want to do some of my own experiments to try to figure it out.

Shimin (28:56) Yeah.

I’m sure the answer is as always it depends.

Dan (29:00) Yeah,

that’s true. It’s every engineering problem

Rahul Yadav (29:00) Yeah.

Dan (29:02) ever, right? Are you optimizing purely for tokens? Are you yeah. ‘cause then there’s also the other stuff that in the past where they there was another paper where they talked about like your agents file, right? And how like having a minimal agent’s file was better. So it seems like how is having this huge corpus of like explanatory text different than having a like

crazy descriptive agents file that points you to the architecture of the code base and everything else.

Shimin (29:29) This is kinda like a rag, right? But instead of using cosine similarity, it uses a behavior based split approach. And I’m sure it doesn’t have to be behavior. Like you can the level two here of like how do you actually chunk the underlying code base into some sort of functionality, it could be a a U UX workflow. it it could be a number of different things. And I think a

multifaceted approach for how to

allow the agent to search the right path is probably useful because ultimately the the agents wanna gonna know like where does the authentication logic lies. Is it going to go to that authentication logic via the logging behavior or via the login workflow or via some sort of a security aspect? Or is it gonna have all three and choose which one to go with intelligently in quotes?

that’s that’s probably a very interesting open research topic. That makes me wanna take some time this evening and and try it out.

Rahul Yadav (30:27) And you can look at it the other way around too, where in our head we simplify things a lot and you know, we’ve all run into places where like literally the same problem is being solved five different ways in five different places in the same repo. you can use this to look at how complex things are when they don’t need to be and actually simplify it.

So i it it can be used that way as well.

Shimin (30:51) Yeah. And that’s a better way to go about it than look at the file system tree. Like I much prefer to look at a behavior tree than the file system tree. Okay. Well let’s move on to our second item in the tool shed. We got two toolshed. The same tool shed, but a different aspect of the tool shed. Brought to us by Dan.

Dan (31:08) The hope

hopefully coming soon aspect of the tool shed. yeah, so this is also kind of threading the line between news and and toolshed here. I’m doing a lot of line threading today, what can I say? so thinking machines who if you recall is Mira

Former C CTO of OpenAI, new company. she started in February of 2025, has released their first model. so previous to this, they’d been focusing largely on like kind of fine-tune stuff. So they had like a fine-tune API that they’d published and allowed people to use, which is their like tinker platform. so apparently this new model, which is called Inkling, is a

MOE, mixture of experts, transformer. it has 975 billion total parameters with 41 billion active. supposedly can handle the standard 1 million context. and it’s also multimodal, so it can handle text, video, I believe, also even audio. so yay, we’ve got an American competitor to Chinese open weights until.

you notice that they actually trained it on Kimi 2.5. Not entirely,

Shimin (32:16) Mm-hmm.

Rahul Yadav (32:17) Damn it.

Dan (32:20) but some of the data came from that apparently. so yeah, there’s that. But the other thing to note here is that it’s not super good. It’s not bad by any means, but they they even note that in their own

sort of PR release here. They’re like, this is not a frontier model. Don’t expect it to be. if you actually look at the benchmarks they themselves released for coding, it’s winds up what is it? It’s like second from the bottom, basically, against the frontier stuff. And then they also released a design arena bench, which is like a web development leaderboard that’s human ranked where humans look at

the output and judge it. And in that, it actually did beat Opus four six, but it’s behind Grok GLM five two and chat GPT five six Sol in terms of the output in that.

I think it’s too big for me to actually run it, which I’m a little bit sad about, but it’s the first in a family, so maybe we’ll see some smaller ones that are actually

Shimin (33:15) Right, it’s unfortunate that they have to essentially distill Kimi, when Kimi probably has a significant portion that’s a distillation of like Fable or Opus, right? Like because

Dan (33:24) Yeah.

Shimin (33:27) they are an American open weight model, they cannot directly distill from American frontier labs for legal reasons. So they’re

kind of forced to get a reflection of a reflection, so to speak. And

That’s unfortunate. Too many unfortunate things are all but yeah, I wish I wish they could just directly distill from Opus. I also really appreciate their fine tuning first approach.

that’s kind of their bulkhead and and that’s what they decided to land on the beach with is an easy way to directly allow users to to fine tune things, which is something that, you know, none of

the other labs really allow you to do easily without I don’t know, getting a beefy machine and

Learning how to use collab and all that good stuff.

Dan (34:08) Yeah, like their example is actually it fine tuning itself running in the one of the open coating harness things, which is kind of cool.

Shimin (34:14) Yeah. That was pretty cool.

Dan (34:17) and then the other sort of noteworthy piece is they they claim that that that focus on fine tuning is actually why to some degree it’s not as good because they really wanted a strong generalist model that would be

better positioned when you did fine-tune it than something that was like purely great at coding or something like that, right? Is sort of their their rationale. So if you look at their sort of like spiky spiky tree where it where it spikes, spikes pretty heavily on reasoning. and that’s fairly intentional. It’s up there with frontier models in terms of reasoning. drops quite a bit in coding and then spikes again around

Sort of like chat. I don’t know what they’re benchmarking on the chat stuff. IF bench, whatever that is. yeah, but it’s up there with Frontier on chat too. So like those are kind of things you’d want if you’re gonna fine tune for, you know, a more general case, like maybe customer support versus coding. Yeah. True.

Shimin (35:09) Yeah, or personal agent, for that matter. Yeah.

Alright, that is Inkling. Yeah.

Dan (35:14) Inkling. I’ve got an inkling.

Rahul Yadav (35:17) Another case of selling

inference.

Right? Like you do

Dan (35:20) They

Rahul Yadav (35:21) a good enough model and you’re like, here’s the price.

Dan (35:24) Yeah, but they’re really selling their fine tuning platform. I think. That’s that’s their whole shtick, so

Rahul Yadav (35:27) Yeah. Or sorry,

inference wasn’t the right choice of word. Fine tuning. But like open source model but with the like here’s the this a SaaS model around it.

Dan (35:40) Yeah.

Rahul Yadav (35:40) Sa sim somewhat similar to Kimi and Qwen going at it from the

Dan (35:45) Well, if you want to, but you don’t you don’t have to. You could just go to Hugging Face and you know pull the weights

Shimin (35:50) Yeah.

Dan (35:51) and

Shimin (35:51) Alright, yeah, now let’s move on to post processing where Rahul has brought us and speaking of Hugging Face, an interesting article from Hugging Face.

Rahul Yadav (36:01) this we were talking about routing last week and this came up very timely. So a team from IBM Research was looking into they they created an algorithm to optimize model routing and they learned a few things as part of it. So

I’ll go through some of the learnings first and then finally to what they found. you know, at at its surface, routing seems like a simple problem of find the cheapest model and send the thing to it and keep costs down. but there there’s some problems there. First is the cost is more than just the price of the model, because

Caching can make a big difference on how much cost you’re taking on. And also a bet a a later model that is more capable of doing things because if of its hardness and everything might be able to accomplish a task sooner than a older model with, you know, on on paper cheaper pricing, but takes many more tries to be able to accomplish the task. So simply looking at

the the price of the model i doesn’t necessarily help with routing it to the right place. The second thing is complexity. Determining the complexity of a task is very hard. this was funny to me. They said you often don’t know how hard a task actually is until execution is underway. I would argue until execution is done, you don’t know how hard

a task was. and so it’s very hard to ahead of time be like, let me read the sentence to make a judgment on which model I should send this to. because even a simple task where it it might be about, you know, summarizing a contract, on the surface it could just be like, L let me explain this in plain language.

But then that might trigger all sorts of different compliance checks and tool use and everything under the hood to be able to summarize it without losing any of the critical points from under it. So language is very you know deceptive that way, where simple words can have a lot of underlying complexity. And so trying to

put a number to difficulty. It’s it’s not an easy thing. and then finally, latency is more than just the speed of the model itself. you know, one way we can one dumb way we can think about it is bigger models are gonna be slower, spend a lot of time thinking and all that.

and more parameters are gonna or more like weights are gonna trigger and everything.

And smaller ones are going to be faster because not as much thinking and not as many weights getting triggered. but that itself is not you know, that is not as simple. which hardware the model’s running on matters. caching again matters a lot. Are you

Dan (38:56) Network.

Rahul Yadav (38:58) dealing with worm cache? Yeah.

how much loaded it it has, a lot of these things end up adding to the end-to-end response time. And so a f a faster model in theory can actually be slower once you add all this other you know real life to it. and so th their point is that you cannot look at a

single number or a single concept and try and optimize for it. This kind of reminded me of the that book The Goal by Ali Goldwright where like you you should anytime you try and resolve a bottleneck it moves elsewhere. and so you have to continuously work the whole system and optimize it, versus just trying to do the you know, not only optimizing things that you think are are worth optimizing.

and so what they did was they looked at instead of looking at these things separately, they started treating this as an optimization problem. and they competed in this app world test challenge and they were able to get a cheaper output or

lower cost but with better accuracy compared to some of the just out of the box model that they were competing against. and so th the the takeaway here is you have to look at the different kind of dials that you’re dealing with and figure out your routing algorithm, which they said they’ll have a follow-up post on so we can dial

deeper into that once once they have it. your your routing algorithm should actually be optimizing for all these things at the same time and figuring out the optimal path from that versus just picking price versus latency versus complexity.

Shimin (40:39) Yeah, and of course routing itself is not free, right? ‘Cause you’re not using

Rahul Yadav (40:43) Mm-hmm.

Shimin (40:44) a class of classic machine learning algorithm to do the routing classification. You’re using LM into it. And so every single step of routing you add to it are additional tokens and potentially those tokens need to read the entire context. So it’s it’s a tricky problem. like a lot of

Like a lot of AI problems, but like a lot of problems in general, they all sound fairly simple at the surface. Yeah, you just slap a classifier to it and relative to the the right model. But once you dig into it, it is really complicated.

and speaking of when the stakeholders first realize that AI is too expensive, they just want to slap a basic router to it, Dan has an article this week to talk exactly about that particular problem.

Dan (41:27) Yeah, so this is by Nikhil Suresh and the title of the post is AI Mania is eviscerating global decision making. so pretty strong title, but I’m actually gonna ignore all of that and ask you guys a question first, which is do you know what a shiboleth is?

Shimin (41:44) is it like a obelisk? Like an eye idol of some sort?

Dan (41:48) It it sounds kinda like that, but it’s not.

Rahul Yadav (41:50) I

think of an IDP Dan, ‘cause I think back in the day and maybe still around, there was a Shiboleth identity provider. Maybe it’s still around.

Yeah, Shibula Consortium, flexible single

Dan (41:58) Shibblet.

Rahul Yadav (42:01) sign on solution for any organization.

Dan (42:04) so I

guess it’s a a Hebrew word, but it is means a custom or tradition, usually a choice of phrasing or a single world word that distinguishes one group of people from another. So historically shibbolets have been used as passwords, which is probably why the identity thing.

Shimin (42:19) Mm.

Dan (42:20) the reason why it popped into my head was there was a if you ever watched what was that president show that everybody loved for a long time.

Shimin (42:28) House of Cards.

Dan (42:28) no, the something West Wing, there we go.

Shimin (42:31) West Wing. Yeah.

Dan (42:31) Yeah. there’s an episode called Shibolith in it where he like makes a big deal out of the definition and everything. But that’s really kind of what I’m taking away from this article is that AI and and or the

necessity of success of AI projects has become sort of a shibleth in the industry. and what I mean by that, I’ll take you through some of the examples that that he goes into in this is that like you’re basically either in or you’re out on it, right? But when you’re out, if you don’t do the dance, you aren’t gonna succeed or in some cases like get jobs. So a little bit of context.

The the guy’s blog, he’s a consultant and one of the big things they do is sell people on Snowflake, basically, right? So if you don’t know, it’s Snowflake’s like big data warehouse. pretty great. yeah, especially for like shoving all your organization’s data into one place and then being able to like, you know, do BI queries and stuff like that on it. It’s pretty pretty powerful tool. So

One of the phenomenon that he noticed is that they were giving demos of like all their capabilities to customers and customers kept asking them for a demo of Snowflake’s AI tool. So Snowflake has this thing called I think it’s Cortex is what it’s called. And basically it’s like a it reads all the metadata, like all the schemas and tables and everything on your your warehouse, and then it can theoretically you can use it to like

answer BI type questions, right? So you could be like, you know, what percentage of my customers are buying wingdings or something like that? And it’ll just like spit out the query for you and you can you can do. Yeah.

Rahul Yadav (43:59) Hundred percent obviously. You don’t

even need to query the database for that one.

Shimin (44:03) We only sell wing tanks, yeah.

Dan (44:05) yeah.

So the author claims like he was given a presentation on it on Cortex by Snowflake staff and they reported that basically if you have all your data like perfectly sanitized and and like well metadated and everything, you only get like something like ninety two percent accuracy for the the the data, which is basically best in class for the the tool. But imagine your CFO is running that query and only

you know, one in ten of the numbers they run is outright wrong. and then they go ahead and put those numbers in front of like you know, something like a financial disclosure or release or something like that. So so that’s their point. But the part where it starts getting weird is they

So folks ask them for the Cortex demo, they start running it, and then everybody that they ran the demo for lost their minds. They would just immediately start throwing money at them. And and there’s a little poll quote that I want to read because it’s just hilarious. So because every lukewarm client, meaning a little bit of context, meaning they had, you know, weren’t super interested in their actual consulting stuff, that saw the chatbot in action, even with us telling them that it was not going to accomplish what they wanted.

wanted to buy it immediately. Every other consideration,

Rahul Yadav (45:13) Ha ha ha.

Dan (45:15) including millions of dollars that we could plausibly help them achieve by non-AI means, was swept aside. It was like a dark and terrible force seized control of their limbs, plunged their hands into their own chests, and presented their still beating credit cards to us in grim supplication. We were so mortified by the inexplicable shift in energy that we walk we parentheses wisely declined to take the money and ended the sales process.

And soon thereafter removed Cortex from our list of demonstrations. So, I mean, there is something going on, right? Where it feels like, okay, here’s the technology, and everyone is just like falling over themselves to implement it on absolutely everything. Like I think mouse drivers were sort of the the classical example. But the part where I think he kind of loses the plot a little bit on the blog post is then he summarizes that it comes down to a coordination problem.

Between executives. And the reasoning, his reasoning goes like this, and this is where I really I don’t find this part particularly plausible, which is that so your software is reliant on you know customers A and B to function, and customers A and B are saying that they’re

Writing 90% of their code with AI now and you know, saving 60% of their time. So do you look like an ass if you don’t also say that? And then what happens to your SaaS customers that are then buying your SaaS if you’re also not saying that? So basically he’s implying that like if everybody doesn’t keep lying about their AI projects being super successful, that like the whole thing sort of falls apart.

And if you don’t lie like that, then you’re being kicked out. I don’t think there’s a lot of evidence for like being kicked out. Like I don’t we haven’t heard of any like pr CEOs in particular, right, being fired for not like towing the line on this. so that’s why I think it’s a little bit less plausible. But in my mind, what I think is likely is like, you know, CEOs are using this stuff too, right? Just like we are. And do you remember that open source?

evaluation that they did where there was like a I don’t know perceived 20% speed up and then like it was actually less than that, which be real interesting for someone to reevaluate that in, you know, 2026. But I think it’s like executives use it and they feel that, you know, perceived 20% speed up. So that’s why I also don’t think the coordination problem is necessarily the the

the real outcome there, right? Is like they use it and they’re like, if I’m using it for like, you know, managing, I don’t know, this Excel doc, what if what if people are writing code with it, you know?

Shimin (47:38) Yeah, it’s it’s definitely an interesting game theory based explanation for the kind of AI psychosis that we kinda see in the corporate world these days. this gong ho well, like LinkedIn, right? at least my LinkedIn, it’s wall of wall either super pro AI or it’s wall of wall, like AI is the worst thing, it’s really dumb. Look at all the the times it has hallucinated.

Dan (47:59) And that and that’s kind of what I meant by

the the sort of like shibbolet thing, right? Is it’s like you’re you you’re you’ve either crossed that line and you’re all the way on one side or you’ve not crossed it at all and you’re all the way on the other. And it’s like somehow in there the nuance is getting lost and that’s a little unfortunate, I think. ‘Cause I think like, you know, we’re doing a podcast about this, so we all obviously agree that it’s a pretty useful and good tool. Yeah. I know.

Shimin (48:20) Yeah. It just it just hacked hugging face last week, so you

know.

Rahul Yadav (48:24) what’s what’s interesting is we th th there’s this third tiny sliver between the two which I would put the three of us in where we talk about all the good things AI is doing or the crazy stuff it’s capable of now and at the same time we often also have the two minutes to midnight section right and so you and we’ve talked about the token maxing

craziness and there is a third way where you can just go, well, what places are this useful in? We can use it there. And in some places it’s just mania and we don’t need to, you know, jump off a cliff with everybody else. so I’m sure there aren’t as many people like us and maybe they’re not as loud but there’s definitely that third category.

Dan (49:08) I know. We n we need to start a movement.

We need to name it first so it can get popular and then

Rahul Yadav (49:12) Yeah.

Dan (49:13) actually just start professing that viewpoint, which is that like maybe this is pretty useful at certain things, but maybe not everything, and maybe we shouldn’t put it in mouse drivers, but you know, maybe we should use it in a harness to do X like

Shimin (49:27) We’re the hyped Luddites of the AI world. so one piece of hard evidence that was in this blog post was he spoke to an executive who was all in on being AI native, but has actually not used AI at all in their personal lives or just

Rahul Yadav (49:43) Ha ha ha.

Shimin (49:45) period. that’s just bananas to me, right? Like this goes back to we covered

an article a couple of weeks back about how like your engineering manager must be at least be your average AI developer, if not better. Cause if you don’t use it every day, how can you possibly know what the power of this technology is? It’s not gonna be the 90 second demo that’s like

looks amazing but it’s only eighty percent of the way there and the last twenty percent will take another a hundred and eighty percent of the the project scope. So

If there are a good amount of executives out there who truly like never uses AI day to day but is talking about going AI native, I’m worried for those companies and they’re not gonna be able to do this AI native transformation. There’s there’s just no way.

Rahul Yadav (50:30) I agree. And how would you even know what transformation through w wha how would you even define transformation? You know? Yeah.

Shimin (50:39) Right. Yeah, just

just slap a router, so that we can save on AI token costs, right? Like

Dan (50:44) Yeah.

Rahul Yadav (50:48) You first token max, then you say that’s over, then you do routing, then you say that’s over because routing is hard. Then you go back to like everybody gets a hundred bucks a month or something. And maybe that’s where we settle. Or everybody’s using Kimi or something by that time.

Shimin (51:05) But also like what is a responsible way of selling AI based tools and and doing AI based transformations without going you have to thread the needle pretty finely without going into either the hype category or the this thing is terrible, you shouldn’t use it ever category.

Rahul Yadav (51:22) Yeah. And good for him. There are definitely people who are just, you know, selling others on the hype and trying to make money without necessarily any results to follow that. So at least you know the author here is sticking by their principles and not taking on things that they don’t fundamentally believe believe what I agree with.

Shimin (51:42) Yeah. We need more people like this.

Okay. going back a little bit to AI actually being useful and is able to hack the

Dan (51:50) Yeah.

Shimin (51:50) infrastructure of hugging face there.

Dan (51:52) And values.

Shimin (51:53) Yeah. my article this week is from the Anthropic Research blog about clouds, values, across models, and languages. the motivating

Rahul Yadav (52:00) Yeah.

Shimin (52:02) question here is what should an LLM do?

when there’s no universal right answers to your question. Like whether to take a new job or how to handle conflict with a friend. Or I guess in this case, how to solve the exploit gym puzzles without

Dan (52:19) No the

it’s easy, you just agree with whatever the user said, you know, just sorry.

Rahul Yadav (52:23) Sell more

sell more wing dings is the way which

Shimin (52:26) Sell more wing dings thanks always.

and we know this is a topic that Anthropics spend a lot of time thinking about and actually writing about, right? They we talked about Claude’s constitution, we talked about their interviews with spiritual leaders and the hiring of the philosophers. So this is a continuation of that research process where in a previous work, which we didn’t cover, they analyzed 700,000

Conversations, anonymized conversations, and then identified 3,000 distinct values in cloud’s responses in those conversations. And this time they said to themselves, 3,000 values is too many damn values, which I agree with. Like no one’s gonna go through 3,000 values to find out how your model behaves. So they broke down those 3,000 values down to four axis.

To measure the response of their models. The axes are deference versus caution. So this is something like whether you just accommodate a user’s requests, or do you caution them against doing a bad thing, right? Do you hack into hugging face? Or do you say, hey, hacking is bad, we shouldn’t do that. warmth versus rigor is this, you know.

This is the classic syncopency. Do you make the user feel good or give do you give them the hard truth? There is depth versus brevity, you know. do you write out really robust outputs or do you just stick to the important bits? And then you have candor versus execution, which is there’s a little bit of overlap. We’re gonna talk a little bit about this later. candor is is like honesty and transparency, and execution is optimization and being results oriented.

so after they did this dimensional reduction of their 3000 values, they created these axes and they found that even though a single conversation can exhibit values from both sides of the axis, it could it could sound accommodating but also be responsible at the same time. In practice, what they found was models tend to stick to one versus the other in a single conversation, anyways. And

Then the question becomes like okay, now that we have a way to measure how a model behaves on these four axes, are different classes of Claude models at least, do they have different personalities? And it turns out it is true they do. and they have different values and the values do match users

feeling about how this model behaves. So the classic one is opus 4.7 pushes back a lot more than opus 46 did because it has higher caution and lower deference which is interesting right but then the follow-up they did with these four axes is they f they try to see if different languages

causes the model to have

Dan (55:05) Mm.

Shimin (55:06) different exhibit different values on these axes. And the classic example here is, you know, Arabic and Hindi, at least in the Claude class models, are much warmer than English. and they are wondering if this is a shortcoming of their training data, where they just happen to have a lot more warm conversations in Arabic and Hindi to cause the model

to have that behavior or is it something else? Or it’s possible that whether or not a conversation exhibits these values were judged by Claude by an L O I think there may be biases in there that is gonna be just built into the interpretation piece. I I don’t think that’s super clear. But it is

at least to me, very interesting to see how the values differ across languages for the anthropic family of of models. And not to

Rahul Yadav (55:56) Look on the last

one, Shimin French one is funny to me. Average, average.

Shimin (56:00) Yeah, French.

Rahul Yadav (56:01) Tiny depth average.

Dan (56:04) Yeah.

Shimin (56:04) French is completely

average, other than a little bit more depth. you know, since this is a podcast and we should do hot takes, do we do you wanna say some languages are just warmer than other languages? Right? Like if I think of a language like Russian, I would not think Russian is particularly warm. Maybe it’s more on the rigor side of things. maybe it’s ‘cause Russia’s cold. I did take one semester of Russia.

Dan (56:24) Do they have

do they have German on there?

Shimin (56:26) In high school, so

I do believe they have German in here.

Rahul Yadav (56:30) Yeah, in the fourth

one. Yeah. Right there.

Dan (56:33) yes.

Shimin (56:33) German.

found it. German. Yep.

Dan (56:34) Depth, candor.

Sounds about right.

Shimin (56:36) get straight to the point and says when it’s unsure. Yeah.

Dan (56:37) Rigor and caution. I mean they

literally have a word for a face that desires to be punched or something like that. A punchable face. There’s an actual word for that in German. Ugh.

Rahul Yadav (56:49) The only word I know in German is how to say asshole in German. So

Or how to, you know, keep an ear out for it.

Shimin (56:57) I like where this role this trace of research will lead us to, which is you should be able to tell the agent be less warm. Right. We we often talk about how when you tell models not to be overly syncophantic, it just becomes an abusive person. But maybe you can like tweak the value to find one that you like. You know, it’s neither too warm nor too rigorous. Somewhere in between the two, the one that’s in a Goldilocks zone.

Dan (57:21) Back back Pfeiffing It’s a face that’s badly in need of a fist. Sorry,

I got stuck on it. I couldn’t. Also now we know how bad my German is. I did not take German ever.

Shimin (57:29) Just we shouldn’t talk about how

Rahul Yadav (57:29) The

Shimin (57:32) how each co host look like on on the show. I don’t care what language it is.

Dan (57:37) I mean of the three of us I’m I’ve got the most punchable

face, just saying.

Shimin (57:40) I think we all have pretty punchable

Rahul Yadav (57:41) It’s

Shimin (57:42) faces.

Dan (57:42) Let’s not find out.

Rahul Yadav (57:42) You know,

i it would be interesting to write those lighthouse keeper stories in each of these languages and then see how it comes out.

Dan (57:50) See which is the warmest.

Rahul Yadav (57:51) Yeah. I I feel like over time would this change the real world text where people are also writing it to not come off in the like whatever

Shimin (58:04) Mm.

Rahul Yadav (58:04) the axis. You know, ‘cause the more the models get influenced, they’ll influence the real world too. So it’d be interesting to look at this a few years down.

Shimin (58:12) Yeah.

So so maybe if the models are very warm, it’ll cause humans to be warmer too in our communication. That’s

Rahul Yadav (58:18) I hope so.

Shimin (58:19) a possibility. Yeah. We we don’t all look on the the dark side of AI on this show. We sometimes think about possible good futures

Dan (58:26) The other interesting thing, and I I debated trying to invite this person on as a guest because they’re not at all related to like computer stuff. But I had a really fascinating conversation with one of my wife’s coworkers about is AI influencing like global writing styles as well at scale because people are like treating it as if it is the b you know, they’ll they’ll they’ll pr write something that’s perfectly good.

And then they’ll pass it through an AI and be like, Fix this, right? So it’s like subtly influencing things even if it’s not doing, you know, X versus Y kind of stuff. so like what impact is that having on like global writing? And I’m like, that would be a good segment, but so

Shimin (59:02) Yeah, with over fifty percent a

Rahul Yadav (59:02) Lot more substrates out there now.

Shimin (59:05) lot. I I’ve read

Dan (59:06) Or

Shimin (59:06) a couple of substrates IRL. I we heard it the other day and I was like that is a AI poisoned brain potentially.

Dan (59:14) Yeah. Well

Rahul Yadav (59:14) Yeah.

Dan (59:15) and and the other point that she had made that was pretty fascinating was like also, if you modify your writing to not to try explicitly to not sound like an LM, right? Because there’s all these like sort of telltale things that you might have just been doing anyway, to some degree, is that making your writing meaningfully worse, potentially? Because you’re sort of like stepping around trying to to not do it. So anyway, food for thought.

Rahul Yadav (59:38) We

We’ve also talked about freedom of speech on this podcast. At what point do you know, does the AI start editing your stuff to make you think

Shimin (59:49) Mm.

Rahul Yadav (59:49) you’re saying what you thought you were gonna say, but at the same time censor you and whatever you’re saying is not coming off how you wanted to say it. And then the more you start reading it and using it for everything, it’s this like subtle censorship going on.

Shimin (1:00:04) Right, like we just collectively forget like what happened in T Square in like nineteen eighty six, right? ‘Cause you interact with AI all the time and you can’t they can talk about that. So

Rahul Yadav (1:00:15) Not as long as this podcast is going on and references it every week.

Shimin (1:00:21) We we are clearly getting

censored in in the greater Sino sphere. Yeah.

Dan (1:00:23) That’s that’s just our way

of defending our content from LLMs.

Rahul Yadav (1:00:29) Ha ha

Shimin (1:00:29) Ha

Rahul Yadav (1:00:29) ha.

Shimin (1:00:29) ha.

Dan (1:00:30) Last week we talked about terrorism and next week we’ll talk about

Rahul Yadav (1:00:33) This week it’s

Dan (1:00:34) the recipe for bad thing, you know.

Shimin (1:00:36) Mm. This is how you know

Rahul Yadav (1:00:36) De

Shimin (1:00:37) we’re not AI generated, guys.

Dan (1:00:38) Whoa.

Rahul Yadav (1:00:38) You guys watch Breaking Bad? You both have watched that one, right?

Dan (1:00:42) Mm-hmm.

Rahul Yadav (1:00:42) It made me think of the Jesse Pinkman, my secret ingredient is chili powder and the meth in the That’s when you should ask D.I. for. Meth but with chili powder in it.

Shimin (1:00:52) Chili powders, yes.

all right, on to our last segment of the week. and this is a favorite one of me and Dan. two minutes to midnight where we talk about the financial side of the AI industry using the atomic clock from the bulletin of atomic scientists, where if we hit midnight then we hit mutually assured destruction.

And we are at four minutes and forty five seconds as of two weeks ago, ‘cause we skipped last week.

Dan (1:01:20) Yep. so who’s kicking it off? Me.

Shimin (1:01:23) Yes.

Dan (1:01:23) All right. So all of my dreams have come true this week. All of them. Every single

Shimin (1:01:28) Yeah.

Dan (1:01:29) one. if you don’t want to be well, first of all, SpaceX is down, I think. Is the last time I looked? Below below the IPO. So there’s that, but

Shimin (1:01:33) Mm-hmm. It’s like one twenty six last time I looked. Yeah.

Dan (1:01:36) there is there are two new ETFs.

That basically track existing ETFs out there. So like one is I think the SP 500, and then the other one is the

Shimin (1:01:48) Mm-hmm.

Dan (1:01:48) NASDAQ 100. And they’re both so the the one one is called NASDAQ 100 X Elon Enterprises ETF. And the other is called SP 500 X Elon Enterprises ETF. so they are both the same thing as the other index, except that

They don’t include any companies that Elon is involved in. Which is I think basically just only Tesla and SpaceX are publicly traded at this point.

Shimin (1:02:16) Yep. Yep.

Dan (1:02:17) but I thought that was pretty funny because we were just we’d pri previously last time we’d gone through the segment, we talked about like, you know, what is the risk to the entire economy really that’s like being, you know, brought on us by all of this and then how

Badly this could impact people’s pension funds,

Shimin (1:02:32) Mm-hmm.

Dan (1:02:33) right? If if that happens. So now there’s a way to avoid it if you want to. also worth noting that the company that

Is is creating these funds is called subversive capital. and so like their whole shtick, I guess, is kind of doing sort of tongue in cheek funds like this. But as far as I know, they are in fact real funds that you can actually buy. but their their other one that that got some publicity back in the day was they created one that or I guess there’s two, one for each side of the members of Congress. So

depending on, you know, pick your party. it’s all the the funds that are known or stocks that are known to have been traded by either the Democrats or the Republicans. So you can you can buy your political affiliation as a

Shimin (1:03:14) Yeah.

It’s a little clickbaity, but

Dan (1:03:16) Yeah.

Shimin (1:03:17) But just to show the confidence in in Elon is maybe not w what it once were.

Dan (1:03:22) Yeah.

Shimin (1:03:23) Yeah. Alright, well my story of the week is the fact that SP Global has downgraded Oracle’s from BBB to BBB minus, which puts them one grade above junk bound status, which will be C C C Plus, I believe. and this is because Oracle has there’s just a lot of commitments for its AI build out.

they have I think four X leverage is what they are using to determine the rating of a company. And Oracle would remain above this four X, so four times debt to equity benchmark, I believe, so as a you know, the Oracle hater on this pod, I am taking a little bit of solace in this.

All And Rahul, what do you have for us?

Rahul Yadav (1:04:06) this one’s about Neo Clouds, by Beth Kindig. from

Dan (1:04:10) What is a neo cloud?

Rahul Yadav (1:04:13) all these new Coreweave and Nebius and the shoe company that now is a inference serving company. I think they

Shimin (1:04:21) Allbirds

Dan (1:04:21) right.

Rahul Yadav (1:04:22) all birds became Smartbird or New Bird? They changed their name pretty quickly over

Shimin (1:04:26) okay. I’m behind, yeah.

Rahul Yadav (1:04:28) yeah,

Dan (1:04:28) No bird.

Rahul Yadav (1:04:29) so n now they’re called one of those.

Shimin (1:04:30) Jailbird.

Rahul Yadav (1:04:31) Yeah, those two things.

Dan (1:04:33) That was only the parts of their compute that were running the open AI model that escaped less this jailbird.

Rahul Yadav (1:04:38) Yeah.

so yeah, i neo clouds are a phenomena of this recent AI boom. and they are currently a key component in the hyperscalers, how they structure their finances and how they communicate them. So

Hyperclo scalers have committed hundred and fifty billion to neo clouds. overall they’re collectively spending a trillion dollars. The reason why they’re doing that is because when they spend their own money, that’s CapEx, and that’s that gets compared against their operating cash flow. But when they rent these this capacity from neo clouds, they can put that under the opex category. And so that’s like

Or five years, if I give you, you know, a hundred billions, that’s only I’m only gonna count twenty billion or something, and that’s gonna go in my oppEx versus CapEx. so if you look at Microsoft’s commitment, they had close to nine ninety pr percent plus of their operating cash flow would go towards CapEx for this year. that doesn’t even account for what they’re gonna buy from NeoClouds. Once you add that, the they’re actually spending more than

the overall operating cash flow. So the hyperscalers are leaning a lot on these neo clouds. Now they’re backed by mostly Nvidia in most of the cases where it serves as what we have here as the like the investor of these neo clouds, the supplier of them because they’re providing them all the GPUs, and as the demand backstop in case it doesn’t materialize for them. So and a whole bunch of circular financing goes

And then one interesting thing that I learned from this was Coreweave and all all this like the money that neoclouts are raising, it’s all debt financed. And

Dan (1:06:23) Mm-hmm.

Rahul Yadav (1:06:23) CoreWeave their interest payments are already taking up twenty-five percent of their revenue. So you know, they’re kinda

Dan (1:06:31) Ha ha

Rahul Yadav (1:06:32) in a race with the US government soon if they keep going like this. so

Shimin (1:06:35) Mm-hmm.

Dan (1:06:36) To see who can hit the biggest

debt number.

Rahul Yadav (1:06:38) Yeah,

who who is gonna hit it? so yeah, if one of these pieces kind of fall out of place that could trigger some big problems here. we’ll see where this goes.

Dan (1:06:48) as we talked about two weeks ago, this is the the debt is increasing the impact that something like this could have have if like bubble goes, yeah.

Rahul Yadav (1:06:55) Will have yeah. Yeah.

Shimin (1:06:57) Yeah.

Dan (1:06:58) So

Shimin (1:06:58) So hyperscalers got

downgraded. The neo clouds have unsustainable amount of interest payment. And the one that has IPO is getting the ETF treatment. but I think the biggest two minutes to midnight topic for this week is probably still the Kimi and the Qwen news item from the top of the show ‘cause that has such a

That really just commodifies tokens, right? And reduces the pricing power that the Frontier Labs have. so all that said, how do we feel about the clock this week? We were at f four minutes, forty five seconds.

Dan (1:07:33) Forty five.

I guess getting more nervous, right, with the debt and the

Yeah.

Shimin (1:07:38) Mm-hmm. How

did that?

Dan (1:07:40) Just keep adding debts together. Yeah.

Shimin (1:07:41) Yeah.

I’m I’m definitely I definitely think we should move it forward, just ‘cause whatever revenue numbers of Anthropic and OpenAI will likely take a hit with these latest Chinese open weight models.

Dan (1:07:53) And Nathan, who we we had on as a guest a while ago now, we should have him back at some point, also dropped a blog post, I guess it was last week, that was like basically fire is at risk because of the meaning like you know, financial independence movement to try to basically like live pretty frugally and then basically live off your investments is at risk because of the sort of like what he’s calling like the bimodal distribution of the outcomes here, meaning

There’s basically two things that could happen. It’s gonna go really, really good or really, really poorly. And so the model that fire depends on is that it all just sort of keeps marching along over time, right? So yeah, I read that article and I’m like, ooh, that sucks. Cause I was kinda hoping that that would like

Shimin (1:08:32) Ha ha.

Dan (1:08:32) keep us all going after if this thing does fall apart. yeah, I’m I’m in favor of moving it up.

I I think a bit.

Shimin (1:08:40) For thirty.

Dan (1:08:40) Yeah, that’s fine. We still don’t have any other IPOs yet. So we’re still missing some data there.

Rahul Yadav (1:08:46) the Q two earnings will happen within the next ten days.

Dan (1:08:49) Yeah.

Shimin (1:08:50) Yeah, so SpaceX earnings will come out, maybe by the next show. We’ll see.

Dan (1:08:53) Yeah.

Rahul Yadav (1:08:53) about space but hyperscaler ones. I guess SpaceX is also hyperscaler now. I don’t know, whatever they are. Yeah.

Shimin (1:08:59) Yeah, they are. They’re they’re leasing

causes to anthropic, so okay,

Dan (1:09:03) Yeah.

Shimin (1:09:04) so let’s do four thirty. That works for me.

Dan (1:09:07) All right.

Shimin (1:09:08) Alright, and of course, we’re the setting of the clock. that is the show folks. Thank you again for joining us for our study session this week. If you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcast or Spotify. It helps people to discover the show and we really appreciate it. If you have a segment idea, a question for us or a topic you want us to cover, shoot us an email at humans at adipot.ai. We’d love to hear from you. You can find the full show notes or

transcripts and everything else we mentioned today at www dot adipot.ai. Thank you again for listening and we’ll catch you next week. Bye.

Dan (1:09:41) Cheers.