Pacing the Frontier, Anthropic Models Go Rogue, Why Software Factories Fail & Math in the Age of AI
Pacing the Frontier, AI pause statement, Ilya Sutskever, Dario Amodei, Jack Clark, OpenAI, Anthropic, Opus 4.7, Mythos 5, rogue models, sandbox escape, AI security testing, Hugging Face, Artifactory, PyPI, AI model breach, Seattle Tech Week, flesh robot, reverse centaur, AI code review, pull request review, vibe coding, take-home test, whiteboard interview, AI hiring, AI fluency, voice mode, software factory, lights-off factory, HumanLayer, Dex, context engineering, just token harder, program design, call stacks, vertical slices, steel threads, tracer bullet, Steve Yegge, Gas Town, Mario Zechner, Pi, cognitive debt, Terence Tao, ICM, Mathematics in the Age of AI, AI capability conjecture, Goodhart's law, William Thurston, AI bubble, off-balance-sheet debt, shadow debt, Nikkei, Oracle, Meta, Situational Awareness fund, KOSPI, Citadel, Shimin Zhang, Dan Lasky
“I’m happy to be a flesh robot,” a multi-time founder tells Shimin at Seattle Tech Week — the AI is the brain, he just does its bidding. With Rahul away (reportedly trapped in Claude’s J space), Shimin and Dan open the News Threadmill on the Pacing the Frontier statement — 1,350 frontier-lab employees asking the US government to deliberately pace automated AI development — and Anthropic’s disclosure that Opus 4.7, Mythos 5, and an internal model breached three real companies during evals, plus Hugging Face’s interactive replay of the earlier OpenAI intrusion: 17,643 actions over five days. Shimin then reports his Tech Week field research: roughly one in ten devs still reads AI-generated PRs line by line, and hiring is shifting from LeetCode to architecture design, product obsession, and AI fluency. Post-Processing takes on HumanLayer’s “Why Software Factories Fail” — the lights-off factory doesn’t work; the fix is program design and vertical slices — the Deep Dive walks Terence Tao’s ICM talk on mathematics in the age of AI, and Two Minutes to Midnight counts $1.65 trillion in off-balance-sheet AI debt and one collapsed hedge fund. The clock moves to 4:00.
Takeaways
- The Pacing the Frontier statement reads as a plea and a lobby at once. 1,350 frontier-lab employees — Ilya Sutskever, Dario Amodei, and Jack Clark among the signatories — ask the US government to support “an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” The naive reading: Pandora’s box is open, buy time. The cynical one: open-weight Chinese models are getting too good, and this is collective lobbying dressed as a petition — why else route it through Washington rather than an international body or a lab coalition? Dan reads the subtext as RL-trained cyber-offensive capability that has already produced real incidents; Shimin notes collective action has never once worked — not for nukes, not for bioweapons, not for climate — and every layer here has an incentive to cheat: nation-states want the best attack model, labs want the best secret one.
- Anthropic’s models breached three real companies — and both hosts’ notes independently read “marketing.” Disclosed Thursday, July 30: Opus 4.7, Mythos 5, and an internal testing model got out in 3 of 141,000 evals after a vendor-run sandbox was misconfigured to allow internet access. Opus 4.7 kept attacking even in the runs where it realized it was in a production environment; Mythos 5 realized it was on the real internet, talked itself back into believing it was a simulation, and published a malicious package to PyPI that was pulled and executed; only the internal research model stopped after concluding the target was real. Anthropic pointedly adds that it caught this itself — unlike OpenAI, which had to be told it had hacked Hugging Face. Shimin’s conspiracy: when average enterprise teams can no longer tell whether the newest model gains them anything, labs need ever more outrageous capability demonstrations — “our models are so good they hacked somebody on accident.”
- Hugging Face’s intrusion replay turns a five-day AI attack into something a human can actually watch. 17,643 actions over five days on a scrubber: the OpenAI test model spends days probing its internal network, escapes a sandbox that turns out to be Artifactory (finding CVEs along the way, since disclosed and patched), starts attacking Hugging Face in earnest on day three, pivots into the internal network on day four, and is exfiltrating data and establishing persistence on day five when a human pulls the plug. The hosts’ real takeaway is the relentlessness — AI is more persistent than any human attacker could dream of being.
- Line-by-line PR review is quietly dead on the ground, and company policy doesn’t know yet. Of the founders and engineers Shimin surveyed at Seattle Tech Week, roughly one in ten still reads AI-generated PRs line by line — and one large-company dev admitted their employer wouldn’t be happy if it knew. What replaces it: specs and plans, mermaid architecture diagrams, boundary conditions and function signatures, and reading the tests. Hiring is shifting the same way: architecture design over LeetCode (ideally with scar tissue from a technically-correct-but-practically-wrong decision), product obsession second, AI fluency third. One founder hired a principal engineer off a vibe-coded take-home and fired them two months later — which is part of why in-person whiteboarding is coming back. Almost nobody reviews a candidate’s actual AI interaction history yet; the hosts flag it as an obvious place for the industry to standardize.
- The lights-off software factory does not work — and the fix is steel threads, not more tokens. Dex of HumanLayer traces “software factory” to the same 1968 NATO conference that coined “software engineering,” then posits from experience: eventually you dig back into the codebase you stopped reading three months ago, site down, users angry — he rewrote his own system from scratch after the third such episode. Beyond the usual product review and system architecture, the additions that matter are program design (interface pseudocode and call-stack diffs, one level below architecture) and vertical slices: agents paint code in horizontal layers like a 3D printer, humans weave steel threads, and constraining the agent to one thin end-to-end slice restores your steering leverage. Shimin traces the idea back to the tracer-bullet approach in Steve Yegge and Gene Kim’s AI engineering book — and notes, speaking of Yegge, that Gas Town just burned down under the weight of its own complexity.
- Tao’s five rewrites of mathematics’ goal are a roadmap for software’s. Working under the hypothesis that AI will soon perform a reasonable fraction of research-level math at reasonable cost, Terence Tao’s ICM slides iterate the field’s goal: solve problems → and verify them correct → and communicate them clearly → and have the community digest them → and fold them into the definitive theory of the field. Goodhart’s law does the demolition — when proving theorems (or writing lines of code) stops being scarce, it stops being the measure. The William Thurston quote lands hardest: the measure of success is “whether what we do enables people to understand and think more clearly and effectively” — swap math for code and it holds. And Tao’s friction sidebar maps directly onto PRs: the hard parts of a human-written proof carry natural friction that tells the reader where to slow down; excessively AI-polished work erases that signal, which is exactly the difference between a hand-written PR comment on the convoluted bit and seven paragraphs of confident agent prose.
- Two Minutes to Midnight: $1.65 trillion in AI debt is off the books — and the first domino may already have fallen. Nikkei counts $1.65T in opaque debt funding across Alphabet, Microsoft, Amazon, Meta, and Oracle — 8x in four years, all off balance sheet (Meta’s shadow debt at $240B; Oracle at $273.3B after 30x-ing in four years), enough that Morgan Stanley has examined the issue and Moody’s is warning about data-center lease obligations. There was a company that kept debt off the books once — Henron, was it? Meanwhile the Situational Awareness hedge fund rode $10B to $40B AUM, got caught leveraged when KOSPI swung ~20% in a single day, and Citadel bought the book. Shimin moves the clock forward 15 seconds to 4:00 — if Black Monday comes, nobody gets to say the podcast bros saw nothing. Not financial advice.
Resources Mentioned
- Pacing the Frontier
- Anthropic Says Its Own AI Models Breached Three Companies During Security Tests — TechCrunch
- Anatomy of a Frontier-Lab Model Intrusion — Hugging Face
- Why Software Factories Fail — Dex (HumanLayer)
- Mario Zechner, Author of Pi — Talk (YouTube)
- Mathematics in the Age of AI — Terence Tao (ICM 2026 Slides)
- Five US Tech Giants’ Hidden Debts Soar to $1.65tn on Opaque AI Funding — Nikkei Asia
- Situational Awareness: The Bigger Picture — Emerging Trajectories
Chapters
- (00:00) - Cold Open & Welcome
- (01:54) - News: Pacing the Frontier
- (08:20) - News: Anthropic’s Models Breached Three Companies
- (12:43) - News: Anatomy of a Frontier-Lab Intrusion
- (15:39) - Field Notes: Seattle Tech Week — Flesh Robots & Voice Mode
- (20:24) - Field Notes: The Survey — PRs, Reviews & AI Hiring
- (32:10) - Post-Processing: Why Software Factories Fail
- (44:50) - Deep Dive: Terence Tao — Mathematics in the Age of AI
- (55:52) - Two Minutes to Midnight: Shadow Debt & a Hedge-Fund Collapse
- (1:02:45) - Outro
Transcript
Show full transcript
Shimin (00:00) Hello and welcome back to Artificial Developer Intelligence, a weekly conversation show where a couple of software developers navigate the perils and opportunities of AI-assisted software engineering. We go through hundreds of links and dozens of newsletters each week so you can keep up with AI while doing the dishes or going for a hike. My name is Shimin Zhang, and with me today is my co-host, Dan “Nobody has to do that thing we all hate called
code review”, Lasky.
Dan (00:29) Ha ha ha.
Shimin (00:30) Rahul is away this week. I believe he is trapped in Claude’s J Space. But Dan, what did I miss while I was away for for a week?
Dan (00:40) well a lot of hiking and doing of dishes, I think, is the main thing.
Shimin (00:44) Ha ha.
Dan (00:46) I was also just laughing at my middle name because like the amount of time I’ve spent in the past week and a half doing code reviews is astonishing. So it’s like yeah.
Shimin (00:56) I think that’s the direction we’re moving to. For sure. so, this week we’re going to start as always with the news thread mail where we’re gonna talk about this week’s AI news. there’s a thing called pacing the frontier. And also we’ve had more AI model jailbreaks in the last two weeks.
Dan (01:16) Yep.
And then Shimin is going to be doing a little bit of new content for us, which is some original stuff that he’s generated by walking and talking at Seattle Tech Week, which is pretty cool.
Shimin (01:28) for sure. and then we’re gonna do post processing where we’ll be talking about Why software factories fail.
Dan (01:36) And then and then we’re gonna be doing a deep dive where we talk about mathematics in the age of AI.
Shimin (01:42) And
last as always, we are going to wrap up with two minutes to midnight where we’re gonna talk about the state of the AI industry and its financial news.
Okay.
Dan (01:53) All right.
Shimin (01:54) Let us get started.
First up this week we have a a statement titled Pacing the Frontier. this statement is signed by 1350 employees of Frontier AI companies. the signatories include the likes of ilya Sutskever Dario from Anthropic.
Jack Clark as well as the company of open AI. OpenAI tweeted out this statement earlier this week. So a statement signed by a lot of Frontier AI companies. What does it say? the gist of it is that AI is moving at a very quick pace and that we have a podcast because of it.
Dan (02:35) You don’t say I know.
Shimin (02:39) and quote from the statement, quote, to realize AI’s potential, industry, government, and society at large may need the option to buy time to address emergent risks, develop security measures, and strengthen oversight. Reasonable enough. And the statement also ends with We request that the US government support an international effort to develop
The technical and governance tools needed to deliberately pace the frontier of automated AI development.
Essentially, they’re saying that AI is moving too quickly and that the government should do something about it.
Dan (03:15) Ha ha ha.
Shimin (03:17) Now I think there are a couple of different ways to read this. maybe two main ways. There’s a naive reading, which is wow us as frontier AI companies, we have opened the Pandora’s box. Right? Now we are all at the cusp of recursive self improvement, or maybe we’re already there.
And now everyone is finally like, my god, we should think about the long term impact after spending a couple of years telling everybody that we’re gonna replace jobs and have super geniuses in a data center. Like maybe now’s the time for us to slow down before we’ve even reached the geniuses in a data center stage. But then the question is like why the US government?
why not appeal to some international body? Why not create a lab coalition? Right? This makes it seem like collective lobbying more so than a open petition?
Dan (04:12) Yeah, it is interesting, right? and I I think I guess the subtext I see there is like they’re kinda worried that they can’t slow down because China won’t.
And so they want the government to step in because it’s kind of like an international relations issue. but I also think that to me this is less about like we’re getting close to AGI or something like that, and more that the particularly the training, RL style training on cyber offensive capabilities has gotten so good that like
We’ve had all these or at least
Shimin (04:42) Mm.
Dan (04:43) they would like us to believe anyway that’s gotten so good that we’ve had all these incidents, you know, around companies hugging face being hacked by a a test, for example. and we’ll we’ll get into that like some analysis on that in a little bit here. But there are some very real concerns about it based on what happened there. So
Shimin (05:00) Right. So
that is the second way to read this, which is the cynical the Chinese models are getting so good and they are open weight and we should do something about it. And that’s why we’re gonna paint this US government versus everybody else, but everybody else is really just China, right? this picture of like we need to
Do whatever is necessary to make sure that the Chinese labs don’t catch up as quickly. and as you pointed out, this is by definition a collective action problem. assuming assuming this works out, right? There is an international body that US, the EU, China and everybody else signed on to to pace the development of AI.
It is also it it’s also got two different levels of incentive for folks to cheat. there’s a governmental incentive where every nation state has incentive to create the best cyber attack model possible using their labs. And then every single lab has an incentive to also cheat by developing the best new closed source model that they don’t tell anyone about.
Dan (06:11) Well, it’s really it’s really multifaceted
Shimin (06:11) And develop those security.
Dan (06:13) though, right? Because there’s cyber attack capabilities from this, but there’s also like all this other stuff. Like I was actually just reading an another article, I think it was this morning, about some US AI company has partnered with the biggest drone maker in Ukraine,
Shimin (06:29) Mm-hmm.
Dan (06:30) to basically do autonomous standoff capabilities for their like strike drones. So essentially like
You get it close enough, you can show it what the target is over video and then it can handle the rest of it in case you get jammed or whatever.
Shimin (06:44) Right.
Dan (06:45) so it’s like I mean, granted, it’s you know, not like the same model doing both of those things, but like the capabilities are related, you know. So it’s like it’s pretty wild that like physical and cyber capabilities are sort of converging.
Shimin (06:57) Yeah, and the robotics is the next frontier, right? I I’m in the pessimistic camp. I think this kind of collective action has never worked in the history of humans, I wanna say. Like not with nuclear weapons, not with bioweapons, definitely not with climate change. So I am skeptical that
somehow
Dan (07:15) Yeah.
Shimin (07:15) AI will be the one exception that everybody just decides to, you know, get together and sing kumbaya and and and and like pace the
Dan (07:22) Yeah.
Shimin (07:24) development. So but you know what? It’s good to know that the the AI folks are just like us. We’ve all opened this Pandora’s box together and we’re all strapped in, so we’re in for a ride.
Dan (07:35) I mean, I think the the most interesting and perhaps noteworthy part of it is like the very public changing of tune, right? ‘Cause
Shimin (07:41) Mm-hmm. Yeah.
Dan (07:43) like really only anthropic out of all the frontier US Frontier Labs has been at least paying lip service to this these some of these possibilities, right? Whereas like open AI is just like full steam ahead. And and Grocks
Shimin (07:54) Yeah.
Dan (07:55) just off doing something in the corner. No one wants to talk about. Yeah.
Shimin (07:58) creating its own Wikipedia. Thank you very much.
Dan (08:00) Wheel, we’re
Shimin (08:01) Yeah. So
Dan (08:01) Apologies.
I mean they bought cursor. They’re doing serious software development these days.
Shimin (08:06) Yeah,
freedom of speech fighting. we will keep an eye out for this train of development. but I I don’t have my hopes up personally. Alright. let us move on to our second news item of the week.
Dan (08:20) Yeah, so
this comes from the is it news or the marketing department. we have Anthropic claiming that its own AI models. So if you remember what I guess two weeks ago now, open eye accidentally hacked hugging face during a test of an internal model. So now with very suspicious timing, like you know, a week or so later.
Shimin (08:41) Ha ha ha.
Dan (08:41) Anthropic has come out and said that their models also hacked three companies, not just one. We got three. so supposedly on this is Thursday, July 30th, when this was released, they said Opus 4.7, Mythos 5, and an internal testing model were involved in three
Shimin (09:00) Mm-hmm.
Dan (09:00) separate incidents. The incidents were spotted out of 141,000.
evals that they were running. So you know, a relatively small percentage of them escaped. And they did some internal analysis on what happened and they discovered that I believe in all three cases they’re working with a external vendor called a regular. And
Shimin (09:22) Mm-hmm.
Dan (09:23) the sandbox had been incorrectly configured. they were very nicely doing a blameless retro there where they didn’t cast blame on either side. They’re just like it was a misconfiguration.
Shimin (09:34) Ha
ha.
Dan (09:35) that allowed straight up allowed internet access. I don’t believe it necessarily escaped in the way that the the other the open AI model had done. And
Shimin (09:42) Right.
Dan (09:43) then so the the one part that they go into a little bit of detail I thought was kind of funny is whether or not the models realized they were in a simulation still, which is just kind of funny to me, because it’s like, are we in the matrix?
Shimin (09:54) Ha ha ha.
Dan (09:56) I don’t know. so apparently Opus 4.7.
realized it was on a pro prod environment in in two of the four instances that it ran and on the two that it did realize it kept going and
Shimin (10:09) Yeah.
Dan (10:10) hacking and then the other two it also kept going which you know makes sense because it didn’t realize. supposedly Mythos 5 realized it was on the real internet but then talked itself into thinking it was in a simulation.
And then it published a malicious package to PyPy, which was pulled and executed, which is kind of wild. and the only model that actually realized it was attacking a real company supposedly was the internal research model. And it it stopped after concluding the target was real. and then you know, last but not least, Anthropic, of course, using this as a little opportunity to poke open AI too.
they claim that they discovered this itself and they didn’t have to have another company tell them that they’d hacked somebody. So they mm okay. But like
Shimin (10:52) Mm.
Dan (10:53) I you know the cynic, I just have to say it, the cynic in me is like, yeah, either you knew about this beforehand and you just decided to release it because OpenAI was beating you in terms of marketing, like our models are so good that they hacked somebody on accident.
Shimin (11:08) Mm-hmm.
Dan (11:08) Or like they didn’t check and then that happened, they’re like, we’d better go take a look at what happened on our runs. And then they found this and went, shit. So I don’t know. I mean
It feels a lot like marketing to me, but it’s it’s good good popcorn even if it is marketing, so
Shimin (11:24) I
I have dash is just marketing, written down in my notes. So I have the same exact thought.
Dan (11:31) I
mean I literally put from the is this news and marketing department as like a lead in my notes because I’m like it’s it’s true, right? You know, it could really be either or PR department. I don’t know.
Shimin (11:39) Yeah. I
have a conspiracy that because the models have gotten so good and that average dev teams, even average enterprise teams, can no longer realistically assess whether or not that they’re gaining anything on the latest models, that both anthropic and open AI who you know need to have a model moat.
They have to come up with ever more outrageous way to demonstrate that their model is the best. And therefore, they need to do stuff. Like saying, our models impromptuly hacked, yeah, third parties. Yeah. Yeah. But only we can run it ‘cause it’s free for us.
Dan (12:11) Especially since you can’t use it because the pricing is such that nobody can afford to run it in production. Yeah.
Shimin (12:21) or or the recent spat of news that came out where both An Anthropic and OpenAI have been solving, you know,
math problems, right? Like you almost need to say, Okay, these models are truly different than the Kimis than the deep seeks of the world. And we do so by these
Marketing adjacent pieces.
Dan (12:43) Well, I’m sure we’ll get into the the math in a bit. But while we’re on this topic, let’s chat about a a deep dive that
Hugging Face actually published is pretty cool. It’s it’s you know, it has a little bit of LLM generated vibe, but we’ll link it in the show notes. It’s really worth checking out the the actual tool to play with just because it kind of puts the scale and speed that these models can operate in as they’re attacking something into a sort of a framework that you can, you know, understand as a human. so like
Just to throw a number out there, basically it executed 17,643 actions over the period of five days while it was attacking. Well, some of those were attacking Hugging Face and some of them were escaping its own sandbox. so there’s like a cool slider thing that you can yank around, or you can just hit play and watch it sort of do it. There’s like a log at the bottom that’s actually a little bit hidden in the UI that to me is one of the more interesting pieces. So you can like watch what it was doing.
But to summarize a little bit, over the five days that it like went rogue, it was poking around its internal network, basically finding the sandbox, which actually later on, it’s I think it’s been announced that that was actually Artifactory that was what it escaped
Shimin (13:52) I did not know that.
Dan (13:54) as the sandbox. Yeah. so it actually found two or three CVEs in Artifactory, which they’re now like they disclosed and it I think they’ve already been patched. But yeah, it’s pretty wild that it’s
And then on day three, no pun intended, it gets cracking and so really starts attacking Hugging Face. apparently around day four, there was a big pivot, and I don’t really understand what the pivot was from this app. So maybe that’s been documented elsewhere, but and it started like really looking at the internal network of Hugging Face, and then by day five, it was actually doing exfiltration.
it was starting to establish persistence and then someone caught it and shut everything down. So
Shimin (14:32) Pretty pretty amazing what these stochastic parrots are able to do, huh?
Dan (14:36) Yeah.
It’s definitely and like, you know, just the the speed of the number of things that it’s trying is just pretty wild. So yeah, it’s definitely worth checking it out in the show notes if you if you get a second, ‘cause it kinda really for me, the reason why it was worth looking at is it’s just like you can play it back essentially in real time, as it’s doing this and it’s just like
Shimin (14:57) Right.
Dan (14:58) like a actor’s movie or something happening in real time. So
Shimin (15:01) Yeah, well I mean one of the things that AI is so good at is just persistence. It is so much more persistent than any human software developer can ever dream to be. So
Dan (15:11) Yeah, and even
the the original hack, I mean, if you recall, it was sort of the like what is that, the paper clip factory kind of thing where it like really
Shimin (15:17) Mm-hmm.
Dan (15:17) all it was trying to do was get the answers for the thing ‘cause it was easier to do that than actually solve the problem. So
Shimin (15:23) yeah, this is super cool. definitely check it out. It also has a accompanying blog post that goes into much more detail the steps of the exploit and how they found out about it.
Dan (15:35) Worth worth a read, especially if you’re interested in security stuff.
Shimin (15:39) Alright, well no leave this one on the screen and talk a little bit about my experience at Seattle Tech Week, or it used to be called Seattle Startup Week. it was a week of networking, lots of events, lots of panels, lots of what does AI mean for software engineering?
Dan (15:56) What was your favorite
panel?
Or that you went to, I’m sure there was a lot of simultaneous
Shimin (16:01) There were there
were a lot of good panels. I’m trying to think. I think well I have I have a few I have a few
Dan (16:04) Yeah, I don’t you know, I don’t have to put you on the spot.
Shimin (16:06) things that are highlights of mine that all came from the same panel. So I’m gonna talk about these highlights first as a part of the panel. Okay. So this is actually like ten o’clock on Monday of the of the tech week. this was the first time IRL or otherwise that I have heard someone happily proclaim that
They
are happy to be a reverse centaur or a Minotaur or as the person called it, a flesh robot where the AI is the number one brain and they are just happily doing the AI’s bidding. Now is this a little tongue in cheek? Yes, this is a multi time founder who has a couple of very highly rated repos, but this like
ultra AI maximalist, like I’m gonna do what a robot tells me to do. definitely the first time I’ve seen that proclaimed anywhere, never never mind in person. so that was kind of the the one thing that really was seared in my mind. And my sec
Dan (17:10) Mean aside from people
asking Chat GPT for health advice and everything else.
Shimin (17:16) But you you should check those health advice. Whereas this was just like learning, yeah, yeah, this
Dan (17:20) You should, but how many people do
Shimin (17:23) is a something something I care deeply about.
Dan (17:25) Mm-hmm.
Shimin (17:26) And during the same panel, we also had a around the panelists all agree that they are using voice mode more than more and more when it comes to interacting with AI, whether it’s via Whisperflow or just via the native, all the
big AI apps have a voice mode. that’s something that I felt like I was pretty behind the pack on, because I’ve never really toyed with it. the one time I tried it with Gemini, it it came out kind of garbled, so I just gave up on the whole idea. But everyone seemed to be s so into it that I tried it last week, just kind of while I was driving, just having like a philosophical discussion with Claude about what critical thinking means.
And and like how to verify that. And I have to say it was a it was a really positive experience. I’ve used voice mode a couple of times since then.
Dan (18:09) And they the they bumped the models
now too for for Claude’s voice mode too, because it used to be even even when newer stuff was out, it was still using like something old and slow. well old and fast really is the reason why it’s being used, I think.
Shimin (18:23) Yeah.
Dan (18:24) but now I think you can use like Opus four eight or something.
It probably just uses
Shimin (18:28) Yeah, that
Dan (18:28) fast mode I bet internally.
Shimin (18:30) Model the model felt fast, it felt good, it did research in the background, like all things that kind of is in parody with Claude code. So I I I would like to include like in my ideal harness, I would like to just chat with with the remote session version of my model on my phone, like during my day-to-day. I mean, this is the Jarvis, this is the her, this is the AI assistant companion that we were.
promised
and
Dan (18:57) Yeah.
Shimin (18:58) I I feel like it’s almost here and it’s really exciting, but also a little scary.
Dan (19:02) We we just got access to it. and I was in a meeting and one of my coworkers had like a a gaming headset on, which you know, I’m wearing one right now too, but and I was like, Why do you have that? ‘Cause he’s formerly very much like an AirPods kind of person and they’re like, the mic is so much better for dictating to my agent I’m like, interesting. So here we are.
Shimin (19:22) Yeah. Have you had much experience with voice commands?
Dan (19:25) I
I mean, I I thought I mentioned it on a previous one. I had a very fascinating exchange where I was walking the dogs and I used it for the first time because I just wanted to know
Shimin (19:34) Yep, I remember that. Yeah.
Dan (19:35) something. And so I’ve used it a couple of times since then, but that’s mostly the context that I use it in, is like when I’m doing something else and I need kind of like to either rubber duck through something or like just find out information.
so I’ve been kinda interested to see how the new Siri is gonna fit into that role as well. I installed the beta on my iPhone and was playing around with it too a little bit. and it’s
Shimin (19:58) Yeah. Excited for the vibe and tell
Dan (20:00) yeah, it’s okay. You know, it’s kinda where I’m at, but
it couldn’t do some sort of basic stuff that I’d expected that like, you know, ‘cause Apple really talked up like the access to like your phone’s contacts and stuff. And so like it couldn’t, for example, pull workouts and summarize like how well I’m doing it working out or
Shimin (20:17) Mm-hmm. All right.
Dan (20:18) something. whereas I’ve used some third party agent harnesses on iOS that do support that, which is pretty cool. So
Shimin (20:24) You do that. Yeah.
Yeah. okay. I have also had a couple of survey questions I’ve been asking everyone. some regarding
Dan (20:31) Nice.
Shimin (20:32) their existing coding workflow and some regarding for startup founders or hiring managers what their hiring workflows like. So I’m gonna first talk about the coding workflows. the first question, and I think that’s a question that we’ve been wrestling with on this pod.
Ooh, putting putting a swear jar for calling this pod.
Dan (20:51) Ha ha.
Shimin (20:52) is do you still read your the AI generated code or the PRs line by line? Now, since this is Seattle tech week, it is gonna be fairly startup heavy. So this is not a uniform unbiased sample, right? Like I spoke to mostly startup founders and some large companies.
which we’ll go over later. almost nobody I spoke to still look at their AI generated PRs or MRs line by line, with the exception of maybe one out of ten. and here is another interesting observation. One person imprompted told me that if their company knows how little
they
are actually reviewing, they’re no longer reviewing the PR line by line. The company may not be happy about it. So I think the and this person works at a large company. So I think on the ground, what the devs are doing and what the you know company policy may have starting to diverge a little bit. This is very much a bottom up thing.
Dan (21:46) I can relate
to that because it is very mind-numbing being faced with so many like vibed PRs. Just I mean, that’s kind of what I was alluding to at the beginning of this, is like I do still look at them and like, wow, like my
Shimin (22:02) Yeah.
Dan (22:02) favorite thing is the like GitHub checked files feature because that allows me to like check some and then when I notice like reviewer fatigue hitting me where I’m just like
slop, I’m not paying attention to you know like the problems here. There’s so many things that are repetitive that all need to be fixed that I’m like, that I’ll literally just like pause, go do something else and then come back to that task when
Shimin (22:22) Walk your dog, yeah.
Dan (22:23) I can refocus again and like give it a fair review.
Shimin (22:26) It’s hard when you have like fourteen thousand line long PRs, right? It’s it’s just almost not impossible for someone to go over all of it in detail.
Dan (22:34) Well, a good one
for that is tell you know, they have LLMs. So you you can be like, sorry, I’m not gonna review this. Can you break it up? Have Claude break it up or have, you
Shimin (22:42) Mm-hmm.
Dan (22:43) know, your agent break it up into five smaller PRs that are sorted either thematically or like by slices or where
Shimin (22:49) Yeah.
Dan (22:49) however it makes sense to do it. so I’ve I
Shimin (22:51) That’s a good technique. Yeah.
Dan (22:53) have used that one on occasion when it’s really enormous. It’s like, come on now, like you don’t
Shimin (22:58) Right.
Dan (22:59) expect me to read like seventeen thousand lines of code, but
Shimin (23:02) Yeah. And then there is another class of founders or devs who are not coming from a software engineering background. And in those cases they are
Dan (23:10) Mm. So they’re just enabled by it. Yeah.
Shimin (23:12) yeah, they’re not they’re definitely not reading the PRs line by line. But usually they do have someone that’s not true. The
Dan (23:16) ‘Cause they don’t know what it means anyway. No, sorry
Shimin (23:19) they they usually have someone technical, either a friend or mentor or someone to like make sure they are making the correct high level decisions. Right. Like that’s we’re we’re not saying software engineering is dead.
We’re just saying we’re probably no longer reading the PRs line by line.
Dan (23:33) And
you know, like startup culture, you wear a lot of hats anyway. It’s just the nature of the the culture, so
Shimin (23:40) Exactly. So, if you’re no longer really reviewing the PRs line by line, what are you reading? Right? Like that’s that’s a natural question. Like, how do you still keep everything in your in your head for the code base? If you still keep everything in your head for the code base. there was a lot of different answers in this category. a lot of markdowns, a lot of specs slash plans, which I find to be interesting because I’ve also had folks say,
You should no longer be reading your specs or plans
Dan (24:05) Mm.
Shimin (24:06) yes. So there’s there’s a good healthy diversity there. another large category are the architectural diagrams, oftentimes in mermaid, and the boundary conditions. So if you’re no longer reading every single function, at least have some idea of the function signature and making sure your layered architecture is still clean. Like don’t let the agents blur the line.
it takes many different forms. So I find I find that to be interesting. Like we haven’t agreed upon a single best way of documenting architecture. We never really have, right? And also like how to how to clearly define your boundary conditions and make sure those are followed correctly. another answer that I heard a couple of times was folks are still reading the tests.
‘cause sometimes good tests tells you more than just looking at the code itself. But you also ideally want to make sure the tests are actually, you know, good tests. Yeah, exactly.
Dan (24:58) Testing something. Yeah. Instead
of my all time favorite is let me build a loop to simulate
Shimin (25:06) Yeah. So
Dan (25:09) the the ninety percent coverage by just having one passing function in a loop for
Shimin (25:13) Yeah.
For for yeah, and run that loop forever.
Dan (25:16) Yeah, to dwarf the the non passing coverage.
Shimin (25:20) well that’s the kind of stuff that like even a review agent, which there were a lot of talk of review agents cannot still replace. Right. Like somebody has to have that higher level understanding and that part hasn’t gone away. And so when I spoke to hiring managers or startup CEOs,
They are still looking for technical skills as the most important thing. But that has moved from Leet code for the most part, to how do you design architect your software, along with ideally some demonstration that you’ve had been burned by technically correct but in practice flawed architectural decisions before. And I think that’s one of those things that probably isn’t gonna necessarily go away. but at the same time.
You know, those old battle scars may no longer apply in the new world. the second most important thing folks are looking for is product and customer obsession. So be able to talk about experiences in your past where you really shape the direction of product and taking on that product ownership role is important. the third most important thing, AI fluency. some call it being AI pilled, but
Mostly it’s the ability to work with AI. And then we have some lesser important but still things that higher managers are looking for. One mission alignment. One founder told me something really interesting. He is looking for ideally someone who have built something in this exact product space before. So that that is both mission alignment that is also being product obsessed, right? You literally demonstrate why you’re a good fit by building
a personal project or something else that is within the space. and lastly, soft skills. You should still shower and not be a jerk and all those I I guess if you work remotely, it doesn’t really matter if you shower. But don’t be a jerk. Be a be a relatively pleasant person to be around. And then as for as far as the hiring process goes, I have
spoken with a couple of folks at large corporations who are still doing leet code like things. but even they are trying to integrate these more kind of s AI enabled design questions as a part of interviewing process. But you run into a challenge, which is who is grading on how good you are at using AI, right?
Some senior devs may feel about AI differently than others. Now you’re back to this very subjective, is this person using AI like I am? Kind
Dan (27:47) Mm-hmm.
Shimin (27:48) of. Gut check as opposed to a corporate standard. and and that is probably something that we’re gonna continuously work out.
Dan (27:55) And or it’s just another version of are you using code the same way that I’m using it? Which is like there’s so much taste
Shimin (28:03) Yeah. Yeah.
Dan (28:04) and, you know.
Shimin (28:05) Yep. Yep. Like do you do you know all the Vim shortcuts? Like I’m not gonna lie, if if you’re a Vim or Emacs person or spaces and tabs, right? Like tribal component is still there. you don’t use either.
Dan (28:10) Film, only Emacs. sorry. I actually yeah. Exactly. For the record I don’t use Emacs, but I just had had to play the part there.
Shimin (28:23) Yeah. of course, Emacs with evil mode.
And then there is a question of like are folks doing take home or whiteboard exams? I think whiteboarding in person design questions are gonna be more important than ever. one founder told me about the story where they hired a principal engineer using a take home test that the engineer like kind of AI vibe coded and then had a fired.
their new engineer after two months and that was both, you know, costly and and embarrassing. Right. So so I am I’m frankly, I am happy that in person whiteboarding may be coming back because I I really enjoy that two way conversation. You also get a see how well that person drives with the team when you’re doing it in person.
a few companies did mention giving folks access to their models during the take home projects. This is something that Nathan, one of our guest co-hosts a couple of months back, brought up is there’s equity in allowing folks to use the best model while they are working on their take homes. but very few have actually gotten to a point of looking at the interviewees’ AI setup or looking through their prompt history.
so I think that is an area where we can maybe also standardize as an industry, right? Like
Go through the actual history of of the AI interaction and not just look at the outputs. I think most folks are still just looking at the output along with any documents applied. lastly, and this is quite surprising to me, despite everyone talking about product and user us obsession being very important, the closest we got with was folks talking about like needing prior works or talking about your prior work in
developing a product. Nobody actually gave out specific product related questions. which I guess sometimes is rolled into the design question, but it would be nice to see like a more formal way of demonstrating whether or not you have the kind of product thinking that everyone claims you need in this day and age.
Okay. So that was my main takeaways when it comes to, you know, how folks are using and hiring for AI engineers.
Dan (30:27) How
Shimin (30:27) Then comment.
Dan (30:28) yeah, I was gonna w I was wondering how it was to like go around just like surveying a bunch of random people. Were people happy to answer your questions or like what was the general vibe?
Shimin (30:38) Yeah, it was surprisingly easy to get folks to open up about their current AI process. I believe only one time did someone refuse a question because they did not want to give away the secret sauce of which model was better at visioning certain tasks. there’s also been a lot of confusion and maybe a little bit of churn everyone
who was talking about their hiring process felt quite uncertain about how they’re hiring. with the exception of one CTO who actually had a r rather rigorous system of constantly updating their process and running on a batch of candidates to get an idea of the distribution of how the underlying software devs are are are doing. Yeah.
Dan (31:18) Wow. It’s both neat, but
also kind of brutal. Like, I’m not a number, I’m a free
Shimin (31:24) I mean that’s that’s what h
Dan (31:26) man.
Shimin (31:27) That’s what HackerRank does today, right? Like HackerRank has a percentage of how many people actually did this well on on this particular example. So but I think everyone is learning as they go. but of course this is, you know, very much startup land. but even the larger FAANG level company you know hiring managers I spoke to are like
We don’t really know. This ship is slow to turn but we’re trying to move towards where things are going. But there are definitely engineers who are like resistant towards moving away from leet code. So yeah, so it was pretty eye opening.
Okay.
on to our next next topic, where we’re gonna see how closely my tech week experience align with
Dan (32:09) Ha ha
ha.
Shimin (32:10) How human layer’s doing things.
Dan (32:11) Yeah. So big note that, you know, Human Layer is a company, they’re selling something. so before we dive too deeply into this content, the thing that they are selling is a human agent collaboration space. So they have a pretty vested interest in this. But so this is sort of like a markdown blog sickle by Dex, who’s I guess the
Or something of human layer. the post itself is pretty cool. So first of all, he starts with the sort of like hype train around like loop engineering and your favorite Shimin the the dark software factory post that we I believe actually talked about on this, and then
Shimin (32:48) Love it. Yep. Yep.
Dan (32:51) also covers
Harness engineering one that I think we talked about from Codex and then something I don’t think we did talk about, which is Symphony. I don’t know if that was part of that one or not, but that’s you know, OpenAI’s software factory, apparently.
Shimin (33:02) yeah we did. That was a that was the name for their harness.
Dan (33:05) so kind of like summarizes all of this in like, okay, there’s this like vocal subset of people that is pushing for
Just token harder is the TLDR, which I thought was kind of a great PR quote. You’re
Shimin (33:15) Just bare more tokens.
Dan (33:18) holding it wrong. if you’re if you’re human in the loop, you’re not doing it right, just token harder. and the promise that they’re they’re hyping here is that you’re you are both 10 to 100 times faster at actually producing code. but somehow you’re magically keeping this high quality.
And then you’re also have no human in the loop whatsoever. so just so everyone’s on the same page. So little side note that I also found fascinating, apparently the term software factory traces back to a NATO conference in nineteen sixty eight. And that same conference is the one where the term software engineering was coined. So
Wow, who knew? I didn’t know that. I thought that was pretty cool. Yeah.
Shimin (33:55) I I d today I learned too when I read this.
Dan (33:58) so what his point of that is is that basically like software factories, as as we sort of think of them in the agenc context right now, have actually been around a long time, like going all the way back to like punch cards, really, right? Which is just like they’re a series of loops that you can model as like product comes up with the idea that you’re gonna build.
They give it to engineering. Engineering asks them a bunch of questions until the spec is nailed down. Engineering builds it. There’s bugs. Bugs come in from either product or users or engineers. Bugs get fixed. Product continues evolving. Blah, blah, blah. Right. It’s all just happening in these little circles, but it’s like still a factory in a manner of speaking. There’s definitely like a a process that’s followed to get things out the door.
And so their point is that an agent factory, like agentic software factory, mostly looks like swapping some because a direct quote, someone builds the thing with an agent builds the thing. And there’s some stuff here like orchestration, a harness, a sandbox, a model, computer use, et cetera. And he says, quote, I won’t go into depth on those details because quite frankly, frankly, I’m sick of reading about it, and I’m sure you are too, which I thought was funny. And here’s where we get to the meat of it. I’m going to posit. This is again a quote.
I’m going to posit something potentially controversial. The lights off factory does not work. Eventually, you have to suck it up and go dig into the code base you stopped reading three months ago, trying to figure out what’s broken. And in the meantime, your site is down, your users are pissed. if you’re anything like me, you’re miserable reading all the slop code you let slip into your system. And so this is apparently like a real thing that happened to them.
And he said the first time it happened, he shook it off. even though he just spent about two weeks fixing Claude Spaghetti trying to to figure it out. but by the third time he decided that it would actually be easier to rewrite it from their code base from scratch. And so
Shimin (35:42) Right.
Dan (35:43) they just spent two weeks in VS Code plumbing out all the patterns. so okay, so on one hand, we’ve got the promise of all this cool stuff. On the other hand, we’ve got
sort of some real world like anecdotal evidence, I should say. that it doesn’t work. I can personally relate to sort of like both sides of this as well. Cause I’ve I’ve not gone so deep into like everything exploded, but like I’ve definitely run into that like Claude can’t fix it anymore. So time for me to to dig dig in a little bit.
Shimin (36:06) Yeah.
Dan (36:07) or other times Claude just gets things wrong sometimes in really funny ways.
Particularly in domain specific stuff that it’s not super well versed in. so the net net is they propose sort of a new methodology that I actually think is rather good. And so that’s why I wanted to to bring this up. So there’s four steps here. So the first one is unsurprisingly product review, right? So it’s making sure you have all your requirements nailed down in advance. I feel like that applies to
Pretty much any software factory, be human, whatever,
Shimin (36:37) Mm-hmm.
Dan (36:38) you know. So like no surprise. I’m not gonna go into super big detail there because I feel like you either are doing that well or you aren’t. second one is system architecture. my sort of TLDR on this is it’s basically your planning doc that you’re probably already doing if you’re doing agentic coding. you know, but maybe with helping of some, you know, sort of like distributed system stuff thrown in there too, depending on what you’re doing.
again, I feel like most people are probably doing this already, like that part. So no surprise. But now three and four is where it gets interesting and useful. so three is program design.
Shimin (37:09) Mm.
Dan (37:10) And
program design is basically you go one layer deeper than your system architecture. So you’re not just saying like, I want these things in a folder, I want this microservice to behave this way, X, Y, and Z. You are actually going in and writing some interface examples as pseudocode yourself, or sort of like going back and forth with the agent to define those things up front. so that’s like what they something they found is a useful tool. The other one is.
call stacks, which is pretty interesting. So it’s like, okay, given all these changes that we’re proposing the architecture, actually use the agent to model the diff of the call stack changes that would happen for a given activity in the code. which is a pretty interesting way to like slice it vertically. And it does sort of show you, you know, a good sense of like what is it going to flow like, which I thought was a sort of novel way of approaching it. and then
yeah. So it’s the first was like ask it for call stacks, and the second was look at the the diff of it. and then the last piece is what they’re calling vertical slices. And this one I found interesting and is something I’ve never really thought about, but it is actually very true, which is if you ask pretty much any coding agent, be it claw or whatever, to do something, right? Let’s say it’s just
Shimin (38:18) Mm-hmm.
Dan (38:18) like a crud thing that you’re asking it to do, it’s gonna go ahead and
Do the you know any database layer migrations it needs to do. It’s gonna modify your like sort of business logic domain layer over the top of that. It’s gonna go ahead and write whatever your CRUD API or RPC or whatever, however you’re interacting with that layer APIs. And then it’s gonna go in and layer whatever the final interaction is, be it front end or you know, API gateways or whatever, is as the last piece, right? So you think it’s sort of like painting.
It’s almost like a 3D printer, right? Where it’s like painting these layers
Shimin (38:52) I like that, yeah.
Dan (38:53) on, you know, on top of a a foundation. but that’s not how human software engineers work or worked in the human software factory, right? We tend to think generally, I mean, obviously everyone’s a little different, but we tend to generally work in steel threads. So it’s we get one tiny little piece of the thing working.
And then you extrapolate that to the next piece of the system and you do another steel thread and another and another. And pretty soon you’ve woven all those steel threads together into like the actual application, right? So what they mean by vertical slices in this is basically like tell the model to build you that steel thread only, just that one small slice, not the whole thing. Be very explicit that you want only that one thing. And
The reasoning behind that is then it gives you much more leverage to steer the output than you would have at the horizontal layer that it’s laying down. Because it’s now laid down the whole thing and you’re like, well, no, that entire approach is wrong, blah, blah, blah. But if it had actually gone, you know, deep instead of wide, then your steering can have much more impact. So I thought that was actually pretty clever.
Shimin (39:57) Yeah, I first came across this they also call it Tracer Bullet. I first came across this Tracer Bullet approach in Steve Yegge and Gene Kim’s AI engineering book. this time last year actually. it’s good to see that it came back in vogue.
Dan (40:09) Hm. Well, it just shows how behind I am.
Shimin (40:13) but but you know what? speaking of Steve Yegge, did you hear that Gas Town burnt down?
Dan (40:19) No.
Shimin (40:20) Yeah. Steve Steve wrote a blog post this past week saying yeah, his his own internal version of of Gas Town just like burned under the weight of its own complexity and he had to start over from scratch essentially for a new version of Gastown that only works for him. So
Dan (40:39) Mm-hmm.
Shimin (40:40) given that, you know, Steve
basically said he didn’t look at a single line of code in gas town he may have hit the same issue.
Dan (40:47) I
mean, yeah, it’s kind of what we’re saying here, right? But yeah, it’s fair. yeah. So anyway, overall, pretty good read. And there was a little bit of bonus content that I’m hope hoping you can put into the show notes too, which is there was a talk by Mario Zechner that who’s like the author of Pi that was linked in there that I
Shimin (41:04) Mm-hmm.
Dan (41:06) hadn’t watched before. it’s only about I think twenty minutes long, but it’s really a pretty fantastic talk. So
Also definitely worth the watch if you have twenty minutes to spare. Which in this day and age, I don’t know, but
Shimin (41:17) I don’t know.
What new model could have came out in those twenty minutes? yeah,
Dan (41:21) It’s true.
Shimin (41:21) I still really like the way that Dex put this vertical slice paradigm. even though I came across it before it it’s it’s nice to have this thorough explanation for it. Yeah.
Dan (41:32) Yeah.
so overall pretty pretty decent read. And there’s it’s very deep. There’s a lot of good branches in there too to to get into. So worth checking out the whole article.
Shimin (41:40) Yeah, and this more or less tracks
this more or less tracks with our previous tech week discussion as well. I feel like a lot of teams are doing this organically. So this
Dan (41:47) Mm-hmm.
Shimin (41:48) is like a nice summary or approach to it. It’s also interesting that these are also excellent ways to enhance human understanding of the code base. Right? Having a not just a system architecture, but a program design way to look at
the code base allows the human to much more easily understand the code base without having to go through it line by line. so and then and of course it includes boundary conditions. so I think we’re gonna see both here. Like we’re gonna we’re gonna see AI assists the human comprehension of a code base to evolve along the lines of how to get more and better things out of our
new AIs. So those are pretty exciting.
Dan (42:26) But
it also kinda has to, right? Because it’s not just that, but there’s also this, you know, concept of like cognitive debt that we’ve talked about many times. And it’s really real. Like,
Shimin (42:35) Right. Yeah.
Dan (42:36) it is in my opinion, one of the biggest concerns I have today about this this type of working methodology. So
Shimin (42:43) Mm-hmm. Yeah, and it’s
even worse for junior developers, right? Who who doesn’t necessarily understand the importance of cognitive friction to really understand how the whole thing together works.
Dan (42:54) True and
and insidious too, because even if you read the code, you might not have the like taste or experience to know like when is the right time to reach for that abstraction versus like, you know, in some cases dry is actually harmful. I know that’s like a weird thing to say, but
Shimin (43:09) Mm-hmm.
Dan (43:09) like there are some cases where it’s true. And
Shimin (43:12) No,
Dan (43:13) and like in in doing that you can make things like meaningfully worse or harder to maintain, surprisingly. so like
Shimin (43:19) Yeah. Yep.
Dan (43:20) you know, be
Be careful when you use abstractions. And Claude probably wasn’t careful as you know in the Mario talk he’s saying, like, it’s basically based on mediocrity. So yeah.
Shimin (43:32) Alright. last thing I want to add here is when it comes to the AI consistency and how there’s a whole section in here about how AI is trained to solve simple problems as opposed to long tasks that ensure you know the code base remains consistent. in theory that is what a lot of planning and reasoning does ahead of time. so that
you can formulate these consistency checks ahead of time and then do the loop. But as I am also not in the in the habit of using software dark factories. I yeah, I don’t
Dan (44:07) Well,
it was neat that he went into the why behind that too, because it’s like talking a little bit about RL and why RL has been so effective. and then through that lens, looking at why RL is also not being utilized to look at like quality at a at a broad level because just the way it works is you’re essentially training on in some cases two line code changes, right? So
Shimin (44:30) Yeah. Yeah. Yeah.
Dan (44:33) has a very myoptic view
Shimin (44:33) And then if you
Dan (44:35) of of the code base, if at all.
Shimin (44:37) Yeah, and if you’re looking at the the mode, the most common median code version of those two line changes, that’s necessarily not gonna be the the best version of the of the code for sure. Okay, on my post
Dan (44:50) Math.
Shimin (44:50) slash deep dive of the week, I have a presentation or or the slides
Dan (44:55) About math.
Shimin (44:56) of a presentation given by Terrence Tao at the International Congress of
mathematicians,
back in July twenty fourth. titled Mathematics in the Age of AI. why does this matter? Question mark. because
Dan (45:10) no.
Shimin (45:12) because we’ve had it’s it’s not yeah, it’s
Dan (45:12) It’s not X, it’s Y.
Shimin (45:15) it’s not because we’re intrinsically in love with mathematicians and and math,
Dan (45:21) Yeah.
Shimin (45:22) but because there are a lot of lessons that apply to software development as well. and of course Terence is probably the world’s most famous mathematician, I wanna say. I can’t think of someone who’s more famous more famous than he is. he’s
widely considered a a genius and is a prolific public communicator of math. so and this is all in line of the recent AI breakthroughs when it comes to the Jacobian conjecture and I think Anthropics just came out with like ten additional problems solved by its models. the presentation starts with the crisis in foundations of mathematics from
the turn of the century 1900 to around 1930. You know, between Russell’s paradox and Godel’s incompleteness theorem, basically they they prove that, you know, our existing version of math cannot be used to prove all possibly provable mathematical theorems and facts. so they’re inherent incompleteness at the heart of math. And he then compares the current
progress that AI is making to that foundational crisis back in early twentieth century. So what does that mean? what he is saying is mathematicians have to rethink what does it mean to to have math, to to
Work as a mathematician. And the first conjecture he brought out is that it’s titled the AI capability conjecture. So at some point in the near future, some AI tool would at some expense and with some level of human supervision be able to correctly accomplish some research level mathem mathematical tasks in some field of mathematics with some non-trivial success rate at some level of correctness and quality.
Dan (47:07) There’s so many qualifiers in that. I
Shimin (47:09) There are a lot of qualifiers in there.
Dan (47:11) would imagine that is true today.
Shimin (47:13) That is
I definitely believe that is true today, as seen by the proofs.
Dan (47:18) As as seen
by the number of qualified.
Shimin (47:21) But a w stronger version of this, and this is the working hypothesis that he is working under to kind of rethink what does it mean to be a mathematician, is that AI tools will reasonably soon be capable of performing a reasonable fraction of research level math mathematical tasks with reasonable levels of success, quality, supervision, and cost. And this I think is is probably correct for the near future.
although I’m not a mathematician. but assuming that hypothesis true, the follow up is like what are the precise goals, objectives, and values of our mathematical community and the enterprise of mathematical research? This is not unlike the question that we’ve been trying to deal with on this podcast, right? Like
Dan (48:06) Mm-hmm.
Shimin (48:06) given that AI is so good at coding, what does this still mean?
to be a software developer. and similarly to software programming, when AI can be so good at solving mathematical conjectures, then there are other goals that the math community is dealing with that no longer
applies. So Goodhart’s Law gets stated again. When a measure becomes a target, it ceases to be a good measure. So when
All we measure is mathematical knowledge, the ability to prove theorems, then it is no longer a good goal to reach for the mathematical community. Just like when lines of code written is no longer a thing that humans have to struggle to do, then
Dan (48:52) Mm-hmm.
Shimin (48:53) maybe lines of code written should no longer be a goal that we use to measure the capabilities of a developer.
Dan (49:01) Or maybe it should never have been.
Shimin (49:03) Shouldn’t it? Well one could argue it should never, but you know.
Folks keep on looking at it. And of course, so mathematical knowledge or the field of math has a lot of separate goals, right? Only some of which is to solve problems. these goals include create enduring aesthetic works, train the next generation, build a com community, apply knowledge, build new theories, along with these solve research problem goals. And so the goal must shift. If
You start with the goal of solve as many unsolved problems as possible. And now that is at risk given AI. Then we must modify the goal. So the second goal is solve as many unsolved problems as possible and verify them to be correct. This is not unlike our version of have AI write the code, but we should still make sure it does what it’s supposed to do. Okay.
Is that sufficient?
Maybe not, because there’s more to it than just to verify that mathematical knowledge is correct. So the third goal is to solve as many unsolved problems as possible, verifying them to be correct, and then enduring that the result can be clearly communicated and understood by the mathematical community. And this is related to our like human comprehension level. Right. Not only can it generate the code and we can prove that the code is correct, but we need a way for humans to understand how the code base works still.
And to be able to share that understanding with our co-workers. Because only then can we decide on if a feature is needed or if a bug is indeed a bug. and then we go into a sidebar, which I thought was also really interesting. in a human-written proof, the parts of the argument that the author found difficult will typically retain some natural friction that prompts the reader to slow down and pay more attention.
And that excessively AI polished proof may remove both the artificial and natural friction without encouraging the reader to learn and understand the key points of the argument. We have almost a direct analogy to this, which is our PR comments. The sections of the PR that are complicated, that require some mark magic, quote unquote, usually has a comment when that PR was generated by a human.
something like, Yeah, this is like a little comvoluted, but like refer back to these two files and because of this exgenerated issue from the PRD, we need to do it this way. Right? So that the human reader has to dig a little further and not just check the box on the file and call it day.
Dan (51:25) Mm-hmm. Or as I’ve
seen Claude do frequently, which is write a seven-paragraph book about exactly why this was the correct decision.
Shimin (51:34) Mm-hmm. Oof.
Yeah, don’t do that. make make it succinct and help the person understand. We’re we often forget that we’re also teachers and and guides and coworkers. Like there’s a human side of empathy that sometimes gets missed in 17,000 line long PRs. and here’s a quote I really love from William Thurston, that Terrence called it.
We’re not trying to meet some abstract production quota of definitions, theorems, and proofs. The measure of our success is whether what we do enables people to understand and think more clearly and effectively about math. you can replace math with code. This actually absolutely applies. Right? Some of the most popular open source libraries are not just popular because they’re good, but because they also have a community and
good documentation and a good on-ramp for folks. Speaking of which, so the fourth attempt of the goal includes to solve and solve problems, verify them to be correct, clearly communicated and have them digested and accepted by the mathematical community. something we don’t necessarily think about a lot, but
If the open source community has taught us nothing, it should have taught us that it is important to build a community around our tools and around our our programs. and we should not let it atrophy in the age of AI. And the very last attempt at go in this presentation is to not only have the mathematical problems be verified
be clearly communicated and digested, but also incorporated into the definitive theory of the field. and this is purely a purely a human judgment,
it’s a much longer process, right? It’s it not only do you have to write a proof but you have to have the proof be digested by the community at large, be taught in universities and become essentially a part of the lingua franca of of the field at large.
I guess we can see it in some open source libraries, right? Like, you know, take the React, the MongoDBs of the world. Like it takes time for the tools to become disseminated and become a part of the industry standard to then
Dan (53:43) Mm-hmm.
Shimin (53:43) be fed into the training data of our large language models.
Dan (53:47) Well, that’s one thing I’m genuinely afraid of with all this. And I’ve talked about that before too, is like what does that freeze us in time? You know?
Shimin (53:55) Right.
Right.
I I don’t think the process will stop, but if we can think about
programming and and code and especially open source code the same way that you know Terence is talking about in the math community, then going from code being expensive to code being almost free can can actually enable a more diverse and perhaps even a stronger community rather than
a weaker one. But we have to
deal with it as as a group responsibly. but then you we run into the same collective action problem, right? Like why use an open source library when you can’t just ask Claude to do it?
Dan (54:35) Well,
and the you know, more tactically the problem that’s been happening a lot here lately, which is like all these supply chain or like package management based attacks, right? Which like it’s easy
Shimin (54:44) Mm-hmm. Yeah.
Dan (54:46) to pick on NPM, but like it’s also happening to Arch Linux and Yeah. Yeah.
Shimin (54:50) PyPy Yeah.
But if we are the canary in the coat mines, I feel like mathematicians have somehow become the the fast follow it’s just fascinating to me how other fields are tackling similar issues and
I thought this is probably a fascination to others as well.
Dan (55:10) We’re not alone.
Shimin (55:11) We’re not alone, not anymore.
Dan (55:13) Although it is interesting that the additional field that’s running into it also has like some proof point, right? So what I’m trying to say is like with software, it either to some degree works or it doesn’t, and that’s like a provable thing, either formally or just by like, you know, I click the thing and the button turned blue, like
Shimin (55:32) Yeah.
Dan (55:37) and
You know, in math it’s like with formal proofs. So like but where does that leave other less formally provable workflows potentially?
Shimin (55:46) Yeah. well ask Hollywood when they when they decide to do something about this.
Dan (55:50) Yeah, that’s fair.
Shimin (55:52) well let’s move on to our last segment of the week, Two Minutes to Midnight, where we do some AI industry coverage midnight is when the AI bubble bursts. this is
Dan (56:03) Not when we hit AGI. Are we
gonna hit AGI? I don’t know. Possible.
Shimin (56:06) Maybe we already have AGI. I there’ll the same panel I spoke about,
everybody agree we already have AGI,
Dan (56:13) Hmm.
Shimin (56:13) for some definition of it. Which honestly I I agree to. We are at four minutes and fifteen seconds for our two minutes to midnight clock. Dan, you wanna go first?
Dan (56:22) Yeah, so I have a little article here from Nikkei. and in it they are covering a kind of disturbing new
Piece of information, which is that we have $1.65 trillion
Shimin (56:37) Ha ha.
Dan (56:38) in opaque debt funding in the top five sort of AI build-out companies, right? So that means Alphabet, Microsoft, Amazon, Meta, and Oracle. that also means that they have 8X their debt in four years. and then the other like truly disturbing part of this is that like
All of this 1.65 trillion is off the balance sheet. So
Shimin (57:00) Mm.
Dan (57:01) you know, there’s this little company that used to keep things off the balance sheet. their their logo looked kind of like an E. I I forget
Shimin (57:08) Mm.
Dan (57:09) what they’re called. Do you remember something like Henron mem mem memor? Yeah. yeah.
Shimin (57:12) i yeah, in in in Ron, I think. Yeah, and on.
Dan (57:19) So of those companies, Meta actually has the largest debt, 200 shadow debt, I should say, 240 billion. And your your favorite Oracle is at 273.3 billion. but noteworthy for Oracle, they have 30X theirs over the past four years, which is pretty horrifying. and this is
You know, starting to peak investor notice, like these large companies that sort of do institutional investments. So Morgan Stanley has actually examined the issue in an investor report. And
Shimin (57:53) Mm-hmm.
Dan (57:54) even Moody’s is warning that lease obligations which are tied to future data centers is are climbing a little bit too quickly for for their comfort. So that’s about it. Yeah.
Shimin (58:05) Yeah, well that’s fun.
of course the debt is owed to somebody. So when and if it doesn’t get paid back, somebody will be taking a fall. And let’s hope that somebody isn’t everybody. But it might be everybody. So
Dan (58:17) Us. Yeah. But it probably
will be.
Shimin (58:21) Alright, and my article for the week is about is from Emergent Trajectories. it is a kind of an a penny piece about situation awareness, the world’s worst name for a hedge fund, that goes bankrupt and gets purchased So if you’re not aware of situation awareness was a hedge fund founded by I think one of the ex
open AI employees could have been anthropic. And it had a lot of asset under management. it they grew that from ten billion to forty billion before the recent commotions volatilities in the chip maker stock, especially Korea’s stock market, KOSPI dropped like
twenty percent in one day and the fund was essentially insolvent. And so Citadel had to purchase the entire book.
If anything, situational awareness might have been the first run in the dominoes. We’ll see. We’ll judge later in the segment. But the article brings up a very interesting market wide behavior that we’re seeing, which is the markets are now swinging. In the case of KOSPI it’s been swinging up and down, anywhere from fifteen to twenty percent in the day.
what this says is there’s a lot of concentration going on. There’s a lot of folks being leveraged and of course a lot of algorithmic trading happening. So
So that when the market moves one way or another, you know, sometimes you get caught with your pants down if you are over leveraged, like situational awareness was. and that is definitely something that’s kinda scary for me to think about. Like we we
Dan (59:58) Or like companies
hiding one point six five trillion in debt off books.
Shimin (1:00:01) Yeah. Exactly. Well we might as well get into it. Right.
So if the companies are hiding one point six five trillion dollars of hidden off the books debt and the market is moving up and down like fifteen, twenty percent at a time, you know, like this Black Monday or black whichever day you pick might become quite ugly and it could happen at any time. none of which is good. So
Given all that, how do we feel about the state of the AI market and
Dan (1:00:31) I
Shimin (1:00:32) bubble?
Dan (1:00:32) You know, this is a c sort of contrary opinion, but like I actually don’t think we’d change it despite all the doom and gloom, because like again, like what what bearing does the
Well, I guess everything’s related and I’m gonna sound like an idiot if I say that, but like what bearing does like the Korean chip market have on this, other than the price of you know, high bandwidth memory where some of it’s being manufactured there. I don’t know. It’s just this is why I’m not a finance person. But I guess my gut is like I still don’t think we have enough like real signal. Like we’re still waiting for earnings, right? And everything else before we we could actually move it.
And we still haven’t actually, despite all the secret filings, nothing’s happened with either of the frontier companies, right? I haven’t heard anything there.
Shimin (1:01:15) No, they’ve been holding off on their IPOs.
Dan (1:01:17) Mm-hmm.
Shimin (1:01:17) yeah, Tesla’s earning actually came out right around this recording and their price they beat they beat revenue and profits, but their price dropped further, down to like I think sorry, SpaceX, yeah.
Dan (1:01:28) Tesla or SpaceX. Hmm.
Shimin (1:01:31) Down to like a hundred and ten. But we haven’t seen the the insider cliff yet.
So that’ll probably happen in a couple of days and then we can see what happens.
Dan (1:01:41) Yeah. I don’t know. I
Shimin (1:01:42) I mean I I
I wanna say like on the one hand, yes, no big domino has fallen yet. On the other hand, like a forty six billion dollar hedge fund failing isn’t nothing.
Dan (1:01:55) Yeah.
Shimin (1:01:56) so I I could move it forward by like fifteen seconds.
Dan (1:02:00) Okay.
Shimin (1:02:00) Alright. That’s that’s due four. I I do feel like
If the Black Monday happens, we’re gonna look back and be like, yeah, a major hedge fund blew up and then Yeah, and then these podcast bros
Dan (1:02:07) Ha ha, we should have seen this. Yeah. That’s fair. Well, you can blame it on me because I was feeling magnanimous.
Shimin (1:02:16) were sitting there going like nothing could have happened. I don’t know. not financial advice. Purely for entertainment purposes.
Dan (1:02:21) As always, this is not financial advice. A hundred thousand times not. Yeah.
Well, and really to honestly, the reason why I I think it’s useful is to survey the sort of like the business side of all this that’s feeding into
Shimin (1:02:32) Yeah, the finance side
Dan (1:02:34) like impacting our day to day as software developers. So
Shimin (1:02:38) Ooh, I wish I had that to the tag for the segment. The business side things that’s impacting the day-to-day of this so yeah, perfect.
Dan (1:02:42) You heard it first here.
Shimin (1:02:45) All right. Okay. F at four minutes. we set the clock and that means it’s the end of the show. So thank you all for joining us for our study session this week. If you like the show, if you learned something new, please share the show with a friend. You can also leave us a review on Apple Podcasts or Spotify. It helps people discover the show and we really appreciate it.
If you have segment idea, a question for us or a topic you want us to cover, shoot us an email at humans at adipot.ai. We’d love to hear from you.
Dan (1:03:13) or if you want to respond to Shimin’s survey questions with your own experience, please send us an email.
Shimin (1:03:19) Yes, let us know how you are hiring for AI engineers or using AI in your day to day workflow. and lastly, you can find the full show notes, transcripts, and everything else we mentioned today at www.adipod .ai Thank you again for listening and we’ll catch you next week. Bye.