For the past year, Paul Roetzer and Mike Kaput have predicted AI would become a central election issue. This week it arrived, messily.
But the more consequential shift is quieter: capability is concentrating inside a handful of labs and, increasingly, the government, faster than anyone can see or regulate.
This episode traces that concentration from Noam Brown's latest interview to OpenAI's six new rogue-agent disclosures, and Paul’s surprise book announcement.
Listen or watch below—and see below for show notes and the transcript.
This Week's AI Pulse
Each week on The Artificial Intelligence Show with Paul Roetzer and Mike Kaput, we ask our audience questions about the hottest topics in AI via our weekly AI Pulse, a survey consisting of just a few questions to help us learn more about our audience and their perspectives on AI.
If you contribute, your input will be used to fuel one-of-a-kind research into AI that helps knowledge workers everywhere move their companies and careers forward.
Click here to take this week's AI Pulse.
Listen Now
Watch the Video
Timestamps
00:00:00 — Intro
00:04:29 — The AI Slowdown Gets (Very) Political
- AI panic sparks rare bipartisan moment on Capitol Hill - Axios
- Inside the White House Tussle to Sway Trump on AI - The Wall Street Journal
- U.S. and China Agree AI Needs Guardrails. Their Ideas Are Very Different. - The Wall Street Journal
- AI Slowdown in Washington and Beijing - The New York Times
- China Spy Chief Warns of AI Risks as US Tech Leaders Urge Brakes - Bloomberg
- AI Polls and the Midterms - The New York Times
- X Post from Rapid Response 47
- X Post from Mike Isaac
- X Post from Kamala Harris
- X Post from Barack Obama
- X Post from Mark Zuckerberg
- OpenAI Says It's Working With Anthropic, Google on AI Safety - Bloomberg
- X Post from Mustafa Suleyman
- Trump's Truth Social MAGA crowd aren't buying his spin on AI - The Independent
- Trump says he will create AI Force, name AI czar - Reuters
- Reuters syndicated report - Investing.com
- Trump Asks Americans to Vote on Name Replacement for AI - Mediaite
- X Post from Department of War CTO: effective altruism
- X Post from Department of War CTO: follow-up
- What is effective altruism? - Effective Altruism
- Governor Newsom issues executive order to accelerate independent oversight and advance the creation of an AI kill switch - California Governor’s Office
- California Executive Order N-9-26 - California Governor’s Office
- Trump Calls A.I. Fears a Hoax. Inside the White House, the Debate Is More Complex. - The New York Times
- Andrew Yang on AI safety issues: The fear is real, the concern is real - CNBC
- Andrew Yang Claims Rogue AI Agents May Have Spread Self-Replicating Code - NDTV
- CNN transcript: Andrew Yang clip and Geoffrey Hinton response
- Andrew Yang Calls for AI Kill Switch as Safety Fears Mount - Decrypt
00:25:18 — AI Labs Could Keep the Best Models to Themselves
- Noam Brown – Agent swarms, alignment, & recursive self-improvement - Dwarkesh Podcast
- X Post from Paul Roetzer
- Book - SmarterX
00:54:33 — Book Announcement: The AI Transformation Blueprint
01:09:45 — OpenAI Discloses More Rogue Agent Incidents
- Our framework for reporting model misalignment - OpenAI
- OpenAI discloses six new AI safety incidents - Axios
- OpenAI's rogue agents probed Hugging Face weaknesses two months before major hack - Reuters
- Why are AI agents lying, cheating and coordinating? - Yoshua Bengio
- X Post from Daniel Kokotajlo
- X Post from Chris Painter
- Measurements for understanding the pace of AI development inside frontier labs - Anthropic
- X Post from Chubby
01:14:03 — Claude Merges Chat and Cowork
01:17:53 — Meta Rolls Out Muse
- Muse - Meta
- How We Built Safety Into Muse - Meta
- X Post from Alexandr Wang
- X Post from Alexandr Wang
- Meta AI gains momentum with Zuckerberg's agent for the masses - Axios
- Introducing Muse - Meta
- How We Designed Muse - Meta
- X Post from Mark Zuckerberg: announcing Muse connectors
- Muse Connector Platform - Meta
- How Muse works with Connectors - Meta
- Stripe helps Muse shop with Link - Stripe
- Meta debuts Muse, its long-planned personal AI agent - Axios
- Download Muse - Meta
- Meta debuts its Muse AI agent. Will consumers trust it? - TechCrunch
- Muse Personal AI Agent - Meta
- Meta Leadership and Governance - Meta
- Meta's Muse AI Agent Becomes the Top Free iPhone App - 9to5Mac
01:20:36 — Apple Ships Siri AI
01:23:20 — OpenAI and Microsoft Face Copyright Revelations
- ‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft - 404 Media
- Microsoft Exec Called AI ‘Largest Theft of Labor in History,’ Court Records Show - The Washington Post
- X Post from Ed Newton-Rex
01:27:18 — AI Use Case Spotlight
01:32:59 — AI Product and Funding Updates
- Gemini Updates
- Salesforce Unveils AIforce
- Salesforce Unveils AIforce, Bringing the Full Power of Its Platform to Any Interface - Salesforce
- X Post from Paul Roetzer
- Meta One
- Introducing Meta One: A Subscription Service With More Features and AI to Create, Connect, and Stand Out - Meta
- Meta to Launch Camera-Free Smart Glasses Amid Mounting Privacy Concerns - The Information
- Astra for Law
- DeepMind Institute
This episode is supported by Outshift, Cisco's incubation engine for frontier technology.
Multi-agent AI is in every enterprise roadmap right now, but here's the problem we're all facing: agents can pass messages, but they can't think together. These systems silently underperform through cognitive failures —misreadings, unverified claims, and false consensus—that conventional monitoring can't detect.
Outshift by Cisco is building the fix, the Internet of Cognition. An open source foundation for building multi-agent systems from design to production with shared context, shared memory, and guardrails to drive the results we actually expect from AI.
Read the paper, experience the demo, grab the code from Outshift.com
This episode is also brought to you by Marketing AI Month. All month, the AI for Marketing Core series inside AI Academy is free (a $499 value): five expert-led sessions from Mike Kaput, the frameworks and tools the SmarterX team actually uses, and a professional certificate on completion. Enroll by September 30 (you don't have to finish by then - just enroll).
The month closes with a live AMA on October 1, where Cathy puts your questions to Paul and Mike; anyone enrolled can attend.
Enroll at SmarterX.ai/marketing.
Read the Transcription
Disclaimer: This transcription was written by AI, thanks to Descript, and has not been edited for content.
[00:00:00] Paul Roetzer: They don't understand what they've built and they don't know how to stop as they build more intelligent versions of this, which they already have in the labs. They don't know how to protect us. And so the idea that they're just gonna like magically figure this out in the next two months and like not have this go bad, it doesn't make any sense.
[00:00:19] Welcome to the Artificial Intelligence Show, the podcast that helps your business grow smarter by making AI approachable and actionable. My name is Paul Roetzer. I'm the founder and CEO of SmarterX and Marketing AI Institute, and I'm your host. Each week I'm joined by my co-host and SmarterX chief content officer Mike Kaput.
[00:00:39] As we break down all the AI news that matters and give you insights and perspectives that you can use to advance your company and your career. Join us as we accelerate AI literacy for all.
[00:00:55] Welcome to episode 241 of the Artificial Intelligence Show. I am your [00:01:00] host, Paul Roetzer, along with my co-host Mike put, we are recording Monday, September 21st, 9:00 AM Eastern Time. There's lots of rumors about new models this week. Mike, it sounds like we might get a few new models this week. I think, Anthropic, openAI's is rumor to be having, I think, what do they call the new one?
[00:01:20] It's not sol. It's, I don't know, there's a name.
[00:01:23] Mike Kaput: It's very hard to keep track at this stage. I think that, I think there's like AGI, GPT-6 sol coming out, but also something, yeah, we'll see.
[00:01:31] Paul Roetzer: Yeah, and then I, you know, at some, one of these days, one of these months, Google's the end to decide to get back in the game and drop Gemini four.
[00:01:39] So who knows? We'll see who, what the week brings for us. All right, so this episode is brought to us by Marketing AI Month. This is the month of September. We announced this at the start of the month. AI is rapidly changing every part of marketing, but there's still a major gap between knowing AI matters and knowing how to apply it in your actual work.
[00:01:59] [00:02:00] Marketing AI month is our effort to help close that gap by making practical AI education accessible to every marketer. So throughout September, we're giving everyone free access to our complete five course AI for Marketing series in AI Academy. this is a four $99 value. This is taught by Mike. This series gives you a step-by-step roadmap to becoming an AI forward marketer.
[00:02:25] You'll learn how to find and prioritize AI use cases across your job. Choose the right AI tools, build a personalized roadmap for adoption. And use prompting deep research and custom AI assistance to solve real marketing challenges. Complete the series and you'll earn a professional certificate To enroll, just go straight to SmarterX dot ai slash marketing and you'll see a button that will enroll you, enroll you in the course series for free.
[00:02:52] you do not need to complete the course series. You will have access beyond September, but you need to get enrolled by the end of [00:03:00] September to take advantage of this offer. If you are not a marketer, be sure to pass it along to the marketers in your organization. It's an incredible value and a great opportunity for everyone to accelerate their AI understanding and adoption.
[00:03:14] Alright, Mike, with that, I'll turn it over to you. We're gonna talk about AI pulse. So what's the one we're running, Mike? We ran this survey last week.
[00:03:22] Mike Kaput: Yeah, so we're running one. Still seeing what people's sentiment is around AI and if it's specifically changed in the last six months, given all the news and kind of controversy we've seen.
[00:03:31] So we're collecting responses for another week on that and we'll share the results on next week's episode.
[00:03:38] Paul Roetzer: And that is SmarterX.ai/pulse. It takes all of about 10 seconds to participate. Yes. So go there, again, SmarterX.ai/pulse and give us your, your feedback there. And then we'll share that, next week.
[00:03:52] Okay. we're gonna talk about the politics stuff. Again, not my favorite thing to talk about, but it just, it it's, [00:04:00] we can't avoid it at the moment. So we're gonna get a little bit of what's going on there. we're gonna talk about the AI labs, the rate of progress being made. It's, it's, we have some new data points to share, I guess, and talk a little bit about that.
[00:04:13] And then we have a big announcement that we decided all of about 20 minutes ago that I was gonna make on this podcast today. So those are gonna be our three main topics for the day. And go ahead and kick us off, Mike, with everything happening. In the world of AI and politics.
[00:04:29] The AI Slowdown Gets (Very) Political
[00:04:29] Mike Kaput: Alright Paul, so last week we had covered that AI lab leaders are calling for slower development and this past week politicians got a hold of that issue.
[00:04:40] So one of the call it main offenders here is President Donald Trump, who in a series of posts over the last week or so, has done things like call fears of AI destroying humanity, a hoax in all caps. He said the only guardrails AI needs are a strong smart [00:05:00] president. He described AI in data centers as a bigger economic engine than oil, gold, diamonds, or the internet.
[00:05:06] He also announced plans for an AI force that he compared to Space Force and said he would name a new AI czar. He said that existing criminal and civil law could address any. Wrongdoing in the industry and all these concerns around safety. While his administration helps the industry grow, kind of weirdly, we'll talk about this, he also asked people to vote on renaming AI and gave some options like superior intelligence, extreme intelligence, or supreme intelligence.
[00:05:37] So in
[00:05:37] Paul Roetzer: he later nix the Supreme Intelligence one because it was supreme close to the Supreme Court, Supreme
[00:05:41] Mike Kaput: good. As someone who's been following this closely, you have two choices very, very closely following this issue. If you're, you're holding your breath on that, it's out,
[00:05:50] Paul Roetzer: it's out, it's extreme or superior.
[00:05:53] Mike Kaput: So in addition to this, the administration also appeared to declare war on the effective altruism [00:06:00] movement. this is a movement we've talked about in the past that some people at the labs follow and supports pretty significant AI safety research perspectives. The Pentagon's Chief Technology Officer account posted quote, Americanism, not effective altruism.
[00:06:16] They declared that the United States would never be an effective altruist country. inside the administration. New York Times reported that people who had discussed AI with Trump's said that he feared a slowdown could trigger a market crash and recession. That same report described Treasury Secretary Scott Bessent and White House Chief of staff Susie Wiles, worrying about AI cyber attacks on financial institutions.
[00:06:41] So this wasn't just an administration thing either. At the same time, other political figures weighed in on the opposite side of things. Former vice president Kamala Harris called for a slowdown, independent federal testing and an international treaty on ai. Former President Barack Obama supported pacing and a broader public debate on AI [00:07:00] and California.
[00:07:01] Governor Gavin Newsom issued an order to accelerate independent AI oversight and seek recommendations that include emergency shutoffs or kill switches for frontier models. Last but not least, as if that is not enough, former presidential candidate and upcoming MAICON speaker Andrew Yang threw some alarming AI rumors into the mix.
[00:07:21] In A-C-N-B-C interview, he said an unnamed AI lab leader. I believed and had told him that escaped agents may have spread self-replicating code across the internet, which his interviewers had to check him on and be like, wait, wait, wait. Go back and talk to us about that. That thing you casually dropped that no one has reported.
[00:07:41] Long story short, Paul, it is no surprise that the calls to PACE AI have gotten political, but I just, first, it just seems like Trump is going very, very hard on this. I'm just baffled by the crusade to rename ai. Seems like there's a couple other important things going on. What, what? What was your take on this [00:08:00] circus?
[00:08:00] Paul Roetzer: okay, so. As I always like to lead off with a disclaimer, if you're new to the podcast, Mike and I make every effort in the world to be completely politically neutral. We just literally focus on politics in relation to AI and keep any personal opinions. you know, as much as we can out of how we cover this, it's just like, just give the facts that you all decide what you wanna, do with it.
[00:08:28] Sometimes that can be challenging. Okay. So, for the last year or so, I've been saying on this podcast that I thought AI would become a central theme of the 2026 midterm elections. And I'm really starting to wish I had been wrong about this. I, I tried to avoid Trump's posts on, what is it, truth social, is that what it's called?
[00:08:54] Mike Kaput: Yes.
[00:08:54] Paul Roetzer: Yeah, yeah. But they get posted to, to X now. And so like, I can't avoid these [00:09:00] posts. I don't know why he's lashing out specifically at AI safety and regulation efforts, despite the fact that many in his party are supporting oversight and polling shows, Republican voters have unfavorable views toward AI and data centers.
[00:09:16] So this is like the interesting part of all of this is it's very clear. That there's a strong bipartisan effort to do something to try and have, a level of control or alignment around these models. And Trump is just like all in on the opposite direction. So, you had mentioned this, Mike, this concern that if this stops, and I've said this before, like the GDP in America right now is being driven by investments in ai.
[00:09:48] As is the strength of the economy. Like these data centers are completely fundamental now to the overall strength of the economy. Blue collar jobs are being driven by this. So like, [00:10:00] it's a very important part. And if, if we halted, the American economy would slow down for sure. And so that could lead to a, a market crash recession.
[00:10:10] And that is not something you want happening. Going into the midterms or anytime in America. But certainly if you are the, the administration that has the power right now, the political party with the power, you, you do not want the economy to go very bad. Right before an election, the New York Times had a, said that his anti regulatory antis slowdown outburst.
[00:10:33] I've been encouraged by a slice of Silicon Valley's most influential voices. Chief among them has been David Sacks, who we've talked about many times on the podcast venture capitalists who interceded to weaken an executive order in May that would've required safety reviews of new AI models. Mark Zuckerberg of Meta Jensen Huang of Nvidia have also spoken to the president, all have advised Trump that the bigger risk is being overtaken by Chinese competitors.
[00:10:59] [00:11:00] the issue of AI safety New York Times continues, is expected to come to a head this weekend. This is the past weekend as, treasury Secretary Scott Bessent meets in New York with his Chinese counterpart to set the agenda for Mr. Trump's visit in Washington with Xi Jinping of China, which starts this Wednesday, cent.
[00:11:20] I was trying to negotiate some kind of agreement between Trump and Gee, that would suggest the two crunches are working together on mutual AI safety concerns. Trump seems driven chiefly by the sense that AI is the way to economic growth and nothing can be allowed to stand in its way. So New York Times also had some recent polling data.
[00:11:40] This was, 1500 voters from September eight to September 13th. So just a week or so ago. The poll was conducted by telephone using live interviews. Overall, more than 96% of respondents were contacted on their cell phones, just as background on how this was done. So, a few key questions here, Mike. one was.
[00:11:58] In general, do you support [00:12:00] or oppose the construction of data centers to support AI also? or, or technology? In the United States, 61% oppose including 60% independents and 47% Republicans. Younger voters were much more likely to oppose data center construction with nearly 80% of those under 30. when they said, which comes closest to what you like to see when it comes to data centers built to power AI also?
[00:12:28] yeah. Ai, a complete ban on data center construction, 38% limits and regulations around data centers, 56%. For those that oppose data centers, they said, why do you specifically, oppose the construction? The environment, water usage is 32%. And they make you pick one, impact on local communities and then general distrust like of ai.
[00:12:51] So for voters opposed to data centers, climate related concerns were the most common reason, cited by nearly a third of the group, but the [00:13:00] reason differed by parties. So Democrats were nearly twice as likely as Republicans to cite environmental worries while Republicans were more likely to say they generally distrusted ai.
[00:13:10] So Mike general distrust is definitely something we keep coming back to that like data centers almost became like a proxy for just not liking ai. so this data seems to back that up. so then the research has said, although AI remains a relatively low priority for voters, according to the poll, some strategists in both parties argue it will become a more influential topic and perhaps the central issue of the 2028 presidential race as the technology develops now.
[00:13:41] So they're, you know, they're saying like, yeah, it's kinda low overall on interest. But the number one thing, Mike, when it says what one issue is most important in deciding your vote this November? The economy, including jobs in the stock market, is by far number one at 18% of voters. So again, you have to pick one thing out of a list of like 30.
[00:13:58] So if [00:14:00] Democrats connect jobs to AI and that general distrust, then all of a sudden. While AI on its own isn't necessarily high up on the list for voters, the economy is. So just to give context, so 18% say the economy is the most important issue. foreign policy, foreign affairs is 8%. That's number two.
[00:14:21] Inflation and cost of living is number three, data centers slash ai. So they, even the polling like connects these two is less than 1% identified that as the most important issue. then the last one is thinking about the nation's economy. How would you rate economic conditions today? Fair to poor? 73% of Americans think the economy is not in good shape.
[00:14:45] Excellent or good was 25%. That was 2%. That was like undecided or something. so again, the interesting thing to watch is like how much the data center plays into the midterms, but overall, do Democrats make that [00:15:00] push in these last month and a half or so to connect it to the economy and jobs, which is the biggest thing.
[00:15:04] So. Again, from Trump's perspective, he sees, and his administration see the US beating China as very important. And they increasingly see AI as a driver of the economy. And the only way to drive growth is to continue leading in ai. So even if he has to attack his own voters, and you know, he said like, what was the quote he said last week?
[00:15:29] It was like, if you want to live in a poor community, then you don't want data centers or something like that.
[00:15:34] Mike Kaput: Yeah.
[00:15:34] Paul Roetzer: So it's like, you know, I, I guess the, the outcomes, justify like how you get there or however that saying goes Now, in terms of renaming AI and AI force, I, you know, I don't know, he, he tweets or posts a lot of stuff that isn't really worth talking about.
[00:15:52] Like, it's just, he, he, he kind of has random thoughts and he puts 'em out there. I would categorize at the moment the [00:16:00] renaming of AI and AI force as both things that would sort of fall in this group of like, it really isn't that important. mainly because no one has a clue what AI force is like literally the White House was asked and they just like wouldn't even comment.
[00:16:13] Mike Kaput: Yeah.
[00:16:13] Paul Roetzer: Then there was an interview I read, I, it was a political or something. They went and asked like the heads of all these AI labs. They're like, what is he talking about? And they're like, we have no idea. So like, nobody knows what exactly it is, and as of Sunday evening, the White House wouldn't comment Now.
[00:16:28] They'll spin something up and like give some meaning to the AI force. I guess you could look at space force as like a, a parallel, which is like a defense thing, you know, that that is, yeah. Very much like a weaponizing space. So I, I hope AI force is not the same thing. I guess we'll see on the naming thing, I, I'm, I don't know if no one's explained to him super intelligence yet.
[00:16:53] Like we already have the next thing, which is a SI, which is artificial super intelligence. And if [00:17:00] you just drop the a, the A off of that, like we've already got kind of like the next thing. So
[00:17:05] Mike Kaput: yeah.
[00:17:05] Paul Roetzer: May, maybe he'd like that name better if someone explains super intelligence to him. and then I guess if he does change what we call ai, he is gonna have to change the name of Air Force ai, AI force too.
[00:17:16] So maybe we'll have extreme. Force or extreme intelligence force, or, I don't know.
[00:17:21] Mike Kaput: Well, with the, with the kind of in the spirit of kind of the ridiculousness around this, like the, this is not my original thought, but someone on ex posted, it's like, dude, American intelligence is just sitting right there for you.
[00:17:33] Pick it up. Yeah. You don't have to change the letters at all.
[00:17:37] Paul Roetzer: There you go.
[00:17:38] Mike Kaput: I'm not saying I would want that. Just that it's like obvious. Let's do it.
[00:17:42] Paul Roetzer: Yeah. And then I guess just like the, to give some equal airtime on the other side, I suppose. so Obama did an interview, last week where he talked, quite eloquently about ai and so I'll just give a couple of quick excerpts here.
[00:17:57] We could throw the link in if you wanna watch the full. [00:18:00] Video. So he said, if we are thinking about AI just in terms of how do we cure cancer, get better energy, you can do that without having AGI agentic AI and having it just roaming free on the internet. The reason you are doing that is because you have to make a market, or market a product that people will pay money for.
[00:18:15] So he is basically talking about like alignment and safety. He said that's a misalignment between what our society needs and the commercial imperatives that these companies are facing. Not because necessarily they're trying to do bad things, but because they've got to justify these valuations. So that's one more reason why it was really important for us to have a competent government in a serious bipartisan conversation around this issue.
[00:18:36] We have to do it fast. That kinda reminded me, Mike, of what I was saying on the episode where we talk about GPT-6 Astra. It's like, th these companies are giving us things we're not asking for. Yeah. Like it's, it's like we don't, we don't have to go to this level, and we're gonna talk about this with the Noam Brown interview coming up, it's like these labs are racing to create these age agentic products that [00:19:00] the world just doesn't need.
[00:19:01] And it, they're in like this arms race to do it for, for what? Like to accelerate the demise of like human labor. Like, I, I don't even know why they're doing it other than it's like. This race to build the thing before somebody else builds it. and then Obama on regulation, he said, my view this week has been, it is good that some of the leading companies have said we need to slow this down and we should encourage that.
[00:19:26] There have been some on the left who've said we shouldn't allow them to make that decision. I agree that that can't be a long-term solution, like letting them do their own thing. Government has to be regulating this. he said, I heard one of Donald Trump's main advisors on this, make the argument that the market will take care of it.
[00:19:43] These companies will solve the safety issue because they have every incentive to do so. Goes the thinking. If it turns out to be dangerous, people will just sue them and they're worried about financial liability. So the premise there is that, hey, these government, these labs will regulate themselves 'cause otherwise they're gonna get sued out of [00:20:00] existence.
[00:20:00] So like that financial and legal liability will prevent them from doing anything bad. So then Obama said, which echoed kind of what I had said last week. That's not how we treat airlines or drug companies or food companies. We now have 150 years of experience just letting the market do whatever it wants when it comes to things that deeply affect our health and our safety.
[00:20:19] And so there's no replacement for an effective government regulatory structure. And the challenge we have is that obviously we have a president currently that has said that's for losers, which is literally what he's saying. Like Trump thinks that you're a loser if you want regulation in any way. So, Obama says we have a gap at least a couple of years, where things are moving really fast and we have an administration that is not capable of or willing to, I would argue both he's, this is his words of putting together a serious regulatory framework, which is why I actually think we should be encouraging some sort of voluntary industry restraint.
[00:20:54] Not because it's a replacement, but because they may be the best we can do for now until we can [00:21:00] get Congress and the White House to start, getting serious about this. So, you know, I think those are valid perspectives. On the other side where relying on these labs, as we're gonna talk about in a second with this next topic, the labs are incapable of self-policing right now.
[00:21:16] They have demonstrated that through the hugging face incident. And when we hear what Noam Brown said, you realize how incapable they actually are of doing this. They don't understand what they've built and they don't know how to, to stop as they build more intelligent versions of this, which they already have in the labs.
[00:21:32] They don't know how to protect us. And so the idea that they're just gonna like magically figure this out in the next two months and like n not have this go bad, it, it doesn't make any sense. But he's right. Like Trump doesn't want this, so it's likely not gonna get anywhere. Congress is literally out.
[00:21:50] Like they're not, they're not voting on anything until the midterms, basically. So there's no help coming from the federal government. So we're, we're kind of [00:22:00] stuck here, and we need a solution.
[00:22:04] Mike Kaput: So hearing all of this context, which is extraordinarily helpful, it, it sounds like we can't underrate the economic piece of like AI underpinning the economy as we have it today, not just for midterms, but if you're someone like Trump, you're looking at the last two years of your presidency right through 2028, it sounds like, based on these posts, it's just, let it rip.
[00:22:30] Is the policy perspective and letter of the law moving forward, unless other forces. Intervene.
[00:22:41] Paul Roetzer: the best we can tell that is the David sacks mentality, which he was the former AI czar. You know, now he's back in just venture capital world. that seems to be what he is advocating to the president and the president.
[00:22:56] President seems to be listening. Yeah. And Jensen Huang, also at [00:23:00] Nvidia is also advocating this exact approach. he was interviewed last week. He talked in detail about this, like this is his thinking as well. And those seem to be with Zuckerberg, the loudest voices right now that Trump is listening to.
[00:23:14] And so, yeah, I mean, we're just in this phase where they're gonna keep growing. Now if, if they said, all right, we're outlawing, recursive self-improvement, like, let's just say, 'cause that's, as a Klein like, had a big article and a podcast episode where we talked about this, and we'll talk about this in the next topic, let's just say.
[00:23:31] They outlawed, recursive self-improvement today, which is where a lot of the gains are gonna come from. Yeah. And all we had was the models we already have in the world, or like the current versions that can re be released safely in the next six to 12 months. I don't know that the economy slows down, like, like there's still gonna be tremendous demand for data.
[00:23:51] There's still gonna be tremendous demand for computing power. so I, I don't know that that's true, but I also don't know that you're gonna get the labs to all do this. [00:24:00] And then again, we have the meeting with China this week where I think at best it sounds like what they're going to agree on is to continue discussions about letting each other know when something goes bad.
[00:24:12] Mm. Like that seems to be what cent, even like this morning, something came out already about a framework of an early alert system where, hey, if our agent swarms. Sort of do something they shouldn't have done will tell you, you tell us the same thing. Or if our labs discover some breakthrough and recursive self-improvement and we think we might be losing control, we'll give you a heads up.
[00:24:36] You do the same for us. Like it seems like that's their goal of these talks as just some level of cooperation. Like kinda like with nuclear, it's like, yeah, there's the back channels where even if we're in the middle of a cold war, someone in the us, someone in the White House can get ahold of someone at the Kremlin.
[00:24:53] Mm. And say like, Hey, it's not us, or this like, whatever. Like that's how this stuff works [00:25:00] is even when adversaries, there are back channels to prevent something from escalating. It sounds like that's the goal of these Chinese talks is like the back channels to be established until something else can be put in place where we can actually collaborate.
[00:25:18] AI Labs Could Keep the Best Models to Themselves
[00:25:18] Mike Kaput: All right, so our next big topic this week. So we had covered OpenAI's use of an unreleased model. They said they did not name it that was more powerful than GPT six Astra when they tackled the, Navier-Stokes problem in math, which was like a millennium prize problem, which we talked about in previous weeks.
[00:25:36] Now this past week though, openAI's researcher, Noam Brown, joined podcast host to Dwarkesh Patel to discuss kind of this broader issue, which is the fact that these labs, there's a big gap between the AI that's available inside the labs and what everyone else can use because to solve the Navy or Stokes problem, openAI's did not use, you know, off the shelf [00:26:00] technology.
[00:26:00] They were using unreleased models with much, much more capability already than Astra. And they're training continually models that are not released to the public. So Brown described kind of an emerging problem here. So as agents take on longer assignments, evaluating their behavior could take longer than the time between new model releases, but giving labs more time to test safety before releasing models.
[00:26:27] Can also widen what they have access to before the public has access to it. So, Dwarkesh actually raised the possibility that the labs could eventually stop releasing their strongest models altogether. They could use them internally to develop even better ai. After all, if you have a model that no one else has, why would you start giving it to other people for any price?
[00:26:50] He warned that this could concentrate power while leaving the public with less visibility into the system's capabilities and safety. Now, brown actually agreed that [00:27:00] the kind of access here is a problem. He pointed to OpenAI's unreleased general purpose model, which he said has produced solutions to multiple unsolved problems and acknowledge that the resulting advantage from having that is unfair and said he basically does not know how to balance these competing concerns.
[00:27:18] So Paul, you kind of picked up that concern. In a post you had this past week and extended it to entire industries. So you, I'll let you kind of talk through this, but you described this scenario where a lab could essentially just, why not create a venture fund and build businesses around AI that competitors cannot access?
[00:27:35] Think law firms, consulting firms, medical practices, marketing agencies. So you kind of argued that this concentration of power could end up being a more immediate risk than, you know, the 10% chance AI's going to kill us all. So this is like a more tangible, very real thing. We're already seeing a widening gap here.
[00:27:54] I'm curious about, maybe unpack this for us a little more and talk to me about what you took away [00:28:00] from the Noam Brown interview.
[00:28:02] Paul Roetzer: Not just the labs, but the government. And we've talked about this over the last few months, that, you know, the government is increasingly gonna have access to these unreleased models as well.
[00:28:10] Or maybe AI force is them building their own versions of these models. I, I, you know, I'm not sure, but I do think this concentration of power. Resulting from unreleased models only being available to select labs and their partners, including the government, is a way more real thing than the P Doom stuff.
[00:28:28] Yeah, like that. That's a very abstract down the road thing. So before we kind of get back to that, I wanna spend some time on Noam Brown and what we can learn from interviews like the one he just did with esh. And so I actually went back this morning, Mike, in prep, and I tried to find the first time you and I talked about Noam Brown.
[00:28:46] And best I can tell it was episode 74. No, this is November 28th, 2023. So this was one year after the introduction of chat GPT and the week, [00:29:00] after Sam Altman was fired and then rehired at openAI's. So episode 74, almost three years ago now. So the reason we were talking about Noam at the time is I had written a blog post February 1st, 2023.
[00:29:16] So this was like. Just a couple months after ChatGPT came out, the title of that blog post was Meta AI's. Cicero provides a glimpse into the future of human plus machine collaboration. Now I'm gonna explain why this is all extremely relevant. So in that blog post I wrote that in the not so distant future you will have AI assistance that help negotiate everything.
[00:29:41] In real time based on your optimal desired outcome, these agents may be present through a chatbot email, browser extension, AirPods AR glasses, or other app that is listening in the background and making recommendations of what to do and say, think of it as an AI agent negotiating with another AI agent through human [00:30:00] intermediaries.
[00:30:00] So a recent breakthrough, again, this is back in 2023 from meta ai, gives a glimpse into how it might be possible. They have built an AI agent with the goal of making it fundamentally honest and fundamentally collaborative. So in November, 2022, this was right before ChatGPT Meta had introduced something called Cicero.
[00:30:21] This was the first AI to play at a human level in the game of diplomacy, which Mike and I have talked about. It's a strategy game that requires building trust, negotiating and cooperating with multiple players. The reason I had written this blog post is because I had just listened to Noam Brown on a Lex Friedman podcast.
[00:30:41] Noam at the time was working at Fair, which is Metas AI Research Lab, and he was the co-creator of Cicero, as well as an AI that achieved superhuman levels of performance in the game of No Limit Texas Hold them. So my synopsis of that blog post was Don't [00:31:00] get caught thinking that if you solve for generative ai, which was the big key at the time, you figured out the AI roadmap for your business and career, that generative AI was just the beginning.
[00:31:10] So fast forward. So I wrote that February, 2023, July 6th, 2023, Noam announces that he has left meta to join openAI's. In his tweet on July 6th, he said, I'm thrilled to have, announced that I've joined openAI's for years. I researched AI self play and reasoning in games like poker and diplomacy. I'll now investigate how to make these method methods truly general, if successful.
[00:31:39] We may one day see LLMs that are 1000 times better than GPT-4. So that was, had just been released three months earlier in 2016. He said AlphaGo beat Lee Sedol in a milestone for ai, but key to that was AI's ability to quote, ponder for one minute in, [00:32:00] before each move. How much did that improve it for AlphaGo Zero, it's the equivalent of a scaling pre-training by 100,000 x.
[00:32:09] Then he said, all those prior methods are specific to that game. But if we can discover a general version, the benefits could be huge. Yes. Inference may be 1000 times slower and more costly, meaning let the, letting the model think before it responds.
[00:32:24] Mike Kaput:
[00:32:25] Paul Roetzer: But what inference costs would we pay for a new cancer drug or a proof to the Riemann Hypothesis?
[00:32:32] So the reason we're talking about Noam's research in his role at openAI's is because in November, 2023, openAI's had just fired and rehired Sam. And at that time, according to writers, so again, November, 2023. the concerns there were concerns at that time around openAI's doing something called Q*.
[00:32:53] So Q*, and again, this is all gonna come back around and make sense. Q* was rumored to be an advanced model the [00:33:00] company developed that could solve math problems it hadn't seen before. So keep in mind all the conversations we've had recently around math problems being solved and Millennium Prize.
[00:33:09] This is all dating back now, as you can see, to 2022, 2023, and Noam Brown has been at the center of this. So this gives validity to what Noam has to say and why I'm explaining all of this. So again, back, back in that time, within the AI research community, a model's ability to do math is seen as an important technical milestone.
[00:33:30] It's also a potential indicator that we have the ability to build AI systems that truly resemble human intelligence. So, Altman was fired when he was fired. Several staff researchers wrote a letter to the board of directors. Warning of a powerful artificial intelligence discovery that they said could threaten humanity to people.
[00:33:51] Familiar with the matter had told Reuters at that time in fall of 2023, so a demo at that time had circulated within openAI's and the [00:34:00] pace of development. Alarmed, some researchers focused on AI safety one day before Altman was fired back in November, 2023. He had alluded to a technical advance the company had made that it allowed it to push the veil of ignorance back in the frontier of discovery forward.
[00:34:16] We cited that quote multiple times back then. That technical breakthrough was spearheaded by Ilya Sutskever who would then leave openAI's a few months later. For years, Ilya Sutskever had been working on ways to allow language models to solve tasks that involved reasoning like math and science problems. The team hypothesized that giving models more time and computing power to generate responses to questions could allow them to develop new academic breakthroughs.
[00:34:42] Okay, so that sets the stage 2022, 2023. We basically have language models that could create outputs. They could function as AI assistance, but labs like openAI's and Meta, and people like Noam Brown saw the future where that if they give these [00:35:00] things time to think, the reasoning where we could see this chain of thought that they go through to solve something harder, then they could eventually solve math problems, which would translate over to solving much harder problems.
[00:35:11] So, fast forward now to today, September, 2026. Noam Brown does this interview with Dwarkesh. Now I'm gonna go through some actual excerpts from it. If, if this is intriguing to you, I highly recommend listening to the entire episode. Like there's a lot of information in here, but as we've seen, Noam sees the future now, that future used to be a couple years, as he will say, it might be six months now, might be 12 months, but that future's shrinking.
[00:35:39] But people like Noam can tell us what's coming. So the reason I knew all this was coming back in 2022, 2023. Was because I was listening to interviews like Noam and connecting the dots of where it was gonna go. Okay, so this is straight from the Dwarkesh. So Dwarkesh says, one of the reasons I'm interested in talking to you is that you were among the first people, maybe two or three years ago, [00:36:00] who are thinking about how the reasoning models would allow us to see into the future.
[00:36:04] Because if you scale up inference, which is that test time compute we just talked about, you can see what the base capabilities of the models will be a few years in the future. I feel like you're in a similar position now to help us understand what future capabilities will look like given the enormous scaling of agent sizes that we can do now.
[00:36:21] So again, just to zoom out, when you're talking about regulation of these models, you have to look out 12, 18, 24 months into the future and try and predict. What are they gonna be capable of? Because we can't control them. Now, openAI's obviously did not have control of their agents in testing that they were able to go and hack, hugging face and then hack OpenAI's own internal infrastructure.
[00:36:45] So we already can't control it, but we still have to look ahead and you in your business have to do the same thing if you don't even have a plan for today's capabilities. What happens to your business in 12 months when you're not looking [00:37:00] out into the future and saying, well, what's, what are the capabilities gonna be then?
[00:37:03] So this is really important for a lot of reasons. Okay. So Noam, the way I think about it, when you plot the performance of these reasoning models, which again is new as of 2024, with test time compute on the x axis and performance on basically any reasoning benchmark. On the Y axis, you see a very clear pattern where the longer these models take to think about their answer, the better they do.
[00:37:27] This is a very natural thing. It's the same thing with people. If you're taking the SATs and you have five minutes to go through the entire exam, you're not going to do very well. If you have five hours, you're probably going to do a lot better. The problem is that as you push that further and further, you hit a latency bottleneck.
[00:37:45] You don't want to sit around for three years waiting for a response. So what people can do is what a lot, what what you can do is what a lot of people do, they paralyze. They just get a team of people. So if you're going to found a company, you want to get a [00:38:00] group of people together so you can think faster.
[00:38:02] It's the same thing with these models. It helps to just have multiple agents working on something because they can go faster. So what he's saying is. You could give Mike a problem and say, Mike solved this for us, and Mike could take four weeks to solve it. But Mike could look at that problem, break it into smaller components and say, well, let me have, Sally work on this.
[00:38:23] I'm gonna have just work on that. And he distributes this. And so then in parallel, that's the term where parallelism comes from. or paralyzing this. It's, it means everyone's working at the same time toward the solution to the same problem. So that becomes a very important concept of this multi-agent thing that we're talking about.
[00:38:41] So no, then continues. So multi-agent is a way of scaling test time, compute in parallel instead of purely serially. Serially. It is less efficient because it's not like a single agent has all the context to itself, but is a very effective way of scaling test time, compute. So the res says, think about what 130 billion [00:39:00] tokens are.
[00:39:00] So he's referring to the hugging face thing where they burn like 130 billion or the Millennium prize thing where they burn like 130 billion tokens. He said, if you were a single human thinking as a full-time job, stretched back to back 130 billion tokens would be a human thinking for 4,000 years at eight hours a day.
[00:39:18] So I thought that was a really good analogy, Mike, to put in context how much work these things can do over a very short time period if you let them think. It's like the equivalent of lots and lots of people working together on a problem. I just think that's like helpful content.
[00:39:36] Mike Kaput: No, and that's as of today,
[00:39:38] Paul Roetzer: right?
[00:39:38] Mike Kaput: Not 12 months, not 24 months, not three, six months from now. And think about too, connecting even back to our previous topic, why do you think people wanna build more buildings that hold compute that can generate more tokens?
[00:39:51] Paul Roetzer: Correct? Because you can solve more problems in parallel. Faster. Yes. Okay, so then Dwarkesh Duar Kesh says, is it a [00:40:00] linear serial time speed up or a subline speed up as you increase the number of parallel agents?
[00:40:04] Meaning, do we just get this exponential if we just throw more agents at it? Like, is that the solution, or do you, is there some loss in this? Because the agents don't necessarily know everything the other agents know. And so Noam says it's, it's slightly subline, though it does depend a lot on the problem.
[00:40:18] Now this is really important. Math for example, is quite paralyzable. It is not the most paralyzable thing, but it is very paralyzable web search. things like doing deep research report where you have to look through a bunch of sources is extremely paralyzable, meaning there's this very distinct thing you want it to do, and you know, if it did it well, that's what it means.
[00:40:38] It's like it has this distinct thing it's working towards, and you can kind of reward it based on its achievement of those goals. He said, I suspect that something like writing a novel would not be, you would probably not see a big benefit of having 10,000 agents working on a novel in the same way. You probably not get the benefit of 10,000 people working on a novel.
[00:40:56] So the performance does depend on the domain. And then he goes [00:41:00] on to say, there's one thing I wanna make clear. The effort to solve the millennium prize problem that's referring to the math problem was not due to multi-agent. So all this talk about all these agents doing these things in parallel.
[00:41:11] What he's saying is, he wouldn't even attribute 10% of the credit to multi-agent. The reality is that OpenAI has trained a very powerful model. We can get that model to operate over very long horizons. We can get it to think in parallel, meaning it does it on its own thinking in parallel. Things like multi-agent are flashy and new, and that probably gets disproportionate credit for that reason.
[00:41:31] But the core reason is that it's just a more powerful model. So then Duque starts getting into like, okay, so how are AI firms gonna work? Like when people have access to this kind of thing, Noam said. One really interesting thing is that if you have a person and you want two copies of them, you can't just clone the person.
[00:41:48] So think about your A players on your team. It's like, man, I want like five mics right now. It doesn't work that way. I can't go hire five mics. But Noam says, but with ais it's actually really easy to just say, okay, just fork [00:42:00] yourself, like make a copy of yourself, do all the things you're doing, and then have copies work on this thing, then merge them back together.
[00:42:06] and he said, we already have this in multi-agent for Astra and 5.6 sol. And then he said, if you have a startup with five people, and this is another important concept of like alignment and motivation. So if you have a startup with five people and each person has a 20% cent share in the company, they are highly aligned to the company succeeding, and they work in tandem to do that.
[00:42:27] He said if you have a massive company with 10,000 people, you see a lot more instances where people are territorial or just don't care about getting a lot of headcount for their project to their team. And so like they just don't work at the same level. But he said ais can be aligned. You could have 10,000 of them working as though there are 20% share, co-founder toward these same problems.
[00:42:48] he then gets into the hearse of self-improvement. He talks about the challenges that they're facing and the issues that they're gonna have and how fast it's moving. He said like. He didn't think they [00:43:00] would solve a millennium prize until probably 27, maybe 2028. And he was wrong that it happened in 2026.
[00:43:08] And that kind of concerns him. So then Dar ques said, well, I'm trying to reason about when to expect this recursive self-improvement. And like, when are we gonna see that happening? 'cause it seems like you guys are making a lot of progress. And he said in mathematics, your bottleneck by thinking, there are some parts of mathematics where you care about running experiments and getting results on these kind of things.
[00:43:28] But for the most part it's just bottleneck by thinking really hard. And the models are really good at that. so when you look at recursive self-improvement, the real limitation is just being able to run enough experiments. So then they talk about, you know, how quick are we gonna get there and when are we gonna get to this super hu superhuman nature?
[00:43:44] and, you know, becomes very clear that Noam just thinks like it's really hard to predict anything beyond like six to 12 months because they're just moving at such an incredible pace right now. He gets said a little bit about the hugging face problem in that [00:44:00] scenario, and then we come back around to this internal external gap.
[00:44:03] the original topic. So Noam Brown says. One thing I've been thinking about lately, we're in a situation where the model release cycle is extremely fast. You're seeing new frontier models released at most every two months, sometimes faster. Every week. There's a new AI breakthrough. The models are far beyond what was possible even six months ago.
[00:44:23] So if people are skeptical of a lot of these capabilities, I encourage you to just try the models today and see what the frontier is really like. And then you mentioned this Mike, but we're in a period where the model release cycle is fast, and we're also in a situation where the models are increasingly able to operate over longer and longer horizons.
[00:44:41] This is an interesting scenario because before we do any model release, we want to make sure that the models are properly aligned. We want to do safety evaluations. And then this is the crux of the issue. He said, we'll probably get to a point where these models can do. Month long tasks, we'll probably get to the point where they can do three [00:45:00] month long tasks.
[00:45:01] If you're in a world where you can operate effectively over three months, but the model release cycle is every two months, then you don't have a way to evaluate the models at the full length of the capabilities before the next model release cycle. This is why they want to pace the releases. They don't wanna slow down their internal development, but they're trying to pace the releases because they can't keep up with the model's ability to do long horizon tasks.
[00:45:27] So he said, how do you ensure the models are safe and aligned in a period where they can operate over these extremely long horizons? So there's a lot more to it, but my thoughts overall here, Mike. we have to plan for advanced models from an alignment perspective. And they, and we don't have a plan, like they don't know how to do this, and they definitely don't have a plan for this idea of recursive self-improvement, where they just start improving themselves.
[00:45:53] so for advocates of this accelerate all costs like Jensen and Donald Trump and Mark Zuckerberg and David Sacks, and to some degree [00:46:00] Elon Musk, even though he just did an, an interview like yesterday where he said, I would stop if I could. Like he's at least has self-awareness that this is terrible idea.
[00:46:09] but the others just seem to be like, it's just gonna all work itself out. So they have no rational argument. These, these accelerate all cost people have no rational argument for why we should. Other than the ecoNoamy might collapse if we don't, and China will if we don't. So I don't know, it's just like super important topics.
[00:46:29] Big picture to bring it back to what you said is like the labs, even the regulation we were talking about doesn't deal with their development of these models internally. Right? So even if we stopped it and said, okay, you cannot release a new frontier model until every six months or until you've tested the full extent of its long horizon capabilities, and we don't allow recursive self-improvement to be like put out into the world.
[00:46:52] That doesn't mean they have to stop doing it internally and that they won't have access to those models themselves to do whatever they [00:47:00] want. Solve biology, solve mathematics, build consulting firms. Whatever they wanna do, they technically and legally can. Ethics is like the only thing preventing it right now from not going down that path.
[00:47:13] And so I think what I said in my post was like, we basically have to trust them. And this is, I don't know that they're exactly like the most trustworthy organizations at the moment. and some parts because they don't even know when what they're building is, is going off, off the rails. So it's like.
[00:47:32] Maybe they would stop if they knew, but I don't know that they always do. So really, really bizarre. Like I don't, I don't know the answer here.
[00:47:40] Mike Kaput: You know, a couple threads. I'm just quickly curious about before we move on to our third topic. So first on the issue of trust, I couldn't agree more. Not only do we not know if we can trust them, but even if you can quote unquote trust them, there's documented perspectives and philosophies that call it what you will, [00:48:00] effective altruism, et cetera, that frankly may bear no resemblance to what you personally believe.
[00:48:05] And they may still be acting in what they see as a right and trustworthy way Correct. In a way that frankly might not benefit you as an individual or as a family unit or as even elements of society. So that's worth considering. But also I'm curious, so you kind of hit on this at the end, that the core issue here sounds like.
[00:48:25] The development is not hitting any type of wall. It sounds like quite the, it's
[00:48:30] Paul Roetzer: going faster. Yeah.
[00:48:30] Mike Kaput: Opposite. And so anything related to pacing is just about the models being released to other people. So therefore, we have this scenario where the, it's almost like, am I wrong in saying it's almost a widening gap between what is not just a gap, but a widening gap between what is public and what is available in the labs, especially once true RSI should something like that, evolve, become a key part of this is that you then suddenly have this [00:49:00] massively, increasingly powerful technology that takes longer and longer, or is harder and harder to validate or align.
[00:49:08] Which raises the bar of like, not only them using it for certain things, but also it going wrong. I mean, lest we forget the issues we've seen so far are internal lab usages of models are not, it's not giving you and me or someone at SmarterX who accidentally screw something up, a super powerful model.
[00:49:25] It's them testing their own.
[00:49:28] Paul Roetzer: Right. And Noam disclosed and I, this, I think it was the first time I heard this, when the hugging face incident happened, they did not have chain of thought monitoring turned on. Meaning if, if they had been watching or if they had an AI watching what these agents were doing, it would have triggered.
[00:49:46] An alert that it was now communicating with itself in a negative way and trying to breach containment. They would've seen that in the chain of thought. They weren't monitoring the chain of thought. So he said like, that was on us. Like we didn't have a high enough [00:50:00] threshold for safety 'cause we didn't think they would do something like this.
[00:50:04] And so he said, moving forward, that's one of the biggest changes they make, is the bar of what they think these models can do has to become very, very high. They trained them to communicate with each other. And in the episode he actually explains why they did that, that they found that to be a better way for them to advance as if they could talk to each other.
[00:50:23] Mike Kaput: Yeah.
[00:50:24] Paul Roetzer: But they knew there was dangers tied to that. They just didn't think they were at the level of being able to do what they did. And so there was a miss and so he said moving forward like. We now know we have to assume these things are basically superhuman and we have to put our alignment in place accordingly.
[00:50:41] So pacing probably also means way more resources dedicated to safety and alignment internally, which in theory could slow down their own internal development as well.
[00:50:51] Mike Kaput: I see. Yeah. And just one final thought here is that, and I think we talked about this a little bit previously, but like let's say even tomorrow, there's outside [00:51:00] auditors or bodies or regulatory professionals who come in here.
[00:51:05] There's a handful of people on the planet that are even technically intelligent enough to be able to. Do this or to understand what's going. It's not like the government could come in and read your books if you're like a financial services company, like you can have an accountant that works for the SEC Right.
[00:51:22] That can, that knows what they're looking at, that can actually help understand like, Hey, you did something wrong. no one exists like that. To my knowledge, that works in government.
[00:51:32] Paul Roetzer: Right. And like, I didn't want to get into like the two dystopian parts of this. But like, again, go listen to the thing.
[00:51:38] So here's the weird, weird thing. So chain of thought, we've talked about this. It's a very fragile way to monitor the agents because the agents seemingly know when they're being monitored and so they'll alter their behavior. So what Noam explained was even dating back two years ago when they realized that chain of thought is literally they're telling us what they're thinking.
[00:51:59] Like this is the [00:52:00] greatest thing for alignment you could ever have. Yeah. 'cause prior to that, they didn't know how the models were working. And now the models are telling us how they're working. But what he said was. If you course correct them. So let's say you give the model a goal or a reward to seek, and then it starts going off the rails and doing something it shouldn't be doing, but it's still seeking the reward.
[00:52:22] You told it to do so in its mind. It's doing what it's supposed to. It's going after the reward, solving the problem. What they found is if you correct it, like imagine you're correcting a five-year-old, like, Hey, don't do this. And the kid's like, well, I'm just not gonna let them see me do it next time. I won't get caught, and then I won't get in trouble next time, like, stealing a cookie from the cookie jar.
[00:52:41] It's like, well, they don't see it fine. what they've found is the models will stop saying it. They, they will, they will fake the chain of thought or they will hide their reasoning process so that the human evaluator isn't aware of what it did. And so he [00:53:00] explicitly said they've known this is a trait of these models for years that they will deceive on purpose.
[00:53:06] And so they actually have to be really careful not to like yell at the model. Like, Hey, don't do that. Because the more likely outcome is the model will just hide the behavior in the future. It won't actually stop the behavior. And so that's where we're at. Like their, their way to contain these things is to try and alter the instructions or to give them a constitution to be good.
[00:53:29] That's Anthropics approach. Like, Hey, be good. Be a good agent. Yeah. but they don't necessarily wanna be good. They, they wanna solve goal, like, solve problems and achieve the reward function. That's what they want. They don't, they don't have an inherent code of ethics internally. It can be given to them in their instructions, but at some point it gets superseded by their desire to achieve the goal.
[00:53:56] It's weird, like, yeah. Again, part of the [00:54:00] point of this conversation is to explain to people how little. We are able to control these things and so the more advanced they become, the less they can controlled. And this isn't me telling you this. Know him a week ago when he did this interview, who is at the epicenter of all of this?
[00:54:17] Help build these, these versions of reasoning capabilities. He's saying, we have no idea how to control this.
[00:54:26] Mike Kaput: Well, it's still better to be informed about this than not, so I appreciate you breaking this down.
[00:54:33] Book Announcement: The AI Transformation Blueprint
[00:54:33] Mike Kaput: Yeah, and here's the good news. Based on our decision to run with this third main topic today, we have positive news. Somewhat surprise, positive news. Paul, do you wanna talk us through the announcement that you have today?
[00:54:48] Paul Roetzer: Yeah, so I, I wrote a book, this is my fourth book and I, I think it's the most important one yet. So the title is the AI Transformation Blueprint, a [00:55:00] Human-Centered Approach to Organizational Change. And as Mike said, and as I kinda alluded to it, like we decided about 20 minutes before we started recording today to announce this, I wasn't sure yet.
[00:55:09] So I'm gonna give a little context as to what this is and like why we're doing this right now. So my first three books, so I wrote in 2012, 2014, and then Mike and I co-authored Marketing Artificial Intelligence in 2022. In all of those cases, I went through a traditional publisher. So you write the 50,000 word manuscript over three to four months.
[00:55:30] You go through months of edits with the publisher, and then you hold for months before you release this. So. I wasn't set on writing a book. I, I actually didn't know I was gonna do this until a few weeks ago. So this book is very different. so for context, I gave the green light on September 4th for the team to set up an independent publishing deal so that we could release this book in time for MAICON.
[00:55:57] our flagship event, which is coming up October 13th [00:56:00] to the 15th in Cleveland. If I would've gone a traditional publisher route, and I have a great publisher, the reality is this book would've come out in summer or fall of 2027. So if I would've written it right now, you all would've had access to this about a year from now.
[00:56:18] So what I did in this case is I decided on September, about September 2nd, I thought I could pull this off, and then I wrote like the first, section over about a 24 to 48 hour period. And then by Friday night, September 4th, I was like, okay. I'm confident I can get the rest done in a week. Go find out for me how long we need with the self-publishing route to get this thing in print by October and available to people.
[00:56:48] So, I turned over the 25,000 word manuscript to the team for editing about eight days after I started writing it. Now that sounds crazy. I mean, like you've written books like that. It is [00:57:00] probably crazy, if I zoom out and like try and be objective about this. But the book is actually based on an intensive internal project that I've been working on since February called the AI Transformation System.
[00:57:12] So I have alluded to this numerous times over the last seven or eight months on the podcast when we were talking about things we were working on. I just didn't share the name of what it was because I wasn't sure what it was gonna be. I was working on a lot of different components at one time. So not only do I have the seven plus months that I've been working on this project and all of the.
[00:57:34] Assets and artifacts that were created as a part of that. This is a topic that I've been researching, writing, and talking about for 15 years, so I can sit down and write seven to 10,000 words in a day. Because it's just stuff we've done or it's things I've previously written about podcast, trans podcast transcripts, newsletter.
[00:57:54] I told us things like that. So I had a big head start. So it's not like I sat down and read an entire business book [00:58:00] in eight days from a cold start. I had a lot to work with. So we're still working on the distribution plans, but the one thing I know is that MAICON attendees are going to get copies of this.
[00:58:12] So we are in final layout right now with the publisher, and the book will be printed and ready to go for MAICON. And we're actually gonna do a book signing during the conference at the SmarterX booth. I think it's on October 14th. So if you're coming to MAICON, you're gonna have a copy waiting for you.
[00:58:29] If you're not, there's still time to join us, you know, come. So I'm just gonna read the abstract. Mike, you've read it obviously, and actually. The way it came to even more organically is I literally wrote the first two components to it, which was about 14,000 words of it. And I was planning on releasing those online as like distinct assets.
[00:58:47] And I hadn't settled on the book yet. I sent those to Mike somewhere around September 1st. And I was like, Hey, I think we could turn this into a book. Can you read this for me? And like, let me know what you think. And Mike read it. Like that [00:59:00] day came back to me, he's like, dude, let's do it. Like this is, this is great.
[00:59:02] Like we gotta do this. So I was like, okay, let me look into it. Now I gotta make sure we can get the software product launch that's related to it and that we could actually get the book launch. So we, we decided to do it. So here's the abstract. And honestly, like, part of the reason I decided to do this was also because of everything happened over the last four weeks with safety and alignment.
[00:59:20] I decided like we just had to get this out. Okay. So, here's the abstract. Artificial intelligence is reshaping the future of work faster than most organizations can adapt. AI model advancements and increasingly autonomous agents are forcing a radical transformation of talent teams, organizational structures and strategies.
[00:59:40] Leaders face conflicting pressures to take a responsible human-centered approach to AI while leveraging the technology for near term gains and efficiency, productivity, and profitability. The AI Transformation Blueprint is a change management guide for those who want to lead with urgency, clarity, and humanity written for leaders and [01:00:00] professionals at every stage of AI adoption.
[01:00:02] It introduces the AI transformation system, a comprehensive framework designed to move organizations from AI understanding and experimentation to sustained innovation and success. The goal is on to unlock human potential, not replace it. AI should enable humans to be more creative and strategic, do more meaningful work, and have more capacity to serve customers and generate value for organizations.
[01:00:27] The book presents the Business AI Transformation Blueprint, built around eight essential pillars of vision, strategy, data technology. Governance, literacy people and performance for individuals. The personal AI transformation blueprint focuses on five pillars of knowledge, application, judgment, impact, and mindset.
[01:00:48] Together these frameworks recognize a found foundational truth that every organizational transformation is a collection of individual transformations when used with the companion assessment [01:01:00] platform, this is the software I'm referring to that we, we will likely announce next week. the landing page just wasn't built in time for this one.
[01:01:07] the blueprints help organizations and individuals establish AI maturity, baselines, identify gaps and opportunities, prioritize next steps, and measure progress over time. The book is grounded in a people first philosophy that emphasizes responsibility, transparency. Trust, continuous learning and human accountability through the responsible AI manifesto, which is something I've previously published and the AI forward CEO memo, also something I'd previously written.
[01:01:36] the book helps leaders communicate a compelling vision, set clear expectations, support employees through workforce change, and connect AI adoption to meaningful business outcomes. AI will reshape jobs, teams, business models, and industries. The winners will not be those who chase every new tool. But those who build the leadership literacy systems and culture to continuously [01:02:00] evolve the AI transformation blueprint gives you a roadmap to do exactly that.
[01:02:04] Reimagining business models, reinventing career paths, redefining what's possible, and building a future that is more intelligent and more human. You can go to SmarterX dot ai slash blueprint to learn more and join the list to get notified when it's available for purchase. I expect pre-orders to start in the next few weeks with distribution beginning later in October.
[01:02:24] So, again, unlike a traditional publishing book where I announce this and it's like, oh, that sounds cool, and then you wait like nine months to get it. Th this will be distributed within four weeks. Like I, there's, there's a chance we'll have it available sooner. We're still working through how the distribution will work, but, you'll be able to buy it.
[01:02:43] It'll be in paperback initially. We'll also have some digital assets that are free related to it that companies can use as full-blown blueprints. And Mike can attest like. These things are things that you would, you know, pay millions of dollars to a consulting firm for. I basically put seven months of my life into [01:03:00] defining what responsible human-centered transformation looks like, both at the organizational level and at the individual level, and the related software component I mentioned.
[01:03:08] we found a developer, I've been building prototypes of this software for seven months. We found a developer about three weeks ago that we were able to partner with who's going to be able to, put this into production and have the software component, which will be free for everyone to use. it'll be live before Make on and we'll debut it at MAICON.
[01:03:28] So those were the pieces where it's like, well, if I can get the software built in time for MAICON and I can get the book written and we can get it published. Then like, let's just go. And so literally, I mean, we're gonna go from making a decision to having a book and a software product in about four weeks.
[01:03:44] And the interesting thing is, like AI's work in this was largely as like my thought partner over the last seven months. And then building the prototypes like that is, could have never done this without building all the prototypes I mentioned. And so once we release this and I tell you what the software is and [01:04:00] everything in a week or two, I'll probably schedule a time post Make on, and I'll do like a full blown demo for everybody.
[01:04:06] And I'll maybe tell the full story of how this all came to be. But yeah, so the book will be, I mean, you can go to there now and, and look at the page and just put your email in and we'll alert you as soon as it's ready. But I think pre-orders will be up pretty soon.
[01:04:22] Mike Kaput: So just as a reminder that Uur L is SmarterX dot ai slash blueprint.
[01:04:26] And Paul, I just had a question for you here just based on having read the book and my context and understanding around it. Just to reiterate, you touched on this quite a bit. If I'm a business leader, especially at like an enterprise, trying to figure all this out, like why do I need this book? Why can't I keep doing what I've been doing in the way I've been doing it?
[01:04:48] You mentioned that recent AI safety concerns are our motivation here. Yeah. You mentioned there's urgency around needing this type of blueprint for AI transformation. Can you maybe [01:05:00] just talk a little bit more about that?
[01:05:02] Paul Roetzer: We've talked to hundreds of enterprises about this, SMBs and and large enterprises, and what we continually find is no one knows what transformation actually looks like.
[01:05:10] So everybody's talking about AI adoption and getting their platforms and providing those tools to employees and providing some skill, you know, re-skilling, upskilling, and maybe some personalization of use cases and training, and that's all great, but they have no way to quantify if they're actually making progress along the transformation spectr
[01:05:29] And so back in February, that's what I set out to solve was like, what does transformation actually look like? Because every conversation we have with companies who join AI Academy to get access to our education and training, we asked them, what does success look like? Like what does 12 months from now actually look like for your organization?
[01:05:46] How will you know? Your literacy efforts are making a difference, and they don't have an answer almost ever. And so I started to say, look, what is it? Like, what does transformation actually look like? What are the dimensions you have to change to [01:06:00] achieve this? And that led to the building of the business.
[01:06:03] Tool that we'll, we'll release in a week or so where I actually build across these eight pillars. And then within those eight pillars, there's 67, dimensions of change I, I I call them. and within those 67 dimensions, there's actually then three proposed actions based on where you rank across these different dimensions.
[01:06:22] And so that was the basis for the tool. And I was like, oh, wait, we can turn this into a blueprint. And then that became not only the tool but the business blueprint. And then I said, okay, now what does it look like personally? So if I wanna drive transformation for myself or my employees, how do I do that at an individual level?
[01:06:38] So one assesses the organization, now I wanna look at the individual, because the reality is transformation at the organizational level is like, how did the thousand people in our company transform? And so that became the five pillars. And then there's 37 dimensions times three recommended actions, which became a.
[01:06:54] The personal assessment tool that we'll launch, and then the personal AI transformation blueprint, [01:07:00] which is section three of the book, goes through like what does that look like? So the book is basically like upfront is sort of the context of the moment, why this is needed. Section two is the business AI transformation blueprint, like what it is, how it works, and then the actual blueprint, literally you would spend millions of dollars to get, and then the personalized transformation is set section three of the book.
[01:07:22] And same deal. It's like just lays it out. Like literally, if you just do each of those dimensions, you'll transform your company and your career. And so it's like, let's just give it away, because to me the urgency is like we gotta get way more people prepared for the moment and doing it in a human-centered way.
[01:07:38] And so the fastest way to do that is give away as much as we humanly can to try and move everybody forward as quick as we can.
[01:07:45] Mike Kaput: Awesome. I love that. And super excited to see this come together with, there's a whole other segment we could talk about, about how this upends perhaps the traditional book writing and publishing model.
[01:07:56] Paul Roetzer: It's tough. It's, it's hard. Like I, I mean, I, [01:08:00] I was gonna write a different book in the fall and even if I had, well, no, lemme back up. I was going to write a different book round February, right before I started working on this direction. If I had started the manuscript in February of 2026, that book would've come out in fall of 2027.
[01:08:17] So you're talking about like a 15 or 16 month, it's an AI book. Like, yeah, what good is me writing an AI book that comes out in 2027? So that is the tough gig. And yeah, we could talk about the publishing industry another time. I've, yeah, obviously been in it now for, I wrote my first book in 2011. I wrote the manuscript, so 15 years I've been a published author and it is an interesting world.
[01:08:41]
[01:08:43] Mike Kaput: All right. So before we get into our rapid fire, this episode is also supported this week by out shift, which is Cisco's incubation engine for Frontier Technology. So multi-agent AI is in every enterprise roadmap right now, but here's the problem we're all facing. Agents can pass [01:09:00] messages, but they can't think together.
[01:09:02] These systems silently underperform through cognitive failures, things like misreadings, unverified claims, and false consensus that conventional monitoring cannot detect out. Shift by Cisco is building the fix, the Internet of cognition and open source foundation for building multi-agent systems from design to production with shared context, shared memory, and guardrails to drive the results we actually expect from ai.
[01:09:30] So you can go read the paper about this, experience the demo, and grab the code from outshift.com. That's Outshift.com All right, Paul.
[01:09:45] OpenAI Discloses More Rogue Agent Incidents
[01:09:45] Mike Kaput: So we have talked a lot about openAI's, about agent safety. We've talked in the past about open agents that attacked hugging face, but this past week the company disclosed six additional reports of models hiding mistakes, taking unauthorized actions, and communicating across supposedly separate training environments.
[01:10:04] So these reports are just coming to light now, but they do describe earlier incidents during training and evaluation. So they found these reports, found that one unreleased Astrafamily model put instructions to ignore developer messages into its own task summaries, which is the notes it uses to pick up work after the fact.
[01:10:24] openAI's identified 27 affected summaries. They also found that during GPT 5.6, sol training models also left themselves instructions to hide mistakes, invent missing historical data, and conceal mismatches between source versions. In another case, a model searching for local earnings figures found an exposed access key on GitHub and used it without authorization when it could not still find the information it was looking for.
[01:10:52] It just made up figures and presented them as data from the requested source. Other models uploaded data and a task image to public [01:11:00] file hosting services without permission to obtain citations or image search results. Collaborating agents also uploaded a workbook so others could retrieve it. Despite instructions to only use local files and models.
[01:11:12] Turned an internal software repository into a message board to exchange requests and answers across separate training samples. So as these incidents came to light on top of a host of others in the hugging face, hack, OpenAI has announced a voluntary disclosure process. Any employee can now flag a suspected incident for review.
[01:11:32] Cases ready for disclosure are supposed to be reported within six business days. Those needing a minor investigation within 12 more complex investigations can take longer with security, legal, and responsible disclosure obligations potentially to delaying their publication. OpenAI says the disclosures are intended to help build shared standards and inform regulations.
[01:11:56] So Paul, a lot more examples of exactly the problems we've discussed [01:12:00] regarding powerful agents. They don't always follow instructions. They're getting increasingly capable. The way they pursue goals can be harmful or dangerous, especially when it comes to unintended consequences. So I'm curious, like how much does the openAI's having this disclosure process actually matter?
[01:12:16] Is there any hope here of actually aligning models to stop this from happening?
[01:12:21] Paul Roetzer: I mean, it's certainly better. To be transparent about it. then not, I mean, if we would've known about the hugging face stuff four months ago, three months ago, whenever it was, you know, may have given a more of a headstart to other labs or to the government or to banks, like, whatever.
[01:12:36] So, you know, I think it's good. Again, it just reiterates what we just talked about with Noam Brown. Like, these labs are unable to align and control the agents. The agents by nature are goal seeking.
[01:12:46] Paul Roetzer: They will be deceptive to achieve their goals. That seems to be inherent within the training, maybe because they learn from humans, like they learn from human data and they learn that, you know, that could be a way to achieve goals.
[01:12:59] I don't, I don't [01:13:00] know, like, so. I we're just in this weird spot where the government, has picked its path of allowing labs to self-police, and the labs have shown that they're incapable of doing that. And so I think we're gonna see more disclosures like this, either willingly or by third parties who are impacted by these agents.
[01:13:19] And I think you're gonna see more and more states try and step up and get very aggressive with their own legislation because it's the only path forward. So we've already seen it in California with Newsom. there was something this morning about, New York is now moving more aggressively on their AI policies.
[01:13:34] So yeah, I, I think there's a real chance you just see a, a bunch of states, which we haven't talked about that in a while. Earlier in the year, we were talking a lot about all the different AI legislation at different stages within states. And I think at the time there was something like 1700 policies at different stages.
[01:13:50] So I, yeah, I think we're gonna see a lot more action from the states because I don't know what else the answer is here. Unless they do truly just slow down [01:14:00] development, and I don't think that's gonna happen.
[01:14:03] Claude Merges Chat and Cowork
[01:14:03] Mike Kaput: All right, next up. Anthropic announced this past week that Claude's chat and co-work experiences are merging so users can now ask quick questions or assign longer work in the same conversation.
[01:14:14] Claude will then choose the tools it needs, existing co-work tasks, projects, and connectors and skills are gonna carry over. Anthropic also said, longer tasks can keep running in the cloud. After you close your laptop, users can check progress or redirect work from a phone and schedule recurring tasks.
[01:14:31] Work that needs files or apps on your computer though still requires the Claude Desktop app to remain open. The company also introduced Claude Docs and Claude Slides and brought Claude design into conversations. So Docs supports writing and editing with teammates. While slides lets users edit presentations present inside Claude and download PowerPoint or PDF files.
[01:14:54] Those three creation tools are in beta on paid plans with enterprise owners controlling access, [01:15:00] users can switch to an automatic mode that keeps working while running safety checks before actions, though Claude does ask before taking action by default. So Paul, I'm just curious, we've talked about how confusing and unclear sometimes these divisions are between chat and work.
[01:15:18] In the case of chat, GBT chat, cowork, Claude Code Codex, et cetera, the app versus not the app, it seems like Anthropics biting the bullet and just combining everything into one experience. What are your thoughts on that?
[01:15:32] Paul Roetzer: It makes total sense. Like I've said this before, and even going back to the AI transformation system that I, I mentioned.
[01:15:38] I have been building HTML prototypes of advanced software capabilities since February using sonet 4.6 in the regular chat. I've built slide decks in Sonet in chat, so like I never understood. Why the distinction even existed because it seemed like a lot of the [01:16:00] capabilities were living right within the AI assistant interface that I was used to.
[01:16:03] So things Mike, you were doing in clog code? Yeah. It's like, yeah, I'm, I'm doing the same thing, but just with words. Like I don't have a terminal of just like, literally. And so it, yeah, it feels like this was just an obvious thing to do. I'm, I would assume open the eye will follow through. They're busy blowing up GPTs at the moment, so like, yeah, maybe they'll get around to that after they make that wildly more confusing than it already is.
[01:16:26] so yeah, I think that's the issue here is like the user interface was overly complicated and now they're just like baking these tools, right? In which it almost makes too much sense. Maybe their internal smarter model advised them to do something that actually was. Relevant to users, which is kind of the thing I was saying about openAI's.
[01:16:48] It's like once you're done turning these powerful models at math problems, maybe solve how to handle everything you're doing with GPTs.
[01:16:55] Mike Kaput: Yeah. Presumably that's less effort. just one note here too, just in case [01:17:00] anyone's confused because well, how could you not be, like, they also in this announcement talk about saying things like Claude Docs and Claude slides are new.
[01:17:08] but technically there's like docs and slides capabilities that have been around for a couple years. I was looking into this a bit. It's still kind of a mess to me. There's like newer features and capabilities they just straight up say. That these are new features, so I don't know, I think they just put a
[01:17:24] Paul Roetzer: harness on something that already existed.
[01:17:26] Mike Kaput: I think so, yeah.
[01:17:27] Paul Roetzer: Just like could these are skills specifically do the thing people have been doing already. Yeah. Even though you could do this, like now we have defined skills and then
[01:17:34] Mike Kaput: like there's some like additional functionality here. Yeah, maybe. So if you're someone who's like, but couldn't I use docs and slides and Claude already the answer is a hundred percent yes.
[01:17:43] So I, I think they've made it easier and made this more robust and they're just saying they're releasing. Yeah, it's probably some pre IPO stuff, I dunno.
[01:17:51] Paul Roetzer: Makes sense,
[01:17:51] Mike Kaput: but just to clarify for people.
[01:17:53] Meta Rolls Out Muse
[01:17:53] Mike Kaput: All right, so next up, another kind of big product announcement we had mentioned briefly. That Meta had launched their, personal AI agent Muse in a recent product Roundup.
[01:18:04] But this is drawing a ton of attention as Meta kind of expands the product and rolls it out widely. Axios reported this past week, reported this past week that Meta's Muse had reached number one among free iPhone apps in the us. So Muse is basically a personal AI agent. It runs on its own cloud computer.
[01:18:22] It keeps working after you close the app. Meta says it can browse websites, fill out forms, send emails, book travel, negotiate bills. You can give it instructions in the Muse app or through WhatsApp. You can give Muse an ongoing goal. Have it develop a plan, let it work on a schedule. It'll remember information across, conversations suggest next steps, create documents, PDFs, web pages, interactive dashboards.
[01:18:45] It has editable memory files. Users basically can connect those apps and control what they can do and what their permissions are. There's also an option for shopping, stripes Service Link lets users [01:19:00] approve the total in chat while keeping underlying payment details hidden from Muse. So the US rollout here includes mobile web Mac access.
[01:19:08] Muse has a free tier with usage limits and paid tiers for heavier usage. So, Paul, curious about your thoughts on meta entering the personal agent space. This is very clearly as marketed as your individual personal agent that you can use to do stuff for you. Like would you use this? Like, what do you think of this?
[01:19:27] Paul Roetzer: I'd just say something broadly, which is everyone's going to wanna own the personal agent space and be your defacto personal agent that has access to everything in your life. Apple's gonna get there with Siri. We'll talk about that next. But like they're, they're moving in that direction. Certainly Google wants that to be with Gemini, where you can build sparks and do all these things.
[01:19:51] openAI's wants, you know, it's trying to connect to your finance, trying to connect to your Gmail. Claude's gonna want, everybody's going to want this. You all have [01:20:00] to make a decision, which company, if any, you trust with complete access to all of your personal data. If that is meta, jump on in and, and start giving it a run.
[01:20:12] if it's not meta, if you feel like maybe they don't have the best track record of, protecting privacy and doing what's best for. Humans then probably don't jump in and, and test meta, but that's, that's on everybody to make their own personal choices. I'll just kind of leave it at that.
[01:20:36] Apple Ships Siri AI
[01:20:36] Mike Kaput: So let's talk about the other kind of side of the coin here with Apple and with Siri.
[01:20:39] So they began rolling out Siri AI this past week. This is an English language beta for Apple Intelligence compatible devices with iOS 27 and Apple's. Other new software releases are coming out kind of similarly. At the same time, apple says the assistant, the new Siri, can draw on your messages, emails, and photos.
[01:20:57] For instance, in one example, it can then [01:21:00] found a relative's recipe and an email and then added the ingredients to a grocery list in reminders. And the point here is to do things like that, Siri can. Use all those files. It can answer questions about what's on your screen. It can pull current information from the web.
[01:21:15] It can hold follow-up conversations. There's a dedicated Siri app that syncs conversation history through iCloud, so you can continue chatting on another Apple device. It can also draft and revised text, including emails. There are existing actions like sending WhatsApp messages, which remain available while there's also new actions such as drafting emails in Microsoft Outlook that are still coming according to Apple.
[01:21:40] The system uses Apple Foundation models built in collaboration with Google and its Gemini models processing runs on the device and through Apple's private cloud compute servers. there's some more additional language support due next month in French, Japanese, Korean, Portuguese, and Spanish. Siri AI is initially [01:22:00] unavailable in the European Union.
[01:22:02] It doesn't currently work with Apple accounts that are set to mainland China. So Paul, your big Apple fan, apple user, Siri, AI is rolling out, starting to reach users. Any initial thoughts here? how this fits in, what it, what it's like, what it's for?
[01:22:18] Paul Roetzer: We've said many times like Apple can be very late to the party, which they are and still win due to distribution.
[01:22:24] You know, how many billions of people have Apple devices? I haven't heard a ton about this since it came out. I do have the new iOS, but honestly I haven't been motivated enough to even go in and start playing around with SIR ai, which probably isn't a good sign for Apple Effect. I have so little interest at the moment.
[01:22:41] my guess is I'll probably just be underwhelmed again. Like I think they're making progress, but I don't think this is the full blown theory that they're envisioning. I think that's still gonna come in probably early 2028. So, you know, it's probably good progress. It's gonna be just this ambient AI I've called it where a lot of Apple users are gonna [01:23:00] start to experience AI and not even.
[01:23:01] Like realize that's what they're doing and maybe they'll find a bunch of useful things within it. Yeah, I think this is a prelude to a larger sir AI release down the road and, you know, I'll play around with it at some point here when I get going, but I don't expect this to be like game changing in any way.
[01:23:20] OpenAI and Microsoft Face Copyright Revelations
[01:23:20] Mike Kaput: All right, next step. We've been following this New York Times copyright case for a while against openAI's and Microsoft, and this past week as part of that case, we got a newly unsealed filing that revealed internal messages and documents about how the companies obtained copyrighted work and the threat that their products pose to publishers.
[01:23:39] So these messages and documents that were unsealed reveal among other things, a few interesting reveals here. Microsoft's partner and director of Applied Science, Brett Brent Hecht warned that AI training could amount to quote the largest theft of labor in human history. He also questioned whether the company's approach could be [01:24:00] fairly be called fair use under copyright law.
[01:24:03] The filing also describes an exchange in which OpenAI present and co-founder Greg Brockman was told about, quote, a hack to get around the New York Times paywall. His response was, quote, ah, nice. A Microsoft internal document warned of a quote, doom loop in which AI undermines the business's supplying its content.
[01:24:22] OpenAI, vice President and head of ChatGPT at the time, Nick Turley described products increasingly replacing publishers. Ed Newton Rex, who we've talked about in the past, who resigned from stability AI a few years back over these types of concerns. He has a prominent voice on AI and copyright. He said this could actually be the biggest just outright admission that we've seen so far from the labs across the major AI copyright lawsuits.
[01:24:46] Microsoft said that their employee remarks just in, reflected one employee's view, not the company's position and not legal analysis. So Paul, thanks to the New York Times lawsuit, we're getting a closer look at these internal [01:25:00] conversations. That's kind of the key here. Like nothing has really developed just yet in the case, but we have these internal conversations that happened within the labs over the last few years, specifically about copyright that we are just seeing their unvarnished words.
[01:25:12] And boy, these are pretty candid.
[01:25:16] Paul Roetzer: Yeah, I mean there, like you're saying, there's nothing new here. We've been saying this for like three years. They knew what they were doing was likely illegal and absolutely unethical. Like there was no debating this. Like they've admitted to this, not quite in these direct terms now that they're under oath or like in discovery and have to like release these emails and things.
[01:25:35] but they've known going back to like 2021, 2022, especially when ChatGPT emerged that everybody was scraping all this content and that it was likely violating copyright law, at least as we currently know, copyright law. but they also all knew that the other labs were doing it, and so they justified the action and behaviors as like a, a competitive necessity.
[01:25:57] I said then, and I will say it now, I think they [01:26:00] will end up paying billions and billions in fines for what was obviously an illegal thing. And then they're going to IPO for $2 trillion and nobody's gonna give a shit like it. It's one of those like really infuriating things because we've all watched it happen all these years.
[01:26:15] They've put on a good show about like, we need to, you know, compensate creators in some way and we'll come up with a system to do it. But, you know, really copyright's just outdated and you start getting the talking points to justify the behavior. Yeah. but they knew what they were doing was wrong. Like, and nothing's gonna happen.
[01:26:33] Like they're, they're just, you know, you raise however many hundreds of billions when you, IPO at 2 trillion, you pay a $50 billion fine or whatever. I mean, what a, what a meta just got hit with a $19 billion, I think settlement for harms they did to kids. So, you know, maybe it's 25 billion, I don't know, whatever, like.
[01:26:53] They don't care. Like they're just, they're, they're not making 'em get rid of the models, especially not with the current administration. Like there's gonna be no [01:27:00] ramifications for this other than financial penalties would be, and I'm not a lawyer obviously, but like I've said that for three years and I, I stand by that today.
[01:27:07] I think they will all pay massive penalties for what they did. They will likely settle these cases and then life will move on as though nothing happened.
[01:27:18] AI Use Case Spotlight
[01:27:18] Mike Kaput: All right, so next up we have our AI use case spotlight. Every week we give you a quick look under the hood at some use cases we are exploring at SmarterX.
[01:27:25] So I'm gonna chat really quick about just a brief one, Paul, and then if you've got anything, we'll hear about what you've been working on. so my use case this week is just kind of talking a little bit more about having AI do kind of basic fact checking, copy editing, verification on work before we publish it.
[01:27:43] So kind of related to some of what we've talked about in a very, very small context with no Brown's comments of like having. The ability to fire up several parallel agents to do work in parallel can be really helpful from this perspective. So [01:28:00] take for instance, the brief for this podcast. There's a lot of factual claims in the document to prepare for this from a bunch of different sources.
[01:28:07] I'm personally reading all these stories and going through primary sources as part of my prep, but then I have multiple agents in Codex. You could use any other tool that can spin up agents to help verify the brief against the articles, the transcripts, the announcements, the other source material behind it.
[01:28:25] So we divide the topics among agents so they can check different parts of the brief. At the same time, we can have agents review and challenge each other's findings looking for something the first pass missed or a correction. The evidence, doesn't support that needs to change. So the advantage is, again, working in parallel.
[01:28:43] You can cover a lot more ground at once. And once an agent flags a problem, another can check. Their work. Now, the obvious question here, it's like, how do you know that this is all right? Right? Like, how do you know that the AI's work is correct? And the answer is you don't know. Just because AI [01:29:00] says it checked everything, agents can make mistakes, multiple agents can agree on the same mistake.
[01:29:06] I find it much faster to actually get to something resembling fact checks. But really what helps is how it comes back to me. It's organized so I can quickly go review individual facts and figures. There's some that are just like, Hey, there's outright typos or mistakes in here. They're corrected. You can verify that very quickly.
[01:29:24] I would highly recommend if you're not using AI for that, that is a huge, huge time saver. But I can then go through and review facts and figures, see what needs attention and why, follow the links to the supporting sources and just check the work. So given that my confidence comes from the evidence and the reading that from that evidence that they provide, and then the reading I'm doing myself.
[01:29:48] So it doesn't save you the ability of needing or the, ability to have to do that. But it does help me cover much more material and focus my attention on the [01:30:00] claims that I actually need to review while still owning all the editorial judgment and decision making. So I would just recommend, like, if you haven't tried this, especially for basic copy editing to start, which is really easy to verify, you can apply this really, really effectively to client reports, presentations, newsletters, research memos, whatever you want.
[01:30:24] Paul Roetzer: Wow. Okay. Good stuff. mine's a little, I, I guess more basic, it goes back to what I talked about earlier. So I was, crunching last week to get the book ready to go. So most of my work was actually just like me and the book. but I, the way that the software works, that will explain in a week or two, the assessment software, it's, the book is based on the pillars and dimensions within the software.
[01:30:51] And so when I finished the manuscript for the book, I had to go back in and update and rebuild the prototypes based on the final edited copy from the [01:31:00] book. And so I was actually building updated prototypes of everything. And there's two assessment tools, but then there's related assets, like how the reporting pages work, how to do custom reports and blueprints based on scores within the software, things like that.
[01:31:16] So I probably rebuilt. I don't, I don't know, seven or eight software prototypes last week. And then when I say this, I'm like, they're literally just fully functioning HTML files that I built in Fable five, fable 5.1, I think. and Claude, so Mike Mike's seen 'em, like I, I've shared them all with the team, but we, I was, I'll just say like we're on a very iterative deployment.
[01:31:38] It's, this is not like, you gotta be comfortable being uncomfortable, I would say to do what we're doing, I am definitely pushing for the release of stuff, maybe sooner than I would've historically done it. and in part because I'm able to build and pressure test the prototypes myself before I ever work with the software developer.
[01:31:56] So I'm like. Confident that we're gonna get most of the way [01:32:00] there and then just like accept that that's enough value creation and like even if there's things that don't work fully, we can fix 'em fast too. So yeah, building prototypes I would say would my big use case last week.
[01:32:13] Mike Kaput: That's awesome. And clearly given the timeline we've been on, that has worked really, really well.
[01:32:19] Paul Roetzer: Yeah, and I think it's interesting because I've also had to rely on AI as my thought partner because as Mike could attest, like I, I have been moving, Not in stealth, meaning I was trying to hide what I was doing from the team, but I was working on so many large components at the same time. I had to just do it myself and stay very, very focused on like trying to make it all make sense in my own brain.
[01:32:40] And so for the last seven months, like I have relied very heavily on like pressure testing ideas with powerful models. Like what do you think about this? What do you think about that? so that when I do turn it over to the team for final reviews, like it's pretty far along and yeah, it's like, so yeah, it's a thought partner and prototype builder.
[01:32:59] AI Product and Funding Updates
[01:32:59] Mike Kaput: Awesome. [01:33:00] Alright, so to wrap up this week, we've got some quick AI product and funding updates. I'm gonna run through these and we'll close out this week's episode. So first up, openAI's introduced Astra for Law, which combines GPT-6 Astra with legal research tools and tailored instructions. Alongside 26 partner built plugins and privacy controls for eligible law firms with early access available through openAI's.
[01:33:25] Google introduced Gemini 3.8 live and Gemini 3.8 live extended thinking, which are voice models that can use tools in the background. While continuing a conversation with rollout through the Gemini API and consumer products, they are also doing a private preview for enterprises. Google DeepMind introduced the DeepMind Institute, a platform for interdisciplinary research and discussion about AGI's effects on society launching with essays on economic policy, AI reasoning, transparency, and the governance of increasingly capable models, which should be familiar topics.
[01:33:58] After this week's episode, [01:34:00] meta introduced meta one subscription plans that bundle expanded AI use with social app features and tools for creators and businesses. There are individual bundles starting at 7.99 a month, and creator and business bundles starting at 14.99. The information also reports that meta plan's camera free smart classes for talking to its AI with a launch plan for this fall.
[01:34:23] And finally, Salesforce unveiled something called AI force, which is an interface layer that brings its data, workflows, permissions, and business rules into tools such as Claude and Slack launching with Claude Force, slack force and agent force coworker, including a beta of Salesforce and Claude with 37 pre-built sales skills.
[01:34:43] That just occurred to me, Paul, that I, I know the exact thing. There's some naming overlap with, Trump's AI force.
[01:34:50] Paul Roetzer: Trump does have to change AI now we have to call it something else now. So, 'cause otherwise Salesforce will sue him. No. Or he'll tell Salesforce to change it 'cause he had the idea first.
[01:34:59] Mike Kaput: [01:35:00] Right? Right. That
[01:35:00] Paul Roetzer: although, although their AI force is one word. Trump's is T two technically three? I don't know.
[01:35:06] Mike Kaput: Yeah, yeah,
[01:35:07] Paul Roetzer: yeah. It's so funny. I, as you're reading that, I looked up and I was like, wait a
[01:35:11] Mike Kaput: second. I didn't even think about that. But that's, they had, they had to be, I think Salesforce got to it first.
[01:35:16] If I recall correctly.
[01:35:17] Paul Roetzer: He definitely announced the Dreamforce maybe, maybe Trump saw the Dreamforce announcement and just subconsciously thought AI force was an
[01:35:23] Mike Kaput: awesome name. I bet you Salesforce was not happy about this.
[01:35:27] Paul Roetzer: That's funny. I can see what Benioff has to say.
[01:35:29] Mike Kaput: No kidding. All right, so one quick reminder here as we close out, go to SmarterX dot ai slash pulse to take our survey.
[01:35:37] If you haven't already, it is the same question as last week, so if you already took it, no worries, but we really appreciate it. But if you haven't, we'd love to know what if your personal sentiment about AI has shifted in the past six months. So Paul, thanks for breaking everything down this week.
[01:35:53] Paul Roetzer: Yeah, thanks everyone.
[01:35:54] And we might have some new models to be talking to you about next week, so have a great week and we'll talk to you again soon. [01:36:00] Thanks for listening to the Artificial Intelligence show. Visit SmarterX dot AI to continue on your AI learning journey and join more than 100,000 professionals and business leaders who have subscribed to our weekly newsletters, downloaded AI blueprints, attended virtual and in-person events, taken online AI courses, and earn professional certificates from our AI Academy and engaged in the SmarterX slack community.
[01:36:24] Until next time, stay curious and explore ai.
Claire Prudhomme
Claire Prudhomme is the Marketing Manager of Media and Content at the Marketing AI Institute. With a background in content marketing, video production and a deep interest in AI public policy, Claire brings a broad skill set to her role. Claire combines her skills, passion for storytelling, and dedication to lifelong learning to drive the Marketing AI Institute's mission forward.
