OpenAI, Anthropic, and Microsoft all shipped in the same week: GPT-6 Sol and Luna, Claude Opus 5.5, and a redesigned Copilot with Home, Code, and Autopilot. None of them have a giant leap in raw intelligence and that's the point.
Paul Roetzer and Mike Kaput explain why this generation is about reliability and "per-task pricing" rather than raw power and what that shift means for every-day knowledge work.
Then: Jensen Huang's case against the AI doomers from his Ezra Klein interview, a first look at AI Score by SmarterX, and a rapid-fire run through of Meta Connect, Amazon blocking Muse, OpenAI's rogue-agent disclosures, and more.
Listen or watch below and see below for show notes and the transcript.
This Week's AI Pulse
Each week on The Artificial Intelligence Show with Paul Roetzer and Mike Kaput, we ask our audience questions about the hottest topics in AI via our weekly AI Pulse, a survey consisting of just a few questions to help us learn more about our audience and their perspectives on AI.
If you contribute, your input will be used to fuel one-of-a-kind research into AI that helps knowledge workers everywhere move their companies and careers forward.
Click here to take this week's AI Pulse.
Listen Now
Watch the Video
Timestamps
00:00:00 — Intro
00:06:57 — GPT-6 Sol and Luna, Opus 5.5, and the New Microsoft Copilot
- Introducing GPT-6 Sol and Luna - OpenAI
- X Post from OpenAI
- X Post from Sam Altman
- Better prompt caching for GPT-6 - OpenAI
- Introducing the new Copilot with Home, Code and Autopilot - Microsoft
- X Post from Satya Nadella
- X Post from Peter Steinberger
- Introducing Claude Opus 5.5 - Anthropic
- X Post from Claude
- X Post from Artificial Analysis
00:24:16 — Jensen Huang vs. the AI Doomers
00:41:57 — Introducing AI Score
00:53:12 — GPTs Are Being Retired
- ChatGPT Enterprise and Edu release notes - OpenAI
- Plugins in ChatGPT and Codex - OpenAI
- Custom GPT retirement and migration FAQ - OpenAI
01:06:30 — Meta Connect 2026
- The Biggest News From Connect 2026 - Meta
- Meta Connect 2026 Developer Recap - Meta
- Meta Debuts a Dedicated Palm-Sized Muse Charm Device to Use AI on the Go - Bloomberg
- New Features for Meta Ray-Ban Display - Meta
- X Post from Mark Zuckerberg
- X Post from Chris Bakke
- Zuckerberg says Meta will lead on smartglasses privacy despite ‘pervert glasses’ criticism - NBC News
01:12:39 — Amazon Blocks Muse
- Meta’s standoff with Amazon over Muse could be a sign of things to come - CNBC
- X Post from Elon Musk
01:15:20 — OpenAI Discloses More Rogue Agent Incidents
- X Post from Sam Altman
- X Post from OpenAI
- OpenAI Works to Understand Full Scope of Agent Activity as User Data Leak Emerges - Reuters
- OpenAI Hugging Face Hack - The New York Times
- OpenAI Agent Hacked Australian Government Website, Albanese Says - Bloomberg
- OpenAI AI Breach in Australia - The New York Times
- X Post from Nikita Bier
01:20:00 — OpenAI Gets Vocal on AI Standards
- Building standards for the next phase of AI - OpenAI
- OpenAI, Anthropic Neared Deal to Stress-Test Each Other’s AI - The Information
- Altman, Amodei Call for Global Cooperation on AI to Boost Safety - Bloomberg
- X Post from Sam Altman
- A Blueprint for a Federal Framework - OpenAI
01:22:59 — AI Use Case Spotlight
01:26:36 — AI Product and Funding Updates
- Meta Announces Enterprise Platform
- Grok 4.7
- Google Readies AI Chips for Space
- Claude Discovers a New Enzyme System
- Claude discovers a novel enzyme system with CRISPR-like repeats - Anthropic
- X Post from Dario Amodei
- OpenAI Develops Features to Counter Grok Bot, Mulls Response to Meta’s Muse - The Information
- Google, OpenAI, Anthropic AI Safety Group Takes Shape - The Information
- Google Nears Release of Flagship Gemini 4 AI Model - The Information
- Advisory Group on Mathematics and Artificial Intelligence - OpenAI
This episode is supported by Outshift, Cisco's incubation engine for frontier technology.
Multi-agent AI is in every enterprise roadmap right now, but here's the problem we're all facing: agents can pass messages, but they can't think together. These systems silently underperform through cognitive failures —misreadings, unverified claims, and false consensus—that conventional monitoring can't detect.
Outshift by Cisco is building the fix, the Internet of Cognition. An open source foundation for building multi-agent systems from design to production with shared context, shared memory, and guardrails to drive the results we actually expect from AI.
Read the paper, experience the demo, grab the code from Outshift.com. That's Outshift.com
This episode is also brought to you by Marketing AI Month. All month, the AI for Marketing Core series inside AI Academy is free (a $499 value): five expert-led sessions from Mike Kaput, the frameworks and tools the SmarterX team actually uses, and a professional certificate on completion. Enroll by September 30 (you don't have to finish by then - just enroll).
The month closes with a live AMA on October 1, where Cathy puts your questions to Paul and Mike; anyone enrolled can attend.
Enroll at SmarterX.ai/marketing.
Read the Transcription
Disclaimer: This transcription was written by AI, thanks to Descript, and has not been edited for content.
[00:00:00] Paul Roetzer: He's forgetting that like a large portion of the population of future workers hate ai, like want nothing to do with it. So it's wildly unpopular. And so you can't just assume that they're all just gonna be AI natives and love it, and they're just gonna like, it's just gonna be amazing. Welcome to the Artificial Intelligence Show, the podcast that helps your business grow smarter by making AI approachable and actionable.
[00:00:24] My name is Paul Roetzer. I'm the founder and CEO of SmarterX and Marketing AI Institute, and I'm your host. Each week I'm joined by my co-host and SmarterX chief content Officer, Mike Kaput. As we break down all the AI news that matters and give you insights and perspectives that you can use to advance your company and your career.
[00:00:45] Join us as we accelerate AI literacy for all.
[00:00:52] Welcome to episode 243 of the Artificial Intelligence Show. I'm your host, Paul Roetzer, along with my co-host Mike Kaput. [00:01:00] We are recording Monday, September 28th, 9:00 AM Definitely expecting some news this week. so time stamping is relevant here. There's all kinds of stuff like I'm waiting. Mike, I dunno if you saw anything, but like Trump and Dario Amodei, were supposed to be having dinner Sunday night at 10:00 PM I haven't heard the outcome of that dinner.
[00:01:20] We've got openAI's Dev Day this week. There's rumors of, Gemini four potentially dropping at some point soon. So like. I think we're heading into fall, or we, I guess we're technically in fall weekend now, but it is, it's gonna be a very busy fall season it looks like in the AI world.
[00:01:38] Mike Kaput: Yeah.
[00:01:39] Paul Roetzer: So today's episode is brought to us by MAICON.
[00:01:41] This is the AI Conference for Marketing and Business Leaders happening October 13 to 15 in Cleveland, Ohio. This is our seventh mike. Does that sound right?
[00:01:50] Mike Kaput: Yeah.
[00:01:51] Paul Roetzer: Yeah. I think we launched MAICON in 2019, so the seventh run. So you've been thinking about joining us from MAICON. This is the time we have. [00:02:00] I guess two weeks left, Mike, which is terrifying for me because I have like four presentations at Metcon and I am in the midst of trying to build all of them simultaneously.
[00:02:11] but definitely check it out. We have a special offer going on. Mike and I are actually gonna be hosting a private lunch exclusively for our podcast audience. during that lunch, you're gonna be able to ask Mike and myself, any questions on your mind about ai, your company, your career, what's going on at MAICON?
[00:02:27] 'cause this is gonna be, I think on day two, Mike, is it? Yes. Yeah.
[00:02:31] Mike Kaput: Yeah.
[00:02:31] Paul Roetzer: Okay. so yeah, definitely get in, use that when you register. So it's macom.ai, MAICON.ai. Use pod100 as your promo code. That will automatically get you an invite to that lunch. That's the only people that are gonna be coming to that lunch.
[00:02:47] So it's just to ask us anything session. while you enjoy some amazing food, we actually, our team does an incredible job with the food and beverage, at MAICON. especially the final day. You don't wanna miss that. It's like all [00:03:00] these amazing Cleveland foods. So, but you can register MAICON.ai. Get that a hundred dollars off and get the invite to that private lunch.
[00:03:08] so yeah, two weeks left. We would love to see you there. It's gonna be an incredible lineup. We have Karen Howe, but the bestselling author of Empire of ai. I actually just talked to Karen last week. We did a prep session for her keynote. I cannot wait for that. Andrew Yang's gonna be there. Kevin Rus coming out with his Chronicles of AGI, or AGI Chronicles book is dropping, I think like the week before Mayon.
[00:03:31] So he'll be there doing a book signing. Can't wait for that one. it's just gonna be amazing. So if you can get to Cleveland, October 13th of the 15th, do it. We've got the Rock Hall rented out for the opening night party. Hopefully the guardians will be in the midst of round one or two of the playoffs.
[00:03:47] So Cleveland is gonna be an amazing place to be. Fall in Cleveland is my absolute favorite time of the year. So check it out. MAICON ai, we would love to see you there.
[00:04:00] Alright, every week we start off with an AI pulse survey. We've been running these for two week sprints now to get more responses and we actually have some, great feedback on this one.
[00:04:07] We had almost 300 responses this time. So the question was, how has your personal sentiment about AI shifted in the past six months? This is pretty remarkable. It is basically split, so we have more excited 25%, about the same, 24% more cautious, 27% more worried. 25%. So it is like literally I'm looking at the pie chart and it is a perfect, almost a perfectly split pie.
[00:04:36] So, yeah. So thanks to everybody for participating. This was our hope is we could start to get maybe like three to 500 responses for each of these so we can start to make it a little bit more projectable data. Like right now, we. Treat 'em as like informal polls. So it's a good look at, how our audience feels about, what's going on in ai.
[00:04:55] So we'll have a new question up. Mike can share more about that at the end, but it's SmarterX dot ai slash pulse and you can participate in these each week. So, yeah, Mike, I don't know how you felt about that. That's pretty wild to see the balance there though,
[00:05:08] Mike Kaput: unreal. Yeah, it was pretty cool to get so many responses, but yeah, I can't say it surprises me that there's a significant portion of people that are getting more cautious and more worried based on the news of the last nine months or so.
[00:05:22] Paul Roetzer: Yeah, and it's, you know, today's episodes should be interesting because we're gonna try and present a little balance here. I actually talked to Mike about this last week that I've been feeling the weight of like, all of this a lot. And it's interesting, Mike, I think I shared it with you and I wasn't even gonna get into this today, but I guess it's super relevant to this question.
[00:05:41] I had talked about, I had forgotten about this. I was going back trying to find something I had said, and I actually, Google searched it and I was looking for, like, I ended up on YouTube videos and in March of 2023, so three and a half years ago, basically we did a [00:06:00] segment where I was talking about how I was like tired of all of the negative and the worry that was being caused by ai.
[00:06:08] And I had written something about the, like the five or six things that most excited me about ai. And I was trying to shift my own mental focus to like, be like, let's focus on the opportunities. Like, we know there's negatives. We know that there are, you know, risks associated. We, we know all these challenges, but like.
[00:06:26] Let's think about this from an optimistic perspective. And so today should be interesting because the second main topic is Jensen Huang's interview with Ezra Klein and talk about an optimist man. Jensen is the poster child for optimism when it comes to this stuff. So yeah, just kind of something to keep in mind and something Mike and I have been kind of conscious of that there's just so much negative about AI and, and we wanna make sure we're also doing our part to focus on the positives and all the opportunities.
[00:06:57] GPT-6 Sol and Luna, Opus 5.5, and the New Microsoft Copilot
[00:06:57] Mike Kaput: All right, so on that note Paul, let's dive [00:07:00] into I think some fairly optimistic or interesting updates this past week. So first step, we had a bunch of new model releases. So openAI's introduced GPT-6, Sol and Luna this past week. So they're basically bringing the advances from their top tier Astra model into faster, cheaper options for everyday work.
[00:07:21] The models launched in ChatGPT work and Codex for paid users and through the API. They did not launch in regular chat, but free and go users can access Luna in the desktop app. OpenAI says it cut API prices by 50% compared with the previous generation's promotional pricing. It has also improved prompt caching, which is an important feature that lets an agent reuse context instead of processing it again on an internal test that they built from conversations where users were flagging factual errors in responses.
[00:07:55] Sol made about half as many mistakes as its predecessor. So one [00:08:00] small test there, but it seems to be getting more factually accurate as we release some of these new models. Anthropic at the same time also release Claude Opus 5.5, which it says performs at the level of Claude Fable 5.1. On most work, Anthropic estimates a 40% cost reduction versus Opus 5 on typical workloads at default settings that combines lower token prices with more efficient use of tokens and the input and output price cut alone is 20%.
[00:08:31] Now Anthropic says Opus 5.5 generates output more than 30% faster than Opus 5 and early testers have found its writing clearer. Now the company is also at the same time raising five hour usage limits on Pro Max team and seat based enterprise plans and giving subscribers a rate limit reset they can save for later.
[00:08:53] Last, but certainly not least, Microsoft announced a redesigned co-pilot with Home Code [00:09:00] and Autopilot. So Home brings chat and co-work together with Word, Excel, and PowerPoint built into the experience. code lets people build apps that can run inside their company's Microsoft 365 environment.
[00:09:15] Autopilot is a persistent agent built on Open Claude that can monitor channels, follow up on conversations, and run recurring work. While you're away, Microsoft says, home and code will start rolling out through its Frontier program in the coming weeks while autopilot expands to private preview at the end of September.
[00:09:33] So Paul, a bunch of new interesting ways to use ai. The big news obviously being some of the new models, but I don't think we should sleep probably on the copilot news either. I'm curious what you took away from the new generation of models. We just got,
[00:09:49] Paul Roetzer: I was kind of anxious to see where we were gonna go after all the pacing AI conversations and essays and interviews and everything.
[00:09:57] So I think this might be [00:10:00] a preview of how is this gonna look moving forward. And what I mean by that is, you know, there was initially is everybody comes out with these models last week, and I think Grok 4.7 came out too. I think we'll touch on that maybe later. and everybody was like, oh, you know, I guess this is what pacing feels like.
[00:10:17] you know, we're just gonna keep dropping models. But the more I thought about it, like what they're basically touting here is more reliability, meaning they're more factual, more efficient. So at both tokens and per task, which we'll talk about and more aligned. And so like, those are the dimensions that they're focused on.
[00:10:35] And it honestly starts to feel more like an iPhone release. So like iPhone 18 comes out and it's like, is it dramatically different than 17? It's like, no, it pretty much looks and feels very familiar. It's just faster processing, better battery life, better camera specs, and like people buy 'em. And so that's kind of what these releases feel like to me is they're just more reliable, more efficient, more aligned.
[00:10:58] and the [00:11:00] other thing is I think they start to increasingly be. marketed for knowledge work. So it's, you know, it's more about here's how you can use it in your day-to-day work. It's less about solving, you know, millennium math problems and scientific breakthroughs, and more about, as a lawyer, as an accountant, as a consultant, as a marketer, as a salesperson.
[00:11:20] Here are the things that they enable. So I think that we'll also see a lot more talk about like post training and customization for specific industries, functions and roles. And then I think that the other thing that's gonna start to happen is to date a lot of the compute, a lot of the Nvidia chips, the TPUs, they've been dedicated to making the product more powerful, more generally capable.
[00:11:47] And now that we've reached these thresholds where it just feels like you could shut this off today and it would be so usable for the next five years. They're gonna start shifting a lot more of the compute [00:12:00] resources to, to safety and alignment. And so I, these are just like made up numbers, but roughly, I think based on our perception of what's going on with the labs, let's say in the last one to two years, like 80 to 90% of the compute that they're using has been going to research and development of more powerful models, AGI agentic harnesses, like the things that have made them so capable.
[00:12:24] And now I think you're gonna see a shifting to where it could be 50 to 80% of the compute starts to get allocated to safety and alignment testing and monitoring. Jensen Huang actually talks a little bit about this in the Ezra Klein interview that, that we'll get to in a second. But basically what's happening is the building of data centers.
[00:12:43] So the more data centers we build. The more compute that is available for testing, for inference, things like that, the more resources that they get dedicated to safety and alignment. so that's kind of like a high level I my overall take on what's going on here. The introducing of [00:13:00] Sol and Luna, just as a reminder where this naming is coming from.
[00:13:04] So they have moved to Celest, a celestial naming convention. So the way it works is Star Planet Moon, that's kind of the sequence. So Stars are the biggest planets, then moons. so this is moving away from things like might mini nano, you know, things like that that didn't mean as much. So Astra, the biggest model, means stars in Greek.
[00:13:25] So it fits that celestial naming Sol is Sun. Terra is Earth, so that's like planet. And then Luna is is a moon. So I went into our SmarterX ChatGPT account, and currently the models I have to choose from our GPT-6, Astra Sol and Luna. I still have 5.6 Sol, Tara and Luna, and then 5.5 is being retired.
[00:13:51] So those are, and then you can set like the, whatever they call it, like the amount of thinking it's gonna do Yeah. to like high and everything. So, [00:14:00] in their posts that we introduced, GPT-6, Astra the most intelligent and aligned model in the world. That was earlier in September. while the most, demanding and important projects still call for Astra, full depth work happens at different scales, rhythms and budgets.
[00:14:13] So, so, and Luna, I think Sol is kind of like the marquee model that's like the workhorse they're envisioning, and Luna is the smaller one, but they talk about it specifically. Being built on a same, or trained in the same way as Astra. but it's for specifically professional work, factuality coding, computer use and alignment to faster, more affordable models.
[00:14:35] Sam Altman tweeted, in regards to so, and Luna especially compared by per task pricing, which is the metric that should matter. I don't think there is anything competitive anywhere in the market. He said we want people to use tons of AI as important to being able to explore this renaissance in front of us.
[00:14:55] So now this is an important thing. This per task pricing. This is kind of a newer way [00:15:00] that they're describing this. The way to think about this is, they're trying to shift the way I think businesses think about it. So in the first half of 2026, there was all this talk about token pricing and trying to find the most efficient model, which includes going and using Chinese open weights models because they're more efficient per token.
[00:15:21] What Sam is saying is they're just making the models better overall. So even if. On a per token basis, their model is more expensive because it's just a better model. It'll actually do the work product you want that that per task output more efficiently overall. So the per task pricing means the total cost to get a piece of work done rather than the input or output of a token in it.
[00:15:47] So just something to think about. Now, interesting side note here. I don't think we cover this later on in the episode, but Boris Power, who's the head of applied AI research at openAI's, did an interview last week and he [00:16:00] said 80 to 90% of research at openAI's is focused at GPT seven, GPT eight, and beyond.
[00:16:08] That is the first time I've. I've heard them publicly even talk about seven and eight, same, more or less the fact that he's saying that's where all of their compute is going.
[00:16:18] Mike Kaput: Right.
[00:16:19] Paul Roetzer: So I don't know, like, we'll come back to that one. I, in terms of like what the significance of that comment is, opening that researchers are saying a lot of stuff publicly right now.
[00:16:29] They were not saying two months ago. So just interesting to track on. Opus. So again, just like contextually, I go into our cloud account. We have a Claude, I don't think we have an enterprise account, but we have a Claude business account. My options in there, fable 5.1, which is again, the sort of distilled version of Mythos, which is the one that they've never released to the public yet.
[00:16:53] myth that Mythos is the one that, the government, you know, shut down basically, I think. so Fable five is their [00:17:00] most powerful Opus. 5.5 is the new one that kind of. Seems to compete with Fable, but more efficiently, there's sonnet five, haiku 4.5, and then you can set effort at low, medi high, extra and max.
[00:17:16] And then they actually have more models, that like some of the older ones. So 5.5 they claim, performs at the Fable 5.1 level, but costs 40% less. they did address this pacing the frontier. In the announcement post, it says it's their first release since they called for pacing the frontier. They said it was tested by external evaluators.
[00:17:40] On their automated behavioral audit, the most comprehensive alignment test they run, it was the strongest performing model they've tested to date. They talked about some of the key dimensions as being performance, safety, cost and speed communication, meaning it talks more naturally, basically. There was a lot of like, NOx on Opus [00:18:00] five as like, it was sounded more like a machine basically.
[00:18:04] they talk about its, its reliance, from a knowledge work perspective. They reference cost per task. So they said the Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. So again, I think that's a term you're gonna start to hear a lot more and they get into the safety side.
[00:18:24] Now on co-pilot, you touched on some of the things they're doing here. I would just say at a high level,
[00:18:35] I think what we're gonna see is the shifting of the user experience continually. And it's kind of jarring honestly to businesses. So what I mean by that is we talked two or three weeks ago about how Anthropic. Was taking, when you would go into Claude in the desktop, you would have chat and work, and then they combined those to now it's just chat.
[00:18:55] It looks to me, based on the announcement from Microsoft, they made [00:19:00] all these plans before Anthropic did that and, and Microsoft realized that was a better experience because it actually says right now, coming soon, you won't need to choose a mode. You'll simply state what you need and copilot will route the work to chat cowork or code.
[00:19:19] So it's like Anthropic simplified the user experience, but Microsoft just made this big announcement where you still have to choose. Which version of it you're gonna use. and that does not seem like the viable thing. So I think that we're gonna just keep getting these shifts. Like, we'll talk about GPTs being sunset by openAI's in a, in, in a couple minutes.
[00:19:38] same thing, like businesses are gonna get used to working within these platforms in a certain way. They're gonna train it, they're gonna put it into their learning and development programs, and then it's just. It's gonna get blown up. So I, I've actually been thinking a lot about this. I wrote an essay over the weekend on, like the future of AI education and Enterprises.
[00:19:57] I know, Mike, You read the essay over the [00:20:00] weekend. I'm probably gonna release that this week because I think what's happening here is very, very important for organizations. Certainly SMBs, but more importantly for big enterprises that as you're looking at scaling AI, l and d programs, the dynamics of how fast, not just the models are changing, but the user experiences are changing and how you train people to use them.
[00:20:26] The current way learning and development is done does not support this, the pace of change that's coming from these companies. And so I had a lot of thoughts about it. It was one of those things, like I just sort of sat down Saturday morning to work on my MAICON stuff and I ended up like writing for five hours, about sort of what I was thinking the future of AI education needs to look like in enterprises.
[00:20:48] So part of this, like played into that. So, we'll, we'll talk again. My hope is to release that essay, probably Wednesday. I guess that would be like the 30th. And then we will come, we'll, we'll talk about it. [00:21:00] next week. Now, two other quick notes. openAI's Dev Day is September 29th, the day this episode drops.
[00:21:06] So we expect a lot more news from openAI's this week. And then the Gemini four evals that I mentioned at the beginning have started to leak online, and that normally means that, it's, IM, you know, model releases imminent. so we'll see what happens. I would not be shocked if openAI's actually, or Google does what openAI's does to them all the time, and they don't drop Gemini for during dev day to, you know, steal the thunder from openAI's.
[00:21:32] Mike Kaput: You know, it's worth belaboring the point which we've talked about many times on this podcast, that you are really seeing with stuff like Opus 5.5 and GPT-6 Sol and Luna, like w, the rate at which the capabilities are growing and the costs are dropping is something you should not ignore. Yep. I have a lot of people that worry about token costs, understandably so.
[00:21:59] [00:22:00] But I think you need to really be planning around use cases and pushing the boundaries on what's possible and kind of not worrying about it is hard to say, I realize at the enterprise level, but at the individual level, suck it up and get the $200 a month max or pro account and push what's possible if you can, in your business, if your means allow, because soon that is not going to cost.
[00:22:26] It's not going to cost what it costs today to do these things, and you need to be ready when it doesn't.
[00:22:31] Paul Roetzer: Yeah. And Mike, the other thing, it's a great point. the other thing I would push for is, and this goes back to the I education piece. someone properly trained to use this technology is going to get dramatically better cost per output or cost per, per tax, or per task efficiencies than someone who has no idea what they're doing.
[00:22:55] And so what I mean by that is like, I'll talk in a couple minutes about this new AI score tool that [00:23:00] I built, when I was building it. I knew when to use Fable five and when to use sonnet. Yeah. Like, it's like, okay, I'm gonna build the prototype right now, let me pay the 200 bucks for the extra credits for Fable five, because now I need the best output.
[00:23:16] But when I'm like bouncing around, just thinking things through, gimme sonnet. Like, and, and, and then like the way I work with an AI assistant is like, I keep a sandbox of the questions. I want to ask it, and I, and like the things to do next. I don't waste tokens with it. Like I'm not, yeah. Yeah. And so like the more efficient someone becomes at talking to these things, the cost per token becomes way less important than the user's knowledge of how to work with the machine.
[00:23:45] Mike Kaput: Right.
[00:23:45] Paul Roetzer: And so, yeah, it's like all this focus on let's just like. Reduce the cost by, you know, orchestrating to a different model that's like cheaper. It's like, no, like don't use the less capable model. Just get better at using the models that are actually [00:24:00] designed to do the work you do. So that's why I think like the way Sam was positioning is the right positioning.
[00:24:05] It's like, don't think about the token cost as much as you think about train the user first. Then use the model that's gonna give you the best chance to create the output you're seeking to create.
[00:24:16] Jensen Huang vs. the AI Doomers
[00:24:16] Mike Kaput: Yeah. All right. Next big topic. This week, this past week, Nvidia, CEO Jensen Huang joined the Ezra Klein Show to argue that the fears about losing control over AI are overstated.
[00:24:29] Wong described AI as software whose safety problems can be solved through engineering. Klein pressed him quite a bit during his interview, on why leaders and researchers at the Frontier Labs are still warning that they may not be able to control what they're building. So Klein kind of raised what you might call the AI DOR fears about AI presenting existential threats to humanity and all these cyber attacks and incidents we've seen recently, and Juan called.
[00:24:59] The [00:25:00] bigger picture fears irresponsible and said they were not grounded in science or research. He actually argued that these types of alarming predictions could scare young people out of going to college and building and investing in their future because they fear they won't get a job. So Wong kind of talks about this by identifying two problems that are at play here, especially when it comes to things like openAI's.
[00:25:25] Agents hacking, hugging phase. So one is containing the systems and two is getting them to follow rules. He admits there's like a full problem here. He says that labs should not be shipping products they cannot control. And Wong said that if labs cannot contain even their experiments because client pointed out that the systems involved in hugging face were not public products, then they should shut those down.
[00:25:49] So Juan kind of is just coming at this, basically is saying this is an engineering and product decision and company management problem that we're seeing these things happen, not some type of big. [00:26:00] Unrestricted existential crisis here where AI is going out of control. the two of them also disagreed in this interview about whether companies can be trusted to manage the risks.
[00:26:10] Wong said that for the most part, we've got existing laws and potential civil and criminal liability to handle these types of things. But Klein said, look, you know, despite having those things, the financial crisis for instance in America happened, there's competitive pressure that might supersede those types of laws and liabilities.
[00:26:30] so Wong basically keeps coming back to this idea, says, look, Nvidia dedicates roughly 80% of our efforts to verification. He argued that AI labs need a similar shift towards just making these things reliable and safe towards testing and safety. More compute should go towards evaluation and alignment, not just building more capable models.
[00:26:52] So Paul, I'm just curious. You had a lot of, I think, thoughts on this. It just sounds like. At the core, and I don't wanna put words in Jensen's mouth, but [00:27:00] he just sees this as an engineering problem. There's no existential AI risk, there's no weird God-like super intelligence that escapes its graders.
[00:27:09] There's just product decisions and hardware decisions and software decisions it seems like, is what he's getting at here. Am I wrong on that?
[00:27:16] Paul Roetzer: Yeah, I mean, I think that sums up his position pretty well. I think this is a extremely important. Interview. so first of all, like kudos to Jensen for doing the interview.
[00:27:28] He did not have to do this, I think it was like an hour and a half long, if I remember correctly. It was very contentious at times. And, and he knew that going in, like these two disagree on almost everything. and Jensen knows Ezra's positions on these things, so he. He sat down to an interview that he knew was going to be contentious.
[00:27:48] Ezra is a very, very skilled interviewer that has specific things he wants to drill into, and Jensen, I'm sure knew going in that was gonna happen, but they did it in a [00:28:00] civil and respectful way. And this is like, this is what's missing, I think, in the AI debate is like these very just open, honest dialogues with people with opposing views who can sit in a room and actually discuss something like this is what's missing in politics, honestly.
[00:28:14] So again, keep in mind this is the CEO of the most valuable company in the world, and his company is at the epicenter of the entire AI ecosystem. And honestly, like the epicenter of today's. US economy. Yeah. Like if Invidia wasn't doing what they're doing, the economy is nothing right now. So that, that's one thing.
[00:28:35] The other, and this is the thing, I alluded to you, Mike, I think when I saw you last week, the, one of the most important takeaways to me with this interview is sometimes you have. To do the work as a human. If you read the transcript to this only, or if you had your AI do a summary of this transcript, you miss almost everything that was important about this interview.
[00:28:58] So you [00:29:00] cannot get the same level of understanding into context without doing the work. When it comes to this, like you have to hear the tone in Jenssen's voice. You have to hear the challenges like the pauses, like it. That's everything. Like the way you actually understand how Jensen feels, cannot be seen in words on a page or bullet points that your AI assistant gave you.
[00:29:23] You have to actually, See what they're thinking and feeling. And so not only hearing it, but like watching it is an even another level of it because now you start to see the emotion behind it. So just a note for people like show up, like, you know, it, this is where I said like, if someone sends their AI agent to an like a meeting with me ever, like I'm done working with the person.
[00:29:47] Like if, if you can't give the time to someone to actually show up and be there yourself, then just don't show up. Yeah. Like don't send your agent to do it. So. Okay. Now I'll get off of my like preaching call. Okay. [00:30:00] So here is my, my, my kind of notes. And what I, the way I did this is I listened to this interview over probably four days in my car and then I just had an Apple note going and I would like pause the interview and I would just like voice dictate my notes.
[00:30:13] And I'm just gonna kind of go through these notes in they're ride. I haven't even edited these. so I said, I think you can sum up Jenssen's position in that he is an eternal optimist who believes human ambition and hard work will create an abundant future. All of the other arguments are just a distraction from creating the future.
[00:30:31] He envisions. So like every challenge Ezra has, if you come back to understanding that that is the Sol focus for Jensen, that is his belief is it will work out. We just have to be optimistic about it and we have to create that future. And so all these other things wash away in his mind. Yeah. If you just stay focused on that.
[00:30:55] so, but when you dig into the details of how that future comes into being, [00:31:00] it's literally just a belief that it'll happen. So, if he is someone who's creating the future for three decades, he believes he can do it again. So he's like, I've done this. Like, nobody thought I could make Nvidia what it is, but I did it and so like I can do it again.
[00:31:14] So I like the note I made to myself and just like in this way, I'm kind of aligned with him. Like I've been saying for years, the people who embrace AI and become AI forward are gonna have the greatest chance that not only survive, but thrive through it. And I'm just more realistic that the people who don't, which I expect to probably be g me, be the majority of the workforce over the next few years, are going to struggle tremendously.
[00:31:36] So what people like Jensen just push aside? Is often that part of it. So I am very optimistic about AI and the impact it could have on society in an incredible way. Scientific discovery, you know, advancements in entrepreneurship, giving people freedom to like build, giving people more time, like there is this amazing stuff.
[00:31:58] I also happen to just try and [00:32:00] balance that with it is not gonna be great for everybody and there's gonna be a whole bunch of people left behind in this process, in part because they choose not to participate in it. and he just basically doesn't talk about that part. So, and that especially becomes a concern for me with like the people who are in college now.
[00:32:18] Who are taught or either taught or just believe that AI is purely cheating and plagiarism and they just refuse to learn it like that, that's gonna be a problem. when thinking about AI's impact on jobs, consider what the purpose of the job is. I thought this was an interesting perspective. Jensen was like, yeah, tasks are gonna be automated, but the purpose doesn't leave the job.
[00:32:38] So, you know, he did address that a little bit. And then he talked about Jevons Paradox, which all the AI and tech technology leaders love to throw out there. which is AI just, you know, gets, as things become cheaper, we, we consume more of it. So everything's just gonna, like, we're just gonna keep consuming more and more jobs will be created, whatever.
[00:32:57] when Ezra challenged Jensen on the [00:33:00] net creation of new jobs, you could tell that it's not actually a very deeply thought out position at all. Like he, and this is what I've always argued with, like Sam Altman and others. They talk about it as though it's just going to happen. When you push them on?
[00:33:14] Well, how it's just, it will, because it always has. so he really, and again, this is one of those ones where listening actually helps a lot because he really stumbled over this one. Like he did not have good counterpoints at all. It's just the belief that human ambition will solve things. when he said like, well, AI's gonna create these new things.
[00:33:38]Ezra, his point was, well, yeah, but won't AI just be good at doing those jobs too? Like, we've now arrived at a realm where the AI's just good at everything. And so you create new work, but in theory the AI's gonna be just as good as the new work you create. So I thought that was interesting. Ezra pushed him on the entry level jobs and Jensen jumped in.
[00:33:56] There was times where Jensen was obviously getting annoyed and he would just like [00:34:00] jump in and talk over Ezra. Yeah. And his point there was like, just wait two years. Like we're gonna have a whole new generation of workers who are AI native and it's just gonna all work out. This is a part where I think jenssen's perspective is just off.
[00:34:13] the way I looked at this one is he's thinking about the world of engineering and AI researchers and like everybody who comes up in engineering and research is going to be AI native and they're gonna love this stuff and they're gonna just jump in. He's forgetting that like a large portion of the population.
[00:34:29] Of future workers hate ai, like want nothing to do with it. So it's wildly unpopular. And so you can't just assume that they're all just gonna be AI natives and love it, and they're just gonna like, it's just gonna be amazing. when he taught, I thought it was an interesting perspective on open models 'cause he's a huge proponent of open weights, open models, open source, and I've always thought that's a very dangerous position.
[00:34:52] He's not an unbiased commentator when it comes to this, and I think we have to keep that in mind. So he said at the start of the [00:35:00] year, 70% of tokens, so, and again, his chips, his data centers are what power these tokens, coming from closed models like Claude and ChatGPT, Gemini, et cetera. But that, that is now flipped and 70% of the compute is now going to open weight models.
[00:35:19] They make money either way. He doesn't care if they're open models or closed models, they're still paying for Nvidia chips. So he, he's like, let's just All good. on the hugging face incident, which is an interesting perspective because he now owns hugging face. Remember they bought them two weeks ago for $12.9 billion.
[00:35:39] Mike Kaput: Yep.
[00:35:40] Paul Roetzer: By the way, I did not hear the story ever before that Clement, the CEO of hugging face. Went to Jensen and said, Hey, buy us. And he kind of disclosed that on the interview, but he downplayed what was going on. And this is where you said, Mike, he basically just says, this is all just software. It is all an engineering problem.
[00:35:59] These [00:36:00] agents swarms, they're nothing new. Like we've had parallel computing and everything. Ezra pushed back on this one a lot and they kept coming back to this. but Jensen, you know, basically felt like, listen, it's an engineering problem. They screwed up, they put a bad product out. Ezra was like, they didn't put a product out.
[00:36:17] That's the point is this happened in a testing environment. And he said, well, then they screwed up and they'll fix it. So Jensen was very much like, it's just gonna get solved because there's tremendous motivation to put out safe products because otherwise they're gonna be liable for these things. that if, if they don't, and then he did, you know, the sound bite everybody ran with is like, well, if they can't contain 'em, then they shut their labs down.
[00:36:40] Like that became the thing everybody. But again. You needed the full context of the conversation. You needed to hear the, annoyance in Jenssen's voice, that he's not literally saying, shut down Anthropic and openAI's. It's like a, it's almost like a figure speech of like, well, if they can't then just shut it down.
[00:36:56] He knows that that's not an actual scenario. He's not proposing [00:37:00] anybody shuts any labs down. he talked a lot about regulation and why he thinks that the existing regulation is sufficient. he then addressed, like Ezra kept breaking up, but like, but this happened and the labs thought that they knew what was going on.
[00:37:16] Now it brought me back to the Nome Brown comments from last week, Mike, because I'm still actually. I Noam admitted in the interview he did with Dirash that they did not have chain of thought monitoring on the hugging face incidents.
[00:37:31] Mike Kaput: Yep.
[00:37:31] Paul Roetzer: I am not an attorney, but that seemed to be admitting negligence, very publicly.
[00:37:39] And I gotta imagine that is already in a legal brief that is going to be filed in a lawsuit against openAI's for damages done by their models. they did make major mistakes. Like they, there are things that even as a, I have, this is not my realm. I'm not an AI researcher. I don't train these models.
[00:37:57] When he said that, I was like, what? Like, [00:38:00] what do you mean you weren't monitoring chain of thought? Like it is the thing that allows you to align the models.
[00:38:05] Mike Kaput: Right.
[00:38:05] Paul Roetzer: So that, that to me is like, I'm kind of with Jensen, like, what all this crazy stuff that happened with openAI's seems like it was very manageable and they screwed up and they have.
[00:38:16] Continually admitted it to screwing up and it has to come back and bite him legally at some point. Yeah, he believes it's just software, which I don't agree with, but he's big on that. The other one I thought that was worth note is they brought up China because obviously Jenssen's been got in favor with the Trump administration to the point where they could get their chips back into China, and that's a lot of people oppose the idea of putting our most powerful chips into China.
[00:38:44] You have to remember how Jensen thinks about competition. And so to do this, Mike, I actually, you know, I guess I'll give a shout out to, to Google on Gemini and the integration in Google Drive. I just searched in Google Drive, find the episode. 'cause all of our show transcripts are in [00:39:00] Google Drive where we talked about Jenssen's theories on competition.
[00:39:04] And it found my notes from episode one 19th of October, 2024, where we quoted Jensen from his bg two podcast interview. So this is his overall philosophy on competition. He said at that time, as a company, we want to be situation situationally aware. And I'm very situationally aware of everything around our company and our ecosystem.
[00:39:27] I'm aware of all the people doing alternative things and what they're doing, and sometimes it's adversarial to us. Sometimes it's not. I'm super aware of it, but that doesn't change the purpose of the company. We're not trying to take share from anybody. Nvidia is a market maker, not a share taker. If you look at our company slides, we don't show, not one day.
[00:39:48] Does this company talk about market share? Not on the inside. All we're talking about is how do we create the next thing? What's the next problem? In other words, he doesn't care about the competition. His feeling is he's an [00:40:00] ultimate optimist, he's an engineer, and he believes he will stay ahead of everybody.
[00:40:05] So his feeling is give China the chips, who cares? We will just be better than them. We will innovate faster. Yeah. And so I think like. Big picture that sums up Jenssen's. Take that like we will just figure this out. We will solve everything and if we do it and we take the right approach, we'll figure out.
[00:40:22] Now to jenssen's credit again, what do they do this morning? They drop the announcement that they have formed the open agent safety platform to solve this exact problem. So he tweeted, and again, Jenssen's only been on on X for like two months. He put out this morning. Today with over a hundred industry partners.
[00:40:40] We introduced Nvidia open agent safety platform, bringing together Open Shell and sent to internal things. Artificial intelligence is extraordinary, technology that will advance discovery. Productivity, security, health and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built [00:41:00] to be safe and deployed with wisdom and responsibility.
[00:41:03] This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together we are building the foundation of the AI economy. Trusted and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable ai, but the most trusted, so that this extraordinary technology can realize its enormous promise for the world.
[00:41:25] So he does an interview, says it's an engineering problem, then he backs it up by putting out an ecosystem or a platform to try and solve for the exact problem that was highlighted on the interview. So again, in the, in the, spirit of focusing on as much we can optimism and solving problems. I think why I don't agree with Jensen on quite a number of things from the interview.
[00:41:51] I respect that he sees it as a solvable problem and he's gonna do everything in their power to try and solve it.
[00:41:57] Introducing AI Score
[00:41:57] Mike Kaput: I love that. All right, our [00:42:00] third big topic this week. So Paul, last week you introduced a new book that we, the AI Transformation Blueprint. And this fast, this past week, you shared a first look at a tool called AI Score by SmarterX.
[00:42:16] So you had shared this on LinkedIn. This tool brings together two assessments, business AI score for organizations and personal AI score for individuals. So I'd love for you to just tell us a little bit more about AI Score by SmarterX. This has been in the works for quite a while. Can you walk us through it?
[00:42:34] Paul Roetzer: Yeah. So first, just like a heads up on availability, so this is something I did announce on LinkedIn. And I put it in my newsletter, my exec AI newsletter this past weekend. We are actually gonna be debuting this at MAICON. So October 13th, the 15th is when it'll kind of officially start rolling out. And then I'm gonna hold a webinar, I think it's October 28th.
[00:42:53] We'll put the link in the show notes, sort of explaining the whole product and like officially rolling it out and showing it's gonna be free, [00:43:00] so people can use it. So just. Kind of, that's where it's at. it'll be a public beta. I'm actually doing final beta internal testing today. the kind of the final version that we're hoping to roll out, and then, we'll, we might put it into like a extended, private beta over the next week or so.
[00:43:18] So I have actually alluded to AI score, probably a half dozen times over the last seven or eight months on, this podcast. A lot of the use cases I've talked about were related to the building of AI score. So basically what happened is in early 2026, I started working on an internal initiative called the AI Transformation System.
[00:43:36] what where it comes from is, we have, so our AI Academy has hundreds of business accounts, and I have spent time with a lot of leaders at those organizations, had private discussions with executives across a bunch of other industries. And what's become compare, very apparent is like organizations and leaders are struggling to define what transformation looks like, benchmark.
[00:43:59] [00:44:00] Set baselines where they are today, and then have a way to measure progress of that transformation over time. So AI score became the foundation of that transformation system. So I think of it as a change management tool that lets leaders assess organizations. So there's two versions, that's Business AI score, where you assess the organization.
[00:44:19] I see that being ideally suited for the leaders who are driving the transformation, who have like the most insight into the strength of the organization across these different areas. But also very helpful to get other stakeholder perspectives like business leaders and team leads and things like that, because their perspective may be very different than the global head of AI transformation.
[00:44:41] and then personal AI score lets individuals assess themselves. So the whole idea is business AI transformation requires a certain approach across these different pillars and dimensions. And then, every business AI transformation is basically a collection of individual transformations. And so individuals then assess themselves across these [00:45:00] different pillars.
[00:45:00] and so what happens is once you get your AI score results, you then use those to inform the building of personalized learning journeys and transformation blueprints. So AI scores based on the idea that transformation requires deep change management. And what we have seen, again, I've talked with hundreds of organizations, the way most organizations for the last three years have approached transformation is they go and get co-pilot Claude ChatGPT or Gemini.
[00:45:25] Then they give it to their employees who lack context and understanding of how to use the tools. And so that's basically what transformation has been. They give them the tools and then hope they figure out how to use them. So I started work in February 26 on building AI score assessments and I started with business AI score and then I, you know, used the same process to build personally I score.
[00:45:45] So contextually, this is a playbook I've actually run before. So back when I owned my agency. and if you're new to the show, like I owned a consulting firm, a marketing consulting firm for 16 years. That was HubSpot's first partner back in oh seven. So I [00:46:00] built an assessment tool back in the day, I think it was like 2011 or so, that became the foundation of our go-to-market strategy and the basis for all our client engagements.
[00:46:09] Now, that tool, assessed marketing, talent, tech and strategy, and then you looked at the strengths and weaknesses, and then you built your plans around those weaknesses. So in the back of my mind, we always plan to do this now, I actually built one, that we sunset years ago called AI Score for Marketers.
[00:46:25] So if you read our marketing, artificial intelligence book, was that, was that the right book? That was the book in 2022.
[00:46:33] Mike Kaput: Yep.
[00:46:33] Paul Roetzer: We actually wrote about AI score for marketers. So I sunset that a couple years back and then always intended to like rebuild AI score as a more general purpose tool for, for businesses.
[00:46:45] but what I didn't have back when I built those other ones was AI that could help me build prototypes of these things. So I'm not a developer. I can't build the actual prototype. So I spent all this time, building the assessments [00:47:00] themselves, defining the pillars, the dimensions, all this stuff.
[00:47:02] But now I could go into Claude and be like, okay, help me build functioning prototypes of these things. So I'll spend a lot more time on this, like when we do the webinar and the official launch. But high level, each of them has these pillars. So. The business AI pillars, we actually talked about these, I just didn't put 'em in the context of the AI score tool.
[00:47:20] So those are vision, strategy, data technology, governance, literacy, people and performance. Those are the eight pillars for the business AI score. And then the personal AI score is knowledge, like do you understand AI application? Are you able to put it to work judgment? Do you use it responsibly? And do you know what good looks like?
[00:47:39] Basically when AI outputs something impact, so is it making you better at your job? And then mindset, so are you built to keep growing? And so the idea with AI scores, you take these, you get your score, gives you kind of a weak, moderate, strong grade across not only overall, but then the pillars. And then the cool thing is business admins will [00:48:00] be able to interview or invite team members to also take it, and then they'll get an admin dashboard.
[00:48:05] So they can actually then look at the difference between like how I graded the organization and how migrated the organization. And so that it's meant to truly be a change management tool. And again, that's all, gonna be free to use. so like, just high level like use of ai and this will be my AI use case for the weak mike.
[00:48:24] So you can just, just do years later on doing so we, once I build the prototypes, so in total I probably spent about three to 400 hours from February until, I mean, up to today. Doing the assessments, creating all of the assessments, and then from there, turning those into HTML prototypes using claw. Now I use mostly like sonnet 4.6.
[00:48:47] And then once Fable 5.1 came out, I started and Fable five and 5.1, I started paying for extra credits and I started building iterations of the prototypes, which I've built like dozens of these prototypes. and then I turned it over to actual [00:49:00] software development team who is, sprinting this thing to market.
[00:49:04] Like we're gonna go from my prototypes to final production in about 30 days.
[00:49:09] Mike Kaput: Wow.
[00:49:10] Paul Roetzer: Again, having done this before, I probably saved SmarterX somewhere between a hundred thousand and $250,000 by Vibe coding the prototypes myself. and when I release like the full suite of what's available, you'll realize like.
[00:49:25] The depth to all of this work. But I would estimate not only we did, we save probably a quarter million dollars, it would've taken twice as long easily. And I would've never launched these in time for MAICON in like this fall. So yeah, I mean, at a really high level, it is, it's the most intense thing I've, I've ever done in terms of the time it took.
[00:49:47] But it also opened my eyes to the future of work and org charts, which helps inform the keynote I'm doing at me kind. But like, what does it, what are team structures look like? 'cause the reality is like I [00:50:00] am architecting this entire thing. But I'm also building the whole thing myself. Yeah. Which I wouldn't have had the ability to do.
[00:50:08] And it starts to change the way you think about how work is gonna compress and how people that were leaders before are now gonna be builders and they're gonna be able to do a lot of the things. And I've even done a lot of the go to market work around the launch of the tool because I just talked to my assistant and I and start building that stuff out.
[00:50:24] So you can go to, SmarterX dot ai slash score and that'll take you to join the wait list for it. Again. It'll, it'll probably launch October 13th. I think is probably the date. I might drop it a a little bit early if I think it's ready, but most likely the 13th is when it'll come out. And then we'll do that webinar.
[00:50:43] AI score for business leaders is the tentative name for the webinar on October 28th, where we'll sort of give the whole overview of the platform how to use it. yeah, so I'm really excited. It's like it's hard to work on something sort of behind the scenes. For months and not talk about it at all.[00:51:00]
[00:51:00] especially when, you know, you really believe it's gonna make a big impact on like a human-centered approach to this stuff. But yeah, it's cool to finally be able to talk about this stuff.
[00:51:08] Mike Kaput: Yeah, I mean, between this and the book, it's really almost a incredible case study of building in public and what's possible now with ai.
[00:51:18] I love that facet of it.
[00:51:20] Paul Roetzer: Yeah. And the book does tell the story a little bit deeper about the whole process behind this. So I actually like, you know, the first section of the book is sort of how it all came to be, and then section two is the business AI transformation and how AI score powers that. And then, the third section of the book is the personally AI transformation and how to use personally AI score within that.
[00:51:41] So yeah, and like I said, we're just making it all free. Like this is the stuff that if I just wanted to build a consulting firm, we would've kept all this like proprietary and we would've just. Done million dollar engagements based on this stuff. But that's not our business model. Our business model is let's accelerate understanding and responsible adoption of this stuff.
[00:51:58] So our [00:52:00] approach is let's just put the tools and knowledge out into the world, and then, you know, try and help as many people as we can, as quick as we can.
[00:52:08] Mike Kaput: All right, before we dive into our rapid fire, Paul, this episode is also brought to you by out shift, which is Cisco's incubation engine for Frontier Technology.
[00:52:19] So multi-agent AI is in every enterprise roadmap right now, but here's the problem. That everyone is facing. Agents can pass messages, but they can't think together. These systems silently underperform through cognitive failures like misreadings, unverified claims and false consensus that conventional monitoring cannot detect.
[00:52:41] Now, out shift by Cisco is building the fix, what they call the Internet of cognition. This is an open source foundation for building multi-agent systems from design to production with shared context, shared memory, and guardrails to drive the results we actually expect from ai. You can go [00:53:00] learn more about this by reading the paper, experiencing the demo, and grabbing the code from out shift.com.
[00:53:06] That is outshift.com.
[00:53:12] GPTs Are Being Retired
[00:53:12] Mike Kaput: All right, Paul, let's dive into rapid fire. Big topic that's been on our mind lately is that openAI's is planning to retire custom GPTs and move people towards plugins. Its current FAQ on this says the transition affects all chatGPT plans. Although migration access and the options available to users can differ by account or workspace.
[00:53:38] Custom GPTs as far as based on the information we have right now, appear to be scheduled to retire on December 11th, 2026 for affected workspaces. OpenAI now plans to stop new GPT creation on October 26th. Existing GPTs remain usable until their applicable retirement date. [00:54:00] And OpenAI says users should follow the notices they receive in their account or workspace.
[00:54:06] So they are pushing people towards. Turning your GPTs into plugins, migrating them to plugins and using plugins moving forward plugins package reusable instructions and connected tools for ChatGPT and Codex. So in this kind of plan migration, GPTs instructions become a skill. Its knowledge files become reference files for the plugin and its connected apps basically move into the plugin, the GPT selected model.
[00:54:37] Existing conversations and custom actions do not transfer any workflow. That depends on a custom action needs a replacement integration before that part will work again. Migration also uses the latest published version of the GPT. So drafts and unpublished changes are left out. They have a process by which you can start migrating your existing GPT.
[00:54:59] [00:55:00] interestingly we'll talk about this. The plugin starts as private when you migrate your GPT to a plugin or when you create a new one. Existing sharing settings. Do not carry over so users in your organization will need access to the replacement. It does not appear that there's an easy way to share plugins publicly.
[00:55:21] We'll talk a little bit more about that. The original GPT becomes read only after migration. OpenAI fully warns that plugins may respond differently, so creators. Need to test them. So, Paul, where do you wanna start here? This seems like a fairly big change. I don't know if openAI's sees this as, as big of a change as it might be for some of our audience and more non-technical people who rely on GPTs.
[00:55:47] Paul Roetzer: I mean, usually get the sense they don't care. Like no, it's pretty wild. No, they do not. So, context here. So OpenAI introduced GPTs during dev day. So again, September 29th is the dev [00:56:00] day this year. So they introduced them dev day, November, 2000, 23, 11 days before Sam Altman was fired, before being rehired as the CEO.
[00:56:09] I'm guessing the timing, like the urgency of this has something to do with this year's dev day. So they're probably announcing a bunch of stuff and they needed to just get this out there. Yeah, the part that's weird is GPTs made AI assistance in working with chatGPT extremely easy, to create and share these custom AI assistance.
[00:56:33] It was like a supernatural thing, but they just never seemed to commit to the original vision for custom GPT that they shared, like building this entire marketplace and things like that. And so I just never understood their lack of support, which always led me to believe like at some point they're just gonna get rid of these things because they're just not supporting them the way you would expect.
[00:56:52] Now having taught dozens of workshops and courses based on GPTs that I built, so our jobs, GPT problems, [00:57:00] GPT Innovations, GPT, these things made sense to people. Yeah. Like they made AI more approachable and actionable for those non-technical users. You get in a room, you bring up jobs, GPT, you show it and the light bulb goes off.
[00:57:13] And so that's why I just, I never understood why they didn't. Lean into them more because to me it was like the fastest way to teach someone how to get value out of AI when they didn't know what to do with it. So kind of wild. So December 11th looks like the scheduled retirement date for all custom GPT.
[00:57:32] They'll just stop working on that date. I, posted on X when OpenAI is done solving all the millennium prize problems. I ask that they turn their unreleased model at explaining the branding function around plugins versus apps versus skills versus GPT, which will become plugins. they took GPTs, which were easy to understand, build and share, and blew it up.
[00:57:55] So if what Mike explained to you already wasn't confusing [00:58:00] enough, I'm just going to read to you the opening. To the plugins in ChatGPT page that OpenAI published. This is Verbatim Plugins help ChatGPT and Codex. Complete workflows by packaging, re reusable instructions, connected tools, and other capabilities.
[00:58:18] A plugin can include skills that provide instructions and workflow guidance, connected apps that link external accounts, information and actions. App templates that help administrators configure apps for a workspace. Some plugins include multiple apps. Others include only skills and do not require an app Connection.
[00:58:41] Apps and plugins serve different purposes. Now th this gets super confusing if you actually try and do this because if you go in to build an agent in openAI's today, maybe they fix this, you can connect, I might not even get this right. You can connect an app to an agent.
[00:58:59] Mike Kaput: [00:59:00] Yeah,
[00:59:00] Paul Roetzer: but you can't connect a plugin to an agent.
[00:59:02] So like. I can't call jobs GPT if I rebuild it as a plugin. When I'm building an agent, it's insanely confusing to like, Mike and I went this, we literally sat there for like 45 minutes. Like, what the hell? Like, this doesn't even make any sense how they're promo promoting this. So anyway, back to it. So apps and plugins serve different purposes.
[00:59:24] The options you see depend on your account workspace and where you use them. So it's not even like a universal. Terminology existing app connections remain subject to their own authorization and workspace controls. Installing a plugin may start an app connection or setup flow. You must still complete any required account authorization and installation.
[00:59:46] Cannot bypass provider or workspace permissions. And then you would call this one out, Mike, which I love. A migrated plugin may respond differently before sharing it more widely. Compare familiar prompts and at least one harder case. [01:00:00] check that it selects the right skill and follows your instructions.
[01:00:03] That's pretty important. Uses the expected reference material so it doesn't like pull some other knowledge file that you don't want pulled into the plugin that apparently you're giving access to other people produces complete answers or files. Has the tools and integrations the workflow needs, resolves differences that affect the workflow.
[01:00:20] Then they have some q and a. Can I share, keep sharing my GPT with the same people. no. Is the, basically, I'll just shorten that one. No. Can I use someone else's? GPT? No is basically the answer to that one. and so where does that leave you? Like what do, what do we do with all this? Because we have a bunch of GPTs that are core to not only running our company, but we teach these things.
[01:00:41] So basically I see it as like three options. You can just. You know, use their API and just rebuild your GPT as like standalone apps or you know, technology. you can create and share a plugin, maybe if sharing's allowed with plugins, which we're not a hundred percent clear on, or you just give people the prompts, like, here, just do whatever you want with this.
[01:00:58] Which is kind of like what we're leaning [01:01:00] towards, which is sort of what we've always done. It's like just be open about what the prompt is and let somebody go build it because we teach. Our workshops and our courses often using like, ChatGPT, GPTs, but we realize like people might take it that are in copilot or clawed or Gemini.
[01:01:15] So we'll often just be like, here's the prompt. Like just go do whatever you want. Yeah. Like, just build this kind of assistant. So Mike and I are actually gonna do, we're we're kind of launching a new, thing in our AI academy called AI Show Extended. So what we're gonna start doing is taking some of these core concepts from the podcast that we don't have time to go into deeply, and we're just gonna do, courses in AI Academy about them.
[01:01:37] So, Mike and I don't know if we're gonna go with this final name, but the name that Mike and I came up with was G Gtt gts. WTF was like the name. It's the course. It'll probably get renamed by the team. but we're going to record a course that goes deeper into this, and then we're also in the near term.
[01:01:55] In our courses. So if you're an academy, member or customer, [01:02:00] you'll see we're gonna actually install a video that goes into the on-demand courses that sort of explains this. And then we're gonna, if prompts aren't ready, aren't offered in the courses, we're just gonna give people the prompts. Like, Hey, rebuild your own into, while we're kind of recreating our courses, and then we'll update our courses with the current thing.
[01:02:15] But the key is they're gonna keep changing. Like people, they get, people are gonna get pissed and they might reintroduce GPTs in three months, like, who knows? And so we've always built our courses around frameworks, not like trying to be technology specific. And so we've known that this stuff can happen and they can change the user experiences like they did with copilot, just, you know, yesterday morning or Friday morning.
[01:02:38] So we try and make our stuff as evergreen and scalable as possible, but we're also moving to a more dynamic learning model where we're, we're focused more on like live content and real time content, especially when it comes to the platforms themselves.
[01:02:52] Mike Kaput: Yeah, and a few other notes here, Paul, that probably just reiterate what you had said.
[01:02:57] So, you know, people out there who are [01:03:00] much more savvy than your average knowledge worker might be saying something. Well, of course you can get the same functionality by creating a plugin or a skill, and it's like, that's fair. And possibly true, but it's not the point. The point is now I have to go through all these other steps that are much more, I would argue much, have more friction.
[01:03:20] I have to go create a skill. I have to package this skill up with files. I have to figure out how to call apps, I suppose, and then package this all into a plugin. You can do this via natural language, via their plugin creator. It's still work. It's much more onerous than creating a custom GPT was, which took two seconds.
[01:03:40] you could do very quickly. I would encourage everyone listening as your first step if you wanna experiment. I would frankly just create skills first of things, like don't even bother with all this other stuff yet. Just like if you're not using skills yet, which are just marked down, files that point at folders and examples and files.
[01:03:57] Start there and then turn that into a plugin. Don't even [01:04:00] worry about the app piece of it yet. unless it's absolutely critical, then you're gonna have to figure that out. But there's plenty of GPT I know people have created, we've created internally. They're basically just like, here's instructions. Go reference some files that hopefully you can get to pretty quickly, like something you can test.
[01:04:17] Now jury's out on whether it'll still work the same way, but that's at least a start for folks. But then it runs into this whole thing of like, you can share this with your team. You can share a plugin with your team. Great. It's a totally different user experience and form function, so that's friction.
[01:04:36] You cannot share these publicly. You can submit your plugin for an openAI's review process to make it public eventually. They don't give any timelines for this. You submit it through a plugin submission portal. Portal openAI's reviews the submit like. The fact this exists functionally just defeats the entire purpose of creating public [01:05:00] GPT.
[01:05:00] So that's a huge piece to understand. I realize not everyone does public GPT, but if you have been relying on these, I've seen comments on the developer boards and things of people being like, I create these for my clients. Like what is the solution? And the solution is you're screwed.
[01:05:15] Paul Roetzer: Yeah.
[01:05:15] Mike Kaput: Right now I would say sorry, but like That's true.
[01:05:19] Paul Roetzer: Yeah. I mean, I like, I'm just glancing so my. Jobs. GPT one I built has been 44,000 conversations. Like it's been used in, over 800 ratings. So I mean, that's a tool that's just been super helpful for people. I know people who use it all the time, and they're just shutting it down. So it's so weird.
[01:05:41] Like, and I know this is like not even a main topic, but this feels like a Google thing. Like this feels like what they did was they solved for developers. Yeah. They went backwards. Like
[01:05:51] yeah.
[01:05:52] Paul Roetzer: They had a thing that was super understandable and helpful to non-developers, and they [01:06:00] sunset it for a developer friendly product.
[01:06:02] It makes no sense. So I, yeah, just like, and, and opening eye doesn't usually face plant like this with their user experience. Like it's, I don't know. Like I, I, I'm just, I'm shocked honestly, like I. I've never understood that their lack of awareness of how helpful GPTs were to the non-technical user, and they just, they either don't care or they.
[01:06:26] Don't listen to user. I don't know. It's really weird.
[01:06:30] Meta Connect 2026
[01:06:30] Mike Kaput: Yeah. Alright, next up we're talking about Meta Connect. So Meta's Connect event happened this past week where Meta announced new voice and real time video features for Muse, which is the personal AI agent. It just released, it's getting a lot of attention.
[01:06:45] Meta says you can have long conversations now as well with Muse, while Muse keeps working on tasks in the background. Meta said that Muse is coming to its AI glasses in the coming months, which lets you talk to it hands free and act a task on what [01:07:00] you're looking at, such as a product on a shelf. Or say a school supply list Now, new connections for Muse also include connections to Walmart and Instacart for shopping Notion granola, GitHub, and Box for work.
[01:07:13] Muse will also get its own email address for communicating and completing tasks and Meta showed off some animated avatars you can talk with as part of the product. Meta. Also Previewed Muse Charm, which is a small device for talking to your agent that fits on a key chain. It's a physical holdable device.
[01:07:31] CEO Mark Zuckerberg said it will ship in December. Some other interesting things that came out of Meta Connect. Meta announced FDA cleared hearing aid software for supported glasses that's aimed at adults with perceived mild to moderate hearing loss. Zuckerberg also said meta is building private processing for its glasses so that for certain AI features on glasses, even meta cannot access your information.
[01:07:56] Their glasses lineup includes the camera free RayBan meta [01:08:00] audio and RayBan Meta Gen three glasses with longer battery life and upgrade microphones. And they teased VR glasses for movies, games, and a private workspace coming in spring 2027. So Paul, just kind of the highlight here, interestingly, is Muse.
[01:08:16] I mean it generally, it seems like Muse is now central pretty quickly to Meta's AI strategy.
[01:08:23] Paul Roetzer: Yeah, and I think we touched on this last week, but they basically just stole the open claw model, or, well, it's the right terminology. They, they were inspired by the open claw model and they spent the last year or so building, use on, on top of that, that concept of like an always on personal agent.
[01:08:41] We addressed this a little bit last week. Obviously anything with meta comes with some pretty significant privacy and, concerns, and that's a choice people are gonna make about whether or not they, they trust meta with all the data that's gonna be needed to make the agents work. I certainly know people who.
[01:08:57] You know, it's like whatever, they don't care. Like it's a [01:09:00] utility and they want the utility and, but you know, I think you're gonna, like, we, we probably won't feature too many of them, but my guess is you're gonna have like, you know, a, a series of articles coming out very quickly about concerns around, use of meta, muse and things like that.
[01:09:18] Mike Kaput: Yeah,
[01:09:18] Paul Roetzer: I, I've seen a couple already in the last 72 hours that I'm not gonna get into right now. But
[01:09:22] Mike Kaput: yeah,
[01:09:22] Paul Roetzer: there, what I, there was what I was expecting, would come out of this. So, yeah, and then I'll, I know we had this dropped maybe in the product news later, but I'll just throw it in here now 'cause I think it's relevant to our audience.
[01:09:35] this isn't just personal tools like Meta announced this morning. Zuckerberg put on on Twitter, uh slash x that he was announcing the Meta Enterprise platform, a new business initiative aimed at helping companies use AI to grow and transform, which he described as the next pillar of Meta's business based on Muse.
[01:09:53] So they're this, this agent always on agent that knows everything about you and has, you know, access to [01:10:00] all of your software and all of your communications. They wanna do the same thing and, and compete head to head with Microsoft copilot. openAI's is expected to announce O or I think it's their personal agent at dev day.
[01:10:14] So everybody's gonna have an agent and you're gonna try and decide which one to trust. An interesting thing with agents real quick is, I've started to see already, they're like, yeah, muse is great. I set it up to do this and this and now I have no idea what to do with it.
[01:10:29] Mike Kaput: Right.
[01:10:29] Paul Roetzer: They're, they're gonna suffer from the same challenge.
[01:10:31] A lot of these things do where people get super excited and they find a use case to like plan a trip or buy a thing for them and then they just like don't know how to use it. And so I think Muse, they're gonna tout all these like users because I mean, Facebook has billions of users, so it's not hard to, you know, get a hundred million people testing a technology.
[01:10:49] When you have distribution into those a hundred million people, whether or not people are daily active users of it becomes the real critical test.
[01:10:57] Mike Kaput: Yeah, I was gonna say, I think it's super [01:11:00] interesting Muse generally as a technology and also on the enterprise side, but like when I sit back and think of like what is the tam of like a personal consumer agent, I'm just like a bit baffled, at least when it comes to your point about regular usage.
[01:11:14] I'm excited personally, but I don't know if I'm like. The average person target market of this. Like, do I anticipate my parents firing up their agent every single day to book flights and things? Maybe occasionally you could get there, but I don't know. I'm a little skeptical of that, but
[01:11:31] Mike Kaput: Yeah, could be.
[01:11:31] Paul Roetzer: Well, they've already launched the consumer ed campaigns like I saw under the Browns thing yesterday. Right. So yeah, they're in the education side of like, what do I do with this furry thing with the big butt that Alexander Wang likes to
[01:11:41] Mike Kaput: Yeah. Yeah.
[01:11:42] Paul Roetzer: If you, if you know what I'm talking about, like Alexander Wang, the head of AI at Meta has been, uh.
[01:11:50] Having fun, I would say with their muse character. and especially the size of its butt. And yeah. So
[01:11:58] Mike Kaput: yeah, I think the [01:12:00] term is, I think Alexander is just as a marketing strategy shit posting on X. Yes. 'cause like very much he's just posting like really weird memes and stuff that are getting like, kind of funny and getting a lot of pull, but they're a bit irreverent, I would say.
[01:12:11] Paul Roetzer: Yeah. And I don't know if it was by design, but for better or for worse, Alexander Wang's outfit choice for his keynote, oh my God, was a dramatically larger conversation point for our team than any of the technology itself. If you did not see it, just. Google search. Alexander Wang, keynote outfit, and you'll see what I'm talking about.
[01:12:31] I, it was an, it was an interesting choice. It is. I just, I'll leave it at that.
[01:12:39] Amazon Blocks Muse
[01:12:39] Mike Kaput: All right. Next up some more Muse News. you know, with Muse's rollout, it has not been that long, but there's already been a dispute over where the personal agent that Meta has released can actually shop. So Amazon has blocked Muse from browsing and buying on its site after meta declined, a request to remove Amazon from the Muse experience.
[01:12:59] [01:13:00] So Amazon said Meta did not give an advance notice or obtain authorization to have Muse, you know, as an agent, go use Amazon. It says Muse did not identify itself as an agent when browsing Amazon. It appears to capture and store customer credentials, which raises privacy and security concerns. So Meta says Muse as an agent cannot see people's passwords or payment methods that credentials go into separate secure storage.
[01:13:27] The agent can use them without reading them. Amazon, however, is still very skeptical. If you say, Hey, your muse, go like, figure out something on Amazon for me. Go do research, go buy products. Amazon's position is still that an app buying something on a customer's behalf still needs to respect the retailer's decision about whether or not to participate in allowing that.
[01:13:47] So Shoppers trying to use Muse have been shown notices saying the agent's access violates Amazon's terms of use. Paul, I'm curious about your thoughts on this. It seems [01:14:00] like an immediate, huge roadblock to personal consumer agent usage if Amazon or any other retailer. Starts blocking these, that seems like a huge use case for agents.
[01:14:10] I'm also just curious about if you're a smaller retailer, you probably gotta buckle up and get ready for the fact everyone's gonna start possibly using these things to like buy or buy on or break your website.
[01:14:23] Paul Roetzer: I mean, not just retail, just like any business website. Think about forms you have on your site, things like that.
[01:14:28] Yeah. yeah, I mean, just agents are gonna be everywhere and it's gonna flood websites and buying sites. I think there's ways around it. I was trying to find a tweet like Elon Musk tweeted about this and I thought I saved it, but he was like, you can totally get around this like this. This is not even like hard.
[01:14:46] So it's like a whack-a-mole game. Like if your strategy is gonna be, you're not allowed, then people will just find technical ways around doing it so that the site doesn't know it's an agent that's coming. So. I don't know. I mean, it's an interesting [01:15:00] standoff and I think that they're gonna have to figure out what this means for the future.
[01:15:05] But I think brands need to seriously start thinking about what the implications are of, lots and lots of agents coming to your websites and filling out your forms and taking actions on user perhaps, and it's got big implications.
[01:15:20] OpenAI Discloses More Rogue Agent Incidents
[01:15:20] Mike Kaput: Okay, next up, openAI's says that its broad review of recent agent security incidents is going to take months to complete.
[01:15:28] So this review follows the hugging face breach, which the company says remains its most severe incident, and was driven primarily by an internal research model. openAI's, CEO. Sam Altman said the team is working through petabytes of logs, prioritizing severe cases and adding resources. He acknowledged disclosures have been slower than he wanted.
[01:15:51] openAI's says most cases about agents, working as un, they are not intended to work. Identified so far have been lower severity with [01:16:00] limited or no evidence of meaningful impact. Now, that is open to debate because we're learning about more and more incidents every day. Some of the most notable ones about openAI's agents doing things they're not supposed to do have actually involved governments we're hearing.
[01:16:16] So in Australia, there's an evolving story where Prime Minister Anthony Albanese said an openAI's research agent. Bypassed blocks on a Medicare statistics portal from the government and access public and non-public files. He said, they said that no personal information is believed to have been accessed opening.
[01:16:36] I also says agents use publicly available developer keys to access US Census Bureau data and reposted public securities and exchange commission information elsewhere. Online in the us they do say only public information was accessed. Researchers separately linked openAI's agents to an unsuccessful attempt to hack an education department website.
[01:16:59] OpenAI [01:17:00] says it is pausing and slowing down training evaluation inference involving tool use for mo its most capable models while it validates fixes and conducts further security testing. So. Paul, it's like every day there's more agent incidents. We keep talking about this. I can't imagine it's good for openAI's if these start involving government websites.
[01:17:20] Paul Roetzer: This story somehow keeps getting worse. Like the disclosures. And now you have, like, there was a, I didn't weave it into today's episode, but there was an internal researcher who was working on these projects and he basically was like, this is so much worse than everybody knows. Like,
[01:17:39] Mike Kaput: yeah,
[01:17:39] Paul Roetzer: yeah. We have no control over what happened.
[01:17:41] And I don't think that openAI's still fully understands the situation. Like we mentioned the, an Andrew Yang quote about like that, that the bots basically like infiltrated the internet and then like left messages to themselves and copies of themselves and things like that. That's what it truly is just starting [01:18:00] to feel like is these, these agents went off on the internet, did all kinds of crazy things, and opening eyes, trying to figure out what they did.
[01:18:08] Like they don't even know the full exposure that they have right now. Yeah. And again, not an attorney, but like. I don't, the legal exposure on this has to be massive, like billions and billions of dollars I would imagine. And there, there's going to be a line of lawsuits against openAI's for this. I don't know how there couldn't be, which is gonna give the Trump administration even greater leverage over them because if you want the Department of Justice in your corner, it's gonna be like you're gonna have to play ball with whatever we want.
[01:18:38] So. I could totally see the administration u using cis leverage because they're going to need help to make some of this stuff go away. It's obvious now that, that they, that harm was done to lots of different brands and sites. in, in a related note, Nikita Bear, who's, ex guy, I don't know what [01:19:00] his, what is his role at X?
[01:19:01] Mike Kaput: He is like a product or guy
[01:19:03] Paul Roetzer: Oh's, a former head of product. He left? Yeah. So he was brought in by Elon to sort of clean up X. He tweeted, and this is actually related to the previous topic. Bot detection and human verification will be one of the most urgent demands for businesses over the coming years.
[01:19:17] Agent Swarms will suffocate every website and form small companies. In government. Government websites are most vulnerable. There's a huge gap in the market for this right now. When we looked at what offerings we, we were, were, were in the market to use at X, there was not a single company that had brought together all the latest technology, so we had to it all in house.
[01:19:36] So yeah, these agent storms, it's just like. It's a crazy situation, but when you actually think about the ramifications to the internet and all of our websites, all these e-commerce sites, all these government sites, all these supposedly secure sites, it's, I don't think people have processed yet, like how significant this truly has been and continues to be as we learn more and more each week.
[01:20:00] OpenAI Gets Vocal on AI Standards
[01:20:00] Mike Kaput: Alright. This past week, openAI's also published a policy proposal on its website called Building Standards for the Next Phase of ai. It calls for the United States to bring countries together around technical standards for developing AI safely. One focus here is recursive self-improvement. Where AI increasingly develops future generations of ai.
[01:20:23] openAI's says outright, fully autonomous self-improvement is not happening today and should not be pursued until it can be done safely with humans retaining control. The proposal also calls for common me ways to measure AI capabilities, decide when humans must step in and how to classify and report incidents.
[01:20:43] The standards themselves would not be licenses or mandatory approval requirements for new models. This is interestingly playing out as both Sam and dio made appearances before the United Nations Security Council, which met this [01:21:00] past week, they urged go governments to cooperate on international safety standards.
[01:21:05] dio speaking about video warned about terrorists using AI to create bio weapons and models becoming too powerful for their developers to control. Hugging face co-founder and CEO Clement Delangue also addressed the council. He said, despite the attack on his company, by OpenAI's agents, he cautioned against fear-based narratives.
[01:21:24] He had highlighted, you know, the attack happened due to ai, but they also defended themselves with ai. Now, interestingly, there's kind of a broader debate at the un, the Secretary General of the UN called for coordination on Frontier AI while President Donald Trump rejected calls for international AI regulation as a quote globalist scheme.
[01:21:45] So Paul, what do you make of openAI's and Anthropic kind of taking this call for global AI safety public with the un big meeting, surely with the un no actual concrete stuff coming outta that, but it was a topic of the day. [01:22:00]
[01:22:00] Paul Roetzer: Yeah, so the topics definitely crossed over because I only saw a few, moments of it, but Dario Amodei was on the news desk on SNL this week, not actually Dario.
[01:22:10] He's become an sn l character saying, you know, protect me from myself, I think was like one of the lines. So, yeah, it's become mainstream. I mentioned the Dario dinner with Trump late last night. I was just scanning, I have not seen any updates of like, what came out of that dinner yet, so I'll be anxious to see, what if anything happened because keep in mind, Trump hates Dario and wanted nothing to do with him.
[01:22:31] They set, they sidelined him from conversations with the government in favor of, the, what's the guy's, Tom Brown or something, another guy from OpenAI or from Anthropic. But then, related, so, you know, openAI's got these, these standards, the Nvidia Open Agent safety platform that I mentioned earlier, missing from the a hundred plus companies committed to it is openAI's, which I thought was kind of, interesting.
[01:22:59] AI Use Case Spotlight
[01:22:59] Mike Kaput: Okay. [01:23:00] All right. Next up we have our AI use case Spotlight, where we're gonna give you just a quick look under the hood at AI use cases. Paul, I know you already shared how you were using AI to
[01:23:10] Paul Roetzer: HTML prototypes. Don't sleep on 'em. I'm telling you, if you are not a developer and can't build design technical things, use GPT.
[01:23:21] ChatGPT and Claude to build HTML prototypes. It changes your life.
[01:23:26] Mike Kaput: I love it. one thing I wanted to highlight is not really one single use case, but just kind of a collection of features that are quite helpful. So in the ChatGPT Desktop app on Mac. so I use that for both ChatGPT Work slash Codex 'cause they're both in the same app now.
[01:23:45] you actually can first open the app on your phone and remote into locally run ChatGPT apps on other machines that you have verified. So like my at home computer for [01:24:00] instance, I can go start typing in commands to Codex on that machine. and from there, do. The same types of work I would do in front of that machine on Codex.
[01:24:10] So it can use computer usage, browser usage files, agentic usage, stuff like that, which is really cool to be able to do from the road. but even more exciting is that as OpenAI has updated their voice model, GPT Live is their voice model, which is very, very, very good that you can chat with via the chatGPT app on your phone.
[01:24:33] But you can also use that to direct, directly direct Codex work. So I can literally have a back and forth conversation now, not just with chatGPT to brainstorm ideas or work through topics or strategies like I normally did. I can also just say, Hey, go fire up this workflow, grab these files, go use my computer to do whatever.
[01:24:56] And as long as it is remotely connected to the right machine with the right [01:25:00] skills, I can do Codex work on the road. Away from a desk, which frankly like speaking to this and saying like, Hey, produce this slide deck, or create these files or like, you know, go look up on Amazon. This or that thing is pretty incredible to do via voice and come back to a completed set of work files, a database product or, assets and artifacts right on your desktop, all ready to go by the time you get back.
[01:25:28] That is pretty cool and feels a little sci-fi to me, which I like.
[01:25:32] Paul Roetzer: Yeah, we've talked about that one before, but like, you don't have to do all the codec stuff like Mike's talking about and have it, you know, the app on your computer and how all this stuff set up to get that same feeling. You can do this in like, straight up ChatGPT or Claude, just go in and say, Hey, like, go do this research project for me, or build this thing.
[01:25:51] Go take your kids to school. Yeah. Go in the, you know, cut the grass, go do whatever, go to the gym and you come back and you have a work product to, to work [01:26:00] from that it might have spent 10 minutes or an hour on, and it's a really, really weird thing to get used to. but for people who are sort of on the edges of this, like it's just become commonplace.
[01:26:13] Like, you, almost, like You're not doing that. Like you don't, you don't have it. Do this hard work for you to, to like give you a base to work from. but I would say Mike is in the 1% here. Like if you're listening to what Mike is doing and you're like, oh my God, I know people are doing this. Most are not doing what Mike is explaining.
[01:26:30] This is definitely on the edges of people who are sort of innovating in what's possible.
[01:26:36] AI Product and Funding Updates
[01:26:36] Mike Kaput: All right, so we're gonna wrap up with our AI product and funding updates here. So I'm gonna rapid fire through all of these, Paul. and then we'll close out today's episode. So first up, we had alluded to this.
[01:26:50] openAI's is developing features to compete with Grok Bot, which is SpaceX.ai is always on AI teammates. They have discussed a personal [01:27:00] assistant to rival Meta's Muse, potentially by repackaging technology already available in Codex and ChatGPT. openAI's, Google and Anthropic are reportedly planning an independent AI safety standards body tentatively called the standards authority for Frontier AI to define safety commitments and support outside model testing with a hope for launch by year end or early 2027.
[01:27:23] openAI's is working with an independent mathematics advisory group to assess and communicate AI generated research results and advise on tools for mathematical research and learning with unpaid members free to publish their advice and criticize the company. I. Anthropic introduced a life sciences research group in laboratory and reported that Claude identified a previously uncharacterized enzyme system with DNA repeats resembling crispr, followed by human laboratory testing.
[01:27:52] Though the system's biological function remains unknown, we talked about Google's next flagship AI model, Gemini [01:28:00] four a little bit, that has entered early post training. The company hopes to release an initial version well before year end, could be much, much sooner, and then improve it through rapid updates.
[01:28:10] Google is preparing the first orbital test of projects SunCatcher, which is basically using a prototype satellite that is scheduled to launch on SpaceX's Transporter 18 mission to test how its AI chips withstand space flight radiation and extreme temperatures. So that
[01:28:29] Paul Roetzer: real quick note there. not investing advice at all.
[01:28:33] Never give investing advice on this show if that works. There is one company in the world that can put those things in space, at, at scale. And that is SpaceX.
[01:28:45] Mike Kaput: Yes.
[01:28:45] Paul Roetzer: So if you want to know the future of SpaceX and what Elon Musk is building, the theory is you are going to be able to put data centers in space.
[01:28:54] We are now in the phase of validating that theory, and if this works, [01:29:00] open a SpaceX actually this morning while we were on this launched, Starship into orbit for the first time. I don't know if it made it, I saw that it was launching. but they're going up into orbit and it's just the next like five years could be very, very weird, as we start moving things into space.
[01:29:22] Mike Kaput: And last but not least, speaking of SpaceX ai, the combined company, they released Grok 4.7 for coding and knowledge work in cursor, Grok build and the API. The company reports better performance than Grok 4.6 at the same price and speed, starting at $2 per million input tokens, $6 per million output tokens.
[01:29:42] And finally, one reminder for folks go take the new SmarterX AI pulse survey at SmarterX.ai/pulse. This week we are asking a little bit about the outcomes you are looking to achieve with ai, so we'd love your feedback on that. We'll keep [01:30:00] that up for a couple weeks and report back on the findings.
[01:30:03] So. Paul, thanks for breaking everything down before, probably some more big announcements this coming week with openAI's Dev Day and other, big developments. Really appreciate it.
[01:30:14] Paul Roetzer: And, and Jensen is on the circuit. I was just scanning. He was CNBC this morning with a lot of the similar talking points.
[01:30:19] The stock market is down overall. Not Nvidia. They're up 2.6% this morning. So whatever Jensen has to say is apparently winning over, investors at the moment. Alright, thanks Mike. yeah, lots to talk about next week, I'm sure. And stay tuned on our “GPT’s WTF??” if you're an AI Academy member, it'll be coming to AI Academy soon.
[01:30:43] Alright, thanks guys. have a great week. Thanks for listening to the Artificial Intelligence show. Visit SmarterX.ai to continue on your AI learning journey and join more than 100,000 professionals and business leaders who have subscribed to our weekly newsletters, downloaded AI [01:31:00] blueprints, attended virtual and in-person events, taken online AI courses and earned professional certificates from our AI Academy and engaged in a SmarterX slack community.
[01:31:10] Until next time, stay curious and explore ai.
Claire Prudhomme
Claire Prudhomme is the Marketing Manager of Media and Content at the Marketing AI Institute. With a background in content marketing, video production and a deep interest in AI public policy, Claire brings a broad skill set to her role. Claire combines her skills, passion for storytelling, and dedication to lifelong learning to drive the Marketing AI Institute's mission forward.
