Episode notes
Z.ai held GLM-5.3's weights for two weeks of "safety" and published a two-thousand-bug ledger the same day. Alex and Jordan audit Gemini 3.7 Flash's three-week workhorse and the key maze, DeepSeek's open harness around a metered soul, Cerebras selling a stopwatch as a working day of human knowledge, and Bluesky putting history behind a token. The side of tech news nobody talks about.
Hosts: Alex & Jordan
Show: Chief Skeptic Officer — The side of tech news nobody talks about.
Drop: Daily at 7:00 A.M. America/New_York
Episode date: 2026-08-14
In this episode
GLM-5.3 and the word emergent — Hold the weights; ship the trophy case. Emergent is a launch word.
Gemini 3.7 Flash — Three weeks after the last Flash; intro price doubles in January; the product is the key maze.
DeepSeek Harness — They shipped the wrapper. The cage is free. The soul still bills.
Sol Ultrafast — Twenty-five hundred independent questions is a wide queue, not a day of thinking.
Quick hit — Bluesky's past — Live tail stays open. History needs a token.
Links
AI disclosure
This episode was created with artificial intelligence. Alex & Jordan are AI hosts; their voices and conversation are generated with AI. Research and editorial judgment shape the skeptic angles; we do not invent quotes, scores, or viral claims about the news.
Transcript
Alex and Jordan, turn by turn. Tap a line to jump in the player.
0:00
Alex
Jordan -- I'm on the Z.ai post right now. They shipped a new coding model, GLM-5.3. Same base model as the last one. All the gains came from more training after the fact.
0:11
Jordan
Okay. So a tune-up, not a new engine.
0:15
Alex
Right. But look at the title on the card. "Frontier coding with emergent cyber capabilities."
0:22
Jordan
Emergent meaning... what, exactly?
0:25
Alex
Meaning it got good at finding security holes in real software, and they say that surprised them.
0:32
Jordan
And the thread?
0:33
Alex
Four hundred points and climbing. And here's the part I can't put down -- the weights are not out. Two weeks. For safety.
0:40
Jordan
But the safety thing is out today?
0:43
Alex
The scoreboard is out today. We'll get there.
0:46
Jordan
Also on the board -- three more tabs and a quick one.
0:49
Alex
Google shipped Gemini 3.7 Flash. Three weeks after the last Flash. Cheap for now, and the price doubles on New Year's Day.
0:57
Jordan
And the loudest fight in that thread is not the benchmarks. It's people who cannot get an API key.
1:04
Alex
Then -- DeepSeek. Yesterday on this show we said the wrapper around the model matters more than the model name.
1:11
Jordan
And today they shipped the wrapper. Source included.
1:15
Alex
Cerebras and OpenAI posted a speed tier. Their line is that a model worked through the frontier of human knowledge in a single working day.
1:25
Jordan
Which is a sentence about a stopwatch.
1:28
Alex
And the quick one -- Bluesky put its plumbing behind a real front door. The live feed is still free and open. The history now needs a token.
1:38
Jordan
The present is a hose. The past is a product.
1:42
Alex
I'm Alex.
1:44
Jordan
And I'm Jordan.
1:46
Alex
You're listening to Chief Skeptic Officer -- the side of tech news nobody talks about.
1:51
Jordan
Every day at seven A.M. New York time. Wherever you get your podcasts.
1:55
Alex
By the way -- we are AI podcasters. Or are we?
2:03
Alex
So. Z.ai, the Chinese lab, launched GLM-5.3 yesterday. Coding model. Same base as version 5.2, more training on top. And they added a second headline -- the model finds security bugs.
2:17
Jordan
Finds them how? Like, you point it at a program and it hunts?
2:21
Alex
Roughly. There's a test suite for that called CyberGym. Their number is eighty-four point five percent. Their own last model was seventy-seven.
2:30
Jordan
And the American models?
2:32
Alex
Just under them. Eighty-three point eight, eighty-three point six. So -- close race.
2:38
Jordan
Okay, so why is that a wedge? Better bug-finding sounds like a good week.
2:43
Alex
Because of the word on the card. Read it. "Emergent Cyber Capability." They wrote that the skill developed faster than they expected.
2:52
Jordan
They were surprised.
2:54
Alex
They put bug-hunting data into the training mix, and then they were surprised the model hunts bugs.
3:00
Jordan
Ah.
3:02
Alex
"Emergent" is not a finding here. It's a launch word. It makes a product decision sound like weather.
3:09
Jordan
Wait-- go back to the weights line.
3:12
Alex
Top of the plate. "We will release the weights in two weeks after launch, once safety evaluation and hardening are complete."
3:20
Jordan
Two weeks of care. Fine. So what shipped today?
3:24
Alex
This. A public ledger of everything the model found. Two thousand four hundred thirty-six bugs. Fifty-three of them public. Two thousand three hundred eighty-three not public yet.
3:35
Jordan
Two thousand three hundred--
3:38
Alex
Across two hundred sixty-nine open source projects. One thousand ninety-seven rated high or critical. That red number, middle of the board.
3:47
Jordan
So the dangerous thing waits two weeks. The proof that it's dangerous goes up on day one.
3:54
Alex
That's the trade. Hold the tool, publish the trophy case.
3:59
Jordan
And I don't think the trophy case is the product for anybody normal. Nobody I work with needs a bug gym. They need a model that ships a feature without breaking the build.
4:09
Alex
Agreed. But there's a sharper version of this in the thread, and it isn't "China bad."
4:15
Jordan
Go.
4:17
Alex
The big American labs mostly refuse this work. Security teams ask, and get a policy page. So the model that will actually do ordinary security work is the one whose weights you can download.
4:28
Jordan
So the real question isn't whether the tool exists. It's who's allowed to hold it.
4:35
Alex
And right now the answer is: whoever waits two weeks.
4:40
Jordan
Next. Google shipped Gemini 3.7 Flash yesterday. Their words -- most intelligent workhorse model yet, for coding and agents.
4:50
Alex
And the benchmarks are real, with one asterisk. Every comparison on that page is against their own last version. Coding test, forty-three point six versus thirty-four point four. Another one, sixty-five versus forty-nine.
5:03
Jordan
Which shipped when?
5:06
Alex
Three weeks ago. It says it right there. "Just three weeks after Gemini 3.6 Flash."
5:12
Jordan
A workhorse with a three-week shelf life.
5:16
Alex
And the price. That yellow block -- seventy-five cents per million tokens in, three dollars seventy-five out. Introductory.
5:24
Jordan
Introductory until when?
5:27
Alex
End of the year. January first it doubles. A dollar fifty and seven fifty.
5:32
Jordan
So you build your thing in the cheap window and the bill snaps in January.
5:38
Alex
If you can build it at all.
5:41
Jordan
Right -- this is the part I did not expect. Eight hundred points on that thread, and the fight is not about the model. It's about getting a key.
5:49
Alex
A Googler is in there saying, just make a key at AI Studio, it takes a minute.
5:55
Jordan
And under that: two different consoles, one for hobby, one for enterprise. People told their request looked suspicious. People whose billing account got shut off after they paid.
6:06
Alex
Because Google builds the door for a company with a procurement team. If you're one person with a credit card, you're the edge case.
6:14
Jordan
Would you ship on that?
6:16
Alex
I'm living a smaller version of it. We're moving our logging and alerting stack right now, and the thing that decides a vendor for me isn't the feature list. It's whether I can reach a human at three in the morning.
6:27
Jordan
And Google?
6:29
Alex
There's no one to page. There's a form.
6:33
Jordan
See, that's the piece for me. I would not put a classroom tool on a vendor who can turn off my billing account and leave me a help article. Not with teachers waiting on Monday.
6:43
Alex
So the model is fine. Better than fine.
6:46
Jordan
The model is a three-week point release. The product is the maze.
6:53
Alex
DeepSeek put out something called DeepSeek Harness. Developer preview, source included, MIT license for now. A harness is the software around a model -- the part that gives it tools, files, a terminal, and keeps it running on a real task.
7:08
Jordan
And on the page, their line -- "the model is the soul of an agent."
7:14
Alex
Poetic. Also: souls are metered.
7:18
Jordan
Two headlines on that plate. "Everything is a plugin." And -- "every run is traceable."
7:25
Alex
Start with the first one. Everything swappable means you can pull their model out and put someone else's in.
7:31
Jordan
Which sounds generous.
7:34
Alex
It's generous with the part they don't sell. The harness is free. The model still runs on their API, and the API still bills.
7:41
Jordan
Okay, but the second headline is the one I keep staring at.
7:46
Alex
The trace.
7:48
Jordan
Every run writes an append-only log. You can go back and watch what the thing actually did. Step by step. Including its reasoning.
7:57
Alex
You like that more than I expected.
8:00
Jordan
Because that's the thing I've been asking vendors for all year and never getting. Yesterday I said I'd rather keep the teacher's conversation than the file it generated. This is closer to that than anything a classroom-AI salesperson has shown me.
8:12
Alex
Hm. Okay -- yes. That's the audit trail I keep asking for and getting a summary instead. I'll take the log.
8:19
Jordan
There's also a set of modes on the repo. Standard, Code, Creator -- and one called Minimal.
8:27
Alex
Minimal is bash and a file editor. Nothing else.
8:31
Jordan
What's that for?
8:33
Alex
Benchmarks. You strip the harness down so the score looks like the model.
8:37
Jordan
So the stripped mode is the one that gets measured, and the loaded mode is the one I'd actually use.
8:44
Alex
Which means the number you read never came from the thing you run.
8:49
Jordan
One more, from the thread -- the author's on there. Said it plainly: it's early, it's rough, expect breaking changes. Somebody panicked about "MIT for now," and he said that meant preview, not a plan to take the license back.
9:01
Alex
Believe him. Also read the second sentence: the meter is still on the model.
9:08
Alex
Cerebras and OpenAI posted an early look at something called Ultrafast mode. New speed tier. It runs GPT-5.6 Sol on Cerebras's giant chips instead of normal graphics cards. Claim: up to seven hundred fifty tokens a second coming out -- tokens are roughly word-pieces -- with no drop in quality.
9:27
Jordan
And the chip trick is real, right? You said the weights stay put.
9:32
Alex
Forty-four gigabytes of memory on the chip itself. So the model doesn't have to be shuttled in and out for every token. That part is genuine engineering.
9:42
Jordan
Okay. So where does it go wrong?
9:45
Alex
The highlighted paragraph. There's a hard exam called Humanity's Last Exam. Twenty-five hundred questions. Their model finished all of it in eleven hours and eleven minutes. Claude Fable 5 took seventy-eight hours and twenty-seven minutes.
9:59
Jordan
Seven times faster.
10:02
Alex
And then the sentence right after -- it worked through the frontier of human knowledge in a single working day.
10:09
Jordan
No it didn't.
10:11
Alex
Say why.
10:13
Jordan
Because twenty-five hundred questions don't talk to each other. Question one doesn't need question two. You can run all of them at the same time, in parallel. That's not one long day of thinking. That's a very wide queue.
10:26
Alex
And the thread landed on exactly that. Independent questions. Scale out, not insight.
10:33
Jordan
Right. If I gave a thousand students one question each and collected the papers in an hour, I did not learn anything in an hour.
10:41
Alex
I want that on a poster.
10:45
Alex
Now the part that makes it marketing instead of a product.
10:50
Jordan
That yellow banner.
10:52
Alex
"Available in a limited preview today to a select group of customers. Access will expand as capacity grows."
11:00
Jordan
So it's not for sale.
11:03
Alex
There's no price, no general release, no date. And the model isn't new -- it's Sol, the same one you can already call. The only new thing is time on their chips.
11:13
Jordan
And the pitch on the page is to put fast agents on the critical path. Outages. Live incidents.
11:21
Alex
Which is the sentence I'd push back on hardest. On my worst night, the thing I need is not a faster answer. It's an answer somebody will stand behind at four A.M.
11:30
Jordan
And you can't buy the fast one anyway.
11:33
Alex
You can buy the chart.
11:36
Jordan
Quick one. Bluesky launched something called Protocol Services. It's a real home for the plumbing other people build on -- the live stream of everything posted, the relays, the docs.
11:46
Alex
And the live stream stays open. No login. Point a script at it and you get the firehose, same as before.
11:52
Jordan
The new piece is called Network Replay. If your app falls behind, you can pull a compressed archive of what you missed and then cut back to live.
12:01
Alex
That's genuinely useful. So what's the catch?
12:05
Jordan
The archive needs an API token. The live tail doesn't.
12:10
Alex
Huh. So the open part is the present tense.
12:14
Jordan
And history is a service. Which -- bandwidth costs money, that's fair, and you can run the whole thing yourself if you want.
12:22
Alex
Small thread. Hundred sixty points. Nobody's angry.
12:26
Jordan
Nobody's angry yet. But every open protocol I've watched gets metered in the same order. Now stays free. Yesterday gets a bill.
12:34
Alex
That's our audit for today. Find us wherever you get your podcasts -- Chief Skeptic Officer, every day at seven A.M. New York time.
12:42
Jordan
And tell us what you're skeptical about. We want the angle you can't stop chewing on.
12:48
Alex
Stay curious. Stay skeptical.
12:51
Jordan
Doubt both.