August 2026
The Ascend Playbook
How China built a frontier AI with no Nvidia.
Got the code from the video?
Enter it with your email to unlock the deeper cut, keep a PDF of this file, and claim founding-member access on the HUMBL app.
The real casualty here isn't a rival chatbot — it's the assumption that the Entity List actually stops anything. Sanctions bet that cutting off the tool stops the output. GLM-5 is the receipt that a determined team just rebuilds the tool instead, on a timeline measured in months, not decades.
Not Nvidia's top line this quarter — its moat's shelf life. CUDA's value was that nobody had a reason to leave it. Zhipu just proved a full alternative stack ships in under a year when there's no other option. Every lab now has a credible reason to at least benchmark the exit.
Two tells. One: whether MindSpore/CANN adoption shows up outside China — that's the real test of whether this is a sovereign-stack story or a genuinely portable one. Two: GLM-5.2's inference speed on the next point release. If the tokens/sec gap to Nvidia closes even partway, 'close, not king' stops being the honest read.
Enter your code + email to unlock this file's bonus.
Not from Instagram? The code's hiding in today's video, not here.
The moat was never the chip. It was the software.
Everyone thinks the chip ban was about hardware. It was about software. For fifteen years, Nvidia's real lock on AI has not been silicon. It has been CUDA, the programming layer every AI engineer on earth learned on.
When you ban a company from Nvidia, you do not just take its chips. You take the entire toolchain: the compiler, the libraries, the debuggers, the frameworks, fifteen years of tooling and community knowledge built on top of CUDA. That is the wall Zhipu actually climbed. Every layer of the Western stack had to be swapped for a younger, thinner Chinese counterpart. The tuned kernels every lab takes for granted did not exist, so they wrote them by hand.
Being cut from Nvidia meant rebuilding fifteen years of tooling from a less mature base.
| Layer | Western stack | China's stack |
|---|---|---|
| Framework | PyTorch | Becomes MindSpore — Huawei's framework, far less mature |
| Low-level | CUDA | Becomes CANN — China's answer to CUDA |
| Chip-to-chip | NCCL | Becomes HCCL — the library that moves data between chips |
| Kernels | CUDA-tuned | Hand-written in Ascend C — the tuned versions do not exist |
The hardware handicap, quantified. Per-chip FP16 throughput — the Ascend 910B is roughly A100-class and far behind the H100, so raw chips could never close the gap. Software and scale had to.
| Chip | FP16 throughput |
|---|---|
| Huawei Ascend 910B | ~320 TFLOPS |
| Nvidia A100 | ~312 TFLOPS |
| Nvidia H100 | ~989 TFLOPS |
He chose to stay.
Tang Jie was born in 1977 in Sichuan, into an ordinary family. He studied Automation before switching to computer science, and earned his PhD at Tsinghua in 2006.
At the crossroads every graduate faced, go abroad or stay, he stayed. He was inspired by Wang Xuan, the scientist who leapfrogged forty years of foreign typesetting technology from inside China.
He built AMiner, a map of the world's scientific research, on a single desktop with roughly 20,000 yuan of prize money. He founded Zhipu in 2019 on three months of free rent.
In 2022 he shipped GLM-130B, a hundred-billion-parameter model, before ChatGPT existed, beating GPT-3 on some benchmarks. His whole career was one argument: China could build foundational technology from scratch. So the ban did not break his thesis. It was his thesis.
Built on one desktop.
Three months free rent.
Before ChatGPT.
Nvidia cut off overnight.
Trained on Ascend.
Open under MIT.
“In the 1970s, when China's conditions were far behind, Wang Xuan solved Chinese laser typesetting and leapt forty years in one step. I was deeply inspired.”
— Tang Jie, on why he stayed
They could not out-spec Nvidia. So they out-engineered the ban.
Four moves. Every one exists because they had less: weaker chips, a slower interconnect, an immature toolchain, and no way to buy their way out.
| # | Move | What it does |
|---|---|---|
| 01 · Toolchain | Rebuild the whole stack | Ran entirely on MindSpore and CANN, and hand-wrote custom operators in Ascend C because the CUDA-tuned versions simply do not exist there. Zero CUDA, zero Nvidia. |
| 02 · Sparsity | Let sparsity beat the weak link | A 744B Mixture-of-Experts fires only about 44B parameters per token, plus sparse attention for long context — less data crossing the slow interconnect means the weak link hurts less. |
| 03 · Slime | Never let a chip sit idle | Decouples generating from training and runs 1,000+ practice runs in parallel — roughly 10x throughput. The crown jewel, and it's open-source. |
| 04 · Scale | Scale beats speed | Each Ascend chip is roughly A100-class, well behind an H100, so they used 100,000 of them, made by SMIC on 7nm without EUV — about 15% more compute time, offset by cheaper chips and state support. |
They could not buy more compute. So they built the machine that wastes none of it.
What it means, and what to watch.
- The sovereign stack is real — silicon to framework to model, now adapted across more than forty domestic chip families. The ban created the thing it meant to prevent.
- The gap collapsed — Chinese open models went from about a year behind the West to a matter of months, on hardware the West wrote off.
- The price floor breaks — a near-frontier model that is free and open forces every paid lab to justify its cost, every month.
- GLM-5.2 ranks around #3 overall, behind the leading Western models. Close, not king.
- Inference is genuinely slower, roughly 17 tokens/sec vs 25+ on Nvidia, and the CANN inference ecosystem is still immature.
- HK$1 trillion+ market cap on ~$100M revenue — a price-to-sales ratio over 1,000x, on a multi-billion-yuan loss.
- On the Entity List for alleged military ties, which Zhipu denies. Hosted data sits under Chinese jurisdiction.
Elon Musk: "A Chinese model this good? Probably Q1 2027." Tang Jie: "Won't take that long." Elon Musk: "Benchmarks aren't real intelligence. The real test shows up in revenue." Tang Jie: "Focus is all we need. Especially focusing on what intelligence truly is."
The ban was built to slow China down. It just proved they never needed America's chips.