July 2026
The Kimi File
How to build frontier AI when you're poor.
Got the code from the video?
Enter it with your email to unlock the deeper cut, keep a PDF of this file, and claim founding-member access on the HUMBL app.
Nobody's pricing this in yet: the real casualty of Kimi isn't a rival chatbot, it's Nvidia's roadmap. A trillion dollars of infrastructure was built on the assumption that intelligence stays expensive to produce. Kimi is the first receipt suggesting that assumption was wrong.
Not OpenAI or Anthropic — they have brand, distribution, and enterprise contracts Kimi can't touch overnight. The squeeze lands on the middle: the API wrapper startups whose entire pitch was 'cheaper than the frontier labs.' That moat evaporates the day cheaper-than-frontier ships from Beijing for free.
Two tells, both public. One: how fast a Western lab quietly cuts API pricing in the 90 days after July 27 — that's the real scoreboard, not the benchmark charts. Two: fork and download counts on the open weights once they land. Adoption curve, not press coverage, is what tells you if this was a moment or a shift.
Enter your code + email to unlock this file's bonus.
Not from Instagram? The code's hiding in today's video, not here.
Why AI costs billions
Every frontier lab hits the same four walls. Money breaks them — reported US frontier runs sit in the hundred-million-dollar class, with over a trillion committed to infrastructure. Moonshot couldn't pay, so they went through each wall a different way.
| Wall | Why it breaks people |
|---|---|
| Chips | Tens of thousands of GPUs for months. China is banned from the best ones. |
| Data | The good text on the internet has largely been read. |
| Crashes | One exploding number destabilizes a run. Weeks of compute burn. |
| Serving | Training is one bill. Answering millions daily, forever, is the bigger one. |
Kimi K2 Thinking reportedly cost 4.6 million dollars to train. Not a typo. [CNBC]
One career. One problem. Memory.
Yang Zhilin. Born 1992, Shantou. Tsinghua, then a Carnegie Mellon PhD in four years. In 2019 he co-wrote Transformer-XL and XLNet — the research that stretched AI's memory window. Google Brain and Meta brought him in. Then he went home. March 2023: he founds Moonshot AI in Beijing, named for his favorite album, The Dark Side of the Moon. 2024: Kimi is #3 in China. January 2025: DeepSeek drops R1 and buries everyone — Kimi falls to #7. Investors say copy DeepSeek, go smaller. He does the opposite. [Papers 2019 · VentureBeat · public rankings]
“Token efficiency is not just about efficiency. It's about improving the upper bound of intelligence.”
— Yang Zhilin, Moonshot AI founder
Seven hacks. Four walls.
| # | Hack | What it does |
|---|---|---|
| 01 | Wake up 2% of the brain | 2.8T parameters, but each word wakes only 16 of 896 experts. Giant knowledge, small-model running cost. |
| 02 | The run that never crashed | MuonClip's QK-Clip tripwire caps exploding values. 15.5T tokens of pretraining, zero loss spikes. |
| 03 | Learn more per word | The Muon optimizer extracts more learning per token than AdamW. Same data, more intelligence. |
| 04 | The experience factory | Thousands of synthetic environments where agents practice tasks; only verified successes become training data. |
| 05 | The model grades itself | Reinforcement on machine-verifiable answers, then self-critique against rubrics. No human graders. |
| 06 | Cheap math, from day one | Trained in compressed low-precision math mid-run onward. Half the memory, twice the speed, tiny quality loss. |
| 07 | The plumbing nobody sees | Mooncake splits reading and writing across machines, recruits idle hardware. 107–115% more requests, Best Paper, 100B+ tokens a day. |
The wall was the strategy.
The table that scared Wall Street
Kimi K3: $3 in, $15 out per million tokens — near half the per-task cost of Claude's flagship. The honest caveat: K3 thinks out loud and burns more tokens per task, so the gap narrows on some workloads. The market did the math anyway — launch week, over a trillion dollars gone from chip stocks. Not because Kimi is the best. Because it's close enough, at these prices, for free. [Artificial Analysis · market data, July 2026]
| Model | Reported training cost |
|---|---|
| Kimi K2 Thinking | $4.6M |
| DeepSeek V3 | $5.6M |
| US frontier runs | $100M+ |
- Anthropic: three Chinese labs, Moonshot among them, allegedly ran ~24,000 fake accounts and pulled 16M+ exchanges out of Claude.
- OpenAI filed similar claims. [Feb 2026]
- Every frontier model was trained by scraping the internet without asking.
- Musk confirmed Grok distilled from OpenAI — the technique is universal; only the target is contested.
- Moonshot disputes the characterization.
Why is copying the internet business, but copying the copier theft?
Free is the weapon
July 27, the full weights go public. Three detonations: the price floor collapses (US labs already adjusted plans), the world becomes his lab (millions of developers improving Kimi for free), and the sanctions logic cracks — if engineering under constraint substitutes for raw compute, a trillion dollars of US infrastructure is priced on a shakier assumption than the market believed.
Most performance numbers are Moonshot's own, not yet independently verified. One independent test flagged K3's hallucination rate near 51 percent — Moonshot says they're fixing it. Anything sent to the hosted version sits under Chinese jurisdiction; serious companies self-host the open weights. He's not a saint. He's a strategist.
The side you can't see is still building.