Why I’m betting on AI distillation — and cheaper models

AI distillation is slashing inference costs, alarming frontier labs and forcing Washington to decide where optimization ends and theft begins.

Why I’m betting on AI distillation — and cheaper models

AI distillation is the generic-drug moment big labs were dreading

From Silicon Valley to DC, the tech world is suddenly obsessed with one concept in AI: distillation. The technique is old. The panic arrived when it started crushing prices.

Kimi K3 showed up with roughly frontier-level coding performance at around half the price of OpenAI’s GPT-5.6 Sol, and I immediately opened a spreadsheet.

I stare at AI pricing pages the way my nonna inspected tomatoes at the market: suspiciously, personally and fully prepared to walk away over an insultingly soft San Marzano.

In my spreadsheet, cheaper inference means better gross margin. In Washington, it apparently means a national emergency.

The fight over AI distillation sounds technical because “teacher-student model compression” photographs poorly at a Senate hearing. Underneath, the argument concerns who gets to charge a premium for intelligence, and for how long.

Frontier labs funded an extraordinarily expensive discovery process. Distillation can reproduce much of the useful behavior without copying the original model weights. AI has reached its generic-drug moment, except there is no settled patent framework to keep everybody civilized.

Copyright law, patents, contracts and trade-secret rules all cover pieces of the dispute. None maps cleanly onto a model learning from another model’s answers.

I understand why the labs are nervous. I would be too.

Google has used this “dangerous trick” for years

Knowledge distillation has existed in machine learning for more than a decade. A large teacher model produces answers or guidance. A smaller student learns to reproduce useful parts of that behavior with less computation.

I think of it as compressing expertise. The student never receives the teacher’s original weights, internal architecture or training dataset. It studies examples of what the teacher does.

Google AI chief Jeff Dean described the method in entirely normal engineering terms during a February 2026 podcast:

Through distillation, which is a key technique for making the smaller models more capable, you have to have the frontier model in order to then distill it into your smaller model.

Nvidia used distillation while training its Llama Nemotron models too. Nobody summoned a Senate committee when Jensen Huang’s company did it. The technique acquired a sinister aura once Chinese labs became good at it.

The results now justify the attention. The BIRD paper, submitted to arXiv on July 17, 2026, gives a fairly wild example using Qwen3-8B. Leichao Dong and his co-authors increased accuracy on MATH-500 from 86.2% to 92.0% while cutting the average response from 3,099 tokens to 1,115.

Better answers with roughly 64% fewer tokens. Every founder paying an inference bill just sat up a little straighter.

BIRD uses self-distillation, meaning the model learns to produce a cleaner, shorter version of its own reasoning. There is no foreign competitor lurking in a trench coat. The model is editing itself after realizing it talks too much, a development I support on behalf of every human trapped in a 90-minute Zoom meeting.

Large reasoning models often burn tokens repeating checks or wandering through dead ends. Distillation can keep the useful capability and throw away the expensive verbal furniture.

That hits frontier labs directly. Their models can cost billions of dollars to develop, yet their outputs can teach cheaper systems how to perform commercially valuable tasks. The lab keeps its weights while watching the scarcity premium evaporate.

Generic drugs work similarly: the manufacturer reproduces a useful result after somebody else paid for discovery, without recreating the original lab notebooks. The analogy gets legally messy in AI, but the market pressure is already here.

Optimization for me, theft for thee

Anthropic has presented evidence that deserves serious scrutiny. In February 2026, the company said DeepSeek, Moonshot and MiniMax created roughly 24,000 fake accounts and generated 16 million exchanges with Claude during industrial-scale distillation campaigns.

Those numbers describe something far beyond a graduate student’s weekend experiment. If Anthropic’s account is accurate, the fake identities and systematic access evasion show deliberate conduct. I have zero philosophical attachment to account farms.

White House science adviser Michael Kratsios escalated the allegation against Moonshot on July 22:

We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.

Kratsios also alleged that Moonshot built an internal platform capable of switching among access methods to avoid detection. Fortune reported that he connected the operation to Nvidia GB300 systems in Thailand, hardware barred from sale to Chinese companies under current US export controls.

That is a specific accusation involving access evasion and possible export-control violations. It needs evidence. As of July 25, the public has received the claim without the technical receipts.

The timeline complicates it further. Anthropic’s Fable became public on July 1. Kimi K3 launched on July 15. Moonshot employee Randy Xian responded with the subtlety of an Italian waiter being asked for ketchup:

Yes, Fable went public on July 1 and K3 launched on July 15. We trained a brand new frontier model in JUST 15 DAYS. Guinness World Record stuff.

That timing cannot clear Moonshot of every possible distillation claim. Earlier access, other Anthropic models or a late-stage post-training process could still matter. It does show why benchmark similarity alone cannot carry the accusation.

I draw the enforcement line around conduct: fake-account farms, credential fraud, security circumvention and deliberate contract evasion. A court or regulator can examine those actions.

Ownership of publicly delivered answers is much harder. Anthropic may prove that somebody broke into the classroom. That would not automatically establish ownership over everything the intruder learned there.

The selective outrage weakens Silicon Valley’s case. OpenAI and Anthropic built their systems using enormous quantities of human-created material, and both have faced lawsuits over that training. CNBC quoted Max Pritt, an attorney representing authors suing AI companies, criticizing the administration for vigorously defending tech companies’ intellectual property while staying largely silent about the creators whose work trained those systems.

Sixteen million exchanges generated through fake accounts belong in a different legal bucket from reading a public webpage. Still, frontier labs now want expensive investment to justify broad control over model outputs. Authors, journalists, artists and programmers made a similar argument about their work. Silicon Valley responded with rather less flag waving.

My honest confession: I sympathize with Anthropic more than my tone suggests. I have spent 20 years building products, and watching a competitor copy months of work can make you physically nauseous. I have also learned that anger writes terrible property law.

AI distillation is crushing the price of intelligence

Kimi K3 matters because buyers now have a cheaper option that appears good enough for serious work. The Associated Press reported that K3 topped Arena’s ranking for front-end coding capability.

Arena CEO Anastasios Angelopoulos gave the release a fairly unambiguous review:

This may be the single biggest release of the year.

Bank of America analysts cited by AP estimated that K3 costs about half as much to use as OpenAI’s GPT-5.6 Sol. That comparison will keep frontier-lab executives awake. Most customers have no desire to fund the theoretically smartest model on Earth.

I have shipped software for two decades through Ad Astrum, from ALYT home-automation hardware to a connected Pascucci espresso machine. Customers care about whether the system clears a reliability threshold at a price their business can support. Benchmark prestige ranks somewhere below “does this break on Friday night?”

Hardware taught me that lesson painfully. A beautiful feature becomes irrelevant when cloud costs eat the margin or an app update ruins device pairing. I used to assume technical superiority bought more time than it does.

It rarely does.

“Good enough, available and affordable” has buried plenty of technically superior products. AI will receive no special exemption because the fundraising deck had impressive GPU diagrams.

SecurityPal founder Pukar Hamal told CNBC that he would consider hosting Kimi K3 on the company’s own infrastructure after checking it for backdoors. His reason was wonderfully unromantic: it could save significant money.

That is how purchasing happens. Security review comes first. Ideology appears around item 14, after uptime and the finance person asking why last month’s token bill resembles Milan rent.

Nvidia sees an upside because cheaper models encourage more usage. In a July 22 interview with Axios, Jensen Huang argued that excellent Chinese open models should be available to American companies. Wider adoption means more demand for chips and data centers.

Huang put the principle plainly:

Distillation, learning from AI, learning from other sources of knowledge, is fundamental to intelligence.

Nvidia earns money as AI consumption spreads. Closed frontier labs earn more when high-end intelligence stays scarce and metered through an API. The same Kimi release looks like market expansion in Santa Clara and margin compression in San Francisco.

When a capability goes from a luxury tasting menu to a €12 plate of pasta, the chef calls it commoditization. The investor calls Washington.

A diagram illustrating AI distillation process, showcasing model efficiency and cost reduction in technology development.

Alt text: AI distillation diagram showing synthetic data and capabilities flowing between American and Chinese teacher and student models.

The knowledge flow already runs both ways

Washington’s clean story about American invention flowing outward has expired. AI development now resembles a group chat where everybody borrows code, papers, synthetic data and the occasional suspiciously familiar idea.

Mira Murati’s Thinking Machines raised $2 billion, then disclosed that its Inkling model drew from DeepSeek-V3’s architecture. Its post-training also incorporated synthetic data generated by Moonshot’s Kimi K2.5, according to Rest of World.

Thinking Machines is hardly a basement operation using random files from Hugging Face. It is one of Silicon Valley’s highest-profile AI companies, founded by OpenAI’s former chief technology officer.

San Francisco startup Anysphere has acknowledged that one of Cursor’s leading products was based on Kimi K2.5. AP reported that SpaceX plans to acquire Cursor for $60 billion.

Chinese model capability is already embedded inside an American software company carrying a proposed valuation larger than Ford’s market capitalization on many trading days. Geopolitical purity gets complicated once the acquisition bankers arrive.

Apple offers an even stranger example. After receiving Chinese regulatory approval, Apple planned to deploy Apple Intelligence in China with Alibaba’s Qwen and Baidu’s Ernie as core components of its local stack. Rest of World reported that the US Department of Defense has designated both companies as Chinese military-affiliated.

An American iPhone sold in China can therefore run approved Chinese AI models, then travel back through LAX in somebody’s pocket. Good luck drawing that supply chain on a Cold War map.

Engineers have shipping deadlines. I care about how a model performs, what its license allows, deployment cost and whether I can customize it. Nationality matters when legal or security exposure appears. A passport has yet to improve coding accuracy.

OpenAI, Anthropic, Google DeepMind, Meta, Alibaba, Tencent, Moonshot, DeepSeek and Zhipu have all published work involving synthetic training data or teacher-student methods. The competitive question concerns which models become teachers and under what permissions.

Blanket restrictions will struggle against the mechanics of distribution. A government can block an advanced Nvidia chip shipment. Quarantining an architecture described in a paper gets much harder after model weights land on Hugging Face, GitHub, cloud providers and local machines.

Europe should be paying very close attention. I grew up in Ivrea, the town of Olivetti, and studied computer engineering at Politecnico di Torino. Watching Europe depend on American closed APIs while Chinese open-weight models set the pricing floor makes me deeply uncomfortable.

On April 9, 2025, European Commission executive vice-president Henna Virkkunen launched the AI Continent Action Plan with a useful bit of urgency:

The global race for AI is far from over. It is time to act.

Correct. Fragmented national strategies will leave Europe renting intelligence from two foreign power centers. European companies need capital and compute, plus a continental market where they can sell AI products without rebuilding the business country by country.

Regulation without European champions produces excellent paperwork and strategic dependency. Bravissimo.

If my model is the moat, mamma mia

A startup whose entire advantage comes from wrapping the smartest API is standing on melting ice. Distillation turns yesterday’s premium capability into tomorrow’s cheap dependency.

I say this as somebody who has built plenty of infrastructure-heavy products. With ALYT, I lived through firmware and hub hardware, plus the mobile app and app-store politics. Every infrastructure advantage eventually migrated upward into the baseline customers expected.

Faster databases did not destroy software companies. They destroyed “we have a database” as a compelling pitch.

AI is following the same path. Durable value lives in proprietary workflow data and customer trust. Domain evaluations matter. So do deep integrations, permission systems and feedback collected from actual use.

Those assets sound less exciting than a benchmark screenshot. They also survive when model prices fall 70%.

Recent research shows how far distillation has moved beyond copying chatbot style. A July 23 paper by Chenhui Gou and four co-authors introduced “Experience Distillation,” which converts an agent’s interaction history into reusable model behavior without requiring new environment calls.

Across 749 software-engineering tasks and six text-adventure games, the method retained at least 64.8% of the gains produced by in-context learning. Direct supervised fine-tuning recovered only 3.8%.

The same approach matched reinforcement-learning baselines while using at least 9.6 times fewer environment samples. For agents that learn through costly experiments or human feedback, that gap can decide whether the unit economics work.

OPOD, another paper submitted on July 23, coordinated separate teachers for text, images and audio. Across 12 benchmarks and three model sizes, it produced the highest average score at every scale.

At the 30-billion-parameter level, the resulting model beat its base model and a jointly post-trained counterpart on all 12 benchmarks. The specialist teachers could then be discarded, leaving one deployable multimodal model.

I expect expensive specialist systems to teach cheaper product models and then disappear from production. Customers will get lower latency and smaller bills. Nobody buys champagne for the teacher.

For founders, my checklist is blunt:

  • Assume model prices will keep falling.
  • Keep the product portable across API providers and open-weight models.
  • Build internal evaluations around customer outcomes rather than Arena screenshots.
  • Retain control of sensitive data and the feedback users generate.
  • Avoid calling “our model” a moat unless the company genuinely trained and owns something defensible.

I run my own Docker stack on Linux for similar reasons. My Ghost site, ERP and analytics run alongside mail, automations and a SvelteKit image interface behind infrastructure I control. Self-hosting occasionally makes me question my life choices at 1:12 a.m., but portability becomes valuable the moment a vendor changes its prices or policies.

Any startup planning around permanent access to one magical model is volunteering to become somebody else’s pricing experiment.

Punish the break-in and leave studying alone

The White House has started drawing a useful boundary. In a July 24 Axios report, Kratsios defended authorized distillation that creates efficient models while condemning covert extraction at industrial scale.

He wrote:

Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem.

I agree. Enforcement can target fake accounts and stolen credentials. It can also cover prohibited automation, access-method rotation, privacy breaches and the circumvention of technical controls.

Sanctions require more than benchmark vibes. Treasury Secretary Scott Bessent has threatened sanctions and Commerce Department Entity List designations if Chinese companies cross into IP theft:

When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.

Fine. The government should disclose enough technical evidence to distinguish an extraction campaign from a model that independently reached similar performance. Watermarks and account patterns could support the case. So could prompt distributions and access logs, without exposing every sensitive defensive detail.

Broad restrictions would punish American startups first. WIRED reported that more than 200 companies, organized through the Little Tech Association and including Y Combinator, wrote to Kratsios and Commerce Secretary Howard Lutnick opposing an outright ban on foreign open-weight models.

Removing affordable alternatives would entrench a small club of American frontier labs. Apparently monopoly pricing becomes patriotic after the correct lobbying meeting.

A separate letter signed by Nvidia, Microsoft, Meta, Palantir, Box and more than 20 other companies warned against premature restrictions on open-weight AI. The signatories described distillation as routine model development:

Distillation, or the practice of using one model’s outputs to help train or improve another, is a widely used technique for model improvement, evolution, and validation.

Researchers would take collateral damage too. Suresh Venkatasubramanian, a former Biden White House adviser, told Axios that losing access to Chinese open-weight models would create a major problem for scientific research. Open weights let researchers inspect and modify systems in ways closed APIs do not permit.

OpenAI co-founder Greg Brockman has framed adversarial distillation as a technical issue. According to Semafor and Axios, OpenAI uses machine learning plus human review to identify mass synthetic-data generation, attempts to score responses and efforts to extract reasoning.

Capable labs should invest there. Better detection can make abusive extraction expensive while leaving ordinary research and authorized training alone.

Huang’s security argument deserves attention as well. He told Axios that relying on one closed model creates a single point of attack and failure. Downloadable models can run inside controlled environments, where companies can inspect them and restrict network access.

I support sanctions when industrial extraction is proven. A vague anti-distillation doctrine would protect incumbents while weakening the startups Washington claims to champion. A flag pin does not improve bad competition policy.

The teacher does not get to retire

By 2028, distillation will sit invisibly inside nearly every serious AI product. Frontier systems will teach cheaper specialists. Agents will internalize lessons from expensive interactions. Multimodal teachers will collapse into single deployable models.

Self-distillation will trim wasted reasoning without any outside teacher. The term itself may fade from product marketing because customers will simply expect lower latency and lower prices.

Frontier labs will remain essential. Somebody has to create capabilities worth compressing, and the strongest teachers will keep setting the ceiling. Their advantage will come from improving faster and running dependable infrastructure, while earning enough trust that customers keep paying for the premium service.

Being first grants no permanent retirement plan. The teacher has to keep teaching.

An American AI strategy that depends on nobody learning from American models has the structural integrity of a wish.

Frequently asked questions

What is AI distillation?

AI distillation is a machine-learning technique in which a teacher model produces answers or guidance that a smaller student model learns to reproduce. The student does not receive the teacher’s original weights, internal architecture or training dataset, allowing useful capabilities to be delivered with less computation.

Why is AI distillation controversial?

AI distillation is controversial because frontier labs invest heavily in developing advanced models while cheaper systems can learn from their outputs. The dispute involves model pricing, intellectual property and access rules, especially when companies allegedly use fake accounts, stolen credentials or technical circumvention to collect outputs at industrial scale.

How does AI distillation reduce AI costs?

AI distillation can reduce costs by teaching smaller models to preserve useful capabilities while using fewer tokens and less computation. It can remove repetitive reasoning, transfer lessons from expensive interactions and combine specialist knowledge into deployable models, producing lower latency and smaller inference bills for customers.

Sources

Related reading