Experts Question White House Claims That China's Kimi K3 Copied Anthropic's Fable

Experts Question White House Claims That China's Kimi K3 Copied Anthropic's Fable

The White House has accused Moonshot, the Chinese company behind the Kimi K3 large language model, of building its AI system by copying Anthropic's Fable model while using export-restricted chips. But independent experts are pushing back, saying the technical evidence doesn't support the allegations.

Michael Kratsios, White House science advisor, described Moonshot's practices as "large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research." His comments come amid reported discussions about potentially banning Chinese open-weight models, a prospect that has unsettled the AI industry.

Moonshot did not respond to questions about its training process, and Kratsios did not provide additional details about the sources behind his allegations.

Officials Point to "Watermarks" in Chinese Models

Kratsios' remarks echoed statements from Treasury Secretary Scott Bessent, who said that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." The specific nature of these watermarks remains unclear, and the Treasury Department did not respond to requests for clarification.

Kimi K3 is currently described as the largest available open-weight large language model, making it a significant player in the increasingly competitive global AI landscape.

Experts Say Distillation Alone Can't Explain Kimi K3's Capabilities

Despite the official accusations, researchers specializing in AI development express skepticism that distillation — the process of systematically querying a model to understand and replicate its capabilities — could fully account for Kimi K3's performance.

Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, pointed to timing as a key problem with the distillation narrative. Fable has only been publicly available since July 1st, and Hancock argued that there simply wasn't enough time to distill sufficient data, train a model, and release it within roughly two weeks.

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock told TechCrunch. "There's just not even frankly time, right?"

Nathan Lambert, an AI researcher at the Allen Institute for AI, offered a similar perspective in a recently released podcast. He noted that distillation has become less impactful as Chinese models approach the frontier of AI capabilities and as training methodologies shift toward reinforcement learning.

Lambert argued that if distillation alone were sufficient to close the gap, competitors could easily replicate models like GLM or Kimi K3 by using their data. The fact that this hasn't happened through supervised fine-tuning alone suggests more sophisticated techniques are at play.

The Technical Reality of Modern AI Training

Distillation typically involves querying a target model to generate training data, sometimes asking the model to articulate its chain-of-thought reasoning. The resulting prompts and responses can then be used to train a new model through supervised fine-tuning, or SFT.

This fine-tuning process is where models can inadvertently pick up identifying characteristics — what Lambert described as a model "picking up its manners." This could explain instances where a third-party model claims to be Claude, Anthropic's flagship product.

However, Lambert emphasized that as models grow more complex, the benefits of SFT diminish. Replicating Fable-level capabilities would likely require reinforcement learning techniques, where a larger model grades a smaller model's responses and adjusts accordingly. These advanced methods demand substantial infrastructure, with large reinforcement learning runs potentially requiring tens of millions of agents.

Using a frontier lab's API for such operations would be, in Lambert's words, "insanely expensive" and potentially a time bottleneck, given the speed of current models. He also noted it might not even deliver a meaningful performance improvement.

A History of Distillation Accusations

The allegations against Moonshot are not entirely without precedent. Earlier this year, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of systematically distilling its models. The company said it identified millions of exchanges between its models and users at those companies through IP addresses and other metadata, describing the queries as "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." Anthropic did not respond to further queries about Fable distillation.

Yet distillation is widely acknowledged as common practice across the AI industry, not only in China. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, calling the practice commonplace. The boundary between distillation and the development of synthetic data sets can be difficult to draw.

Hancock cautioned against underestimating Chinese AI teams' technical expertise, noting that one of Moonshot's founders was a PhD student at Carnegie Mellon University. "These are legitimate researchers and engineers doing solid work," he said. While American models contribute to global progress, Hancock believes China's development would continue even if U.S. progress stalled. "They're not just riding coattails here," he added.

Chip Smuggling Concerns Add Another Layer

Kratsios' allegations extended beyond distillation, claiming that Moonshot had obtained advanced Nvidia Grace Blackwell 300 chips and accessed GB300-equipped servers in Thailand. These chips are banned from export to China.

Sam Bresnick, a research fellow at Georgetown's Center for Security and Emerging Technology, confirmed that a black market for such chips exists. In May, the founder of Supermicro, a U.S. server builder, was indicted for allegedly smuggling advanced chips into China.

Bresnick advocated for "know your customer" laws for data centers worldwide, arguing that companies conducting large training runs on state-of-the-art hardware should be subject to reporting requirements. The Biden administration's Department of Commerce proposed federal know-your-customer rules for data centers in 2024, but no visible progress has been made under the Trump administration. Exporters shipping advanced chips abroad are technically required to ensure they are used only for approved purposes.

The intersection of distillation allegations, chip smuggling concerns, and potential regulatory action against Chinese open-weight models represents a growing flashpoint in the global AI race. As governments weigh restrictions and companies deny wrongdoing, the debate over who truly owns AI capabilities — and how they are developed — shows no signs of slowing. What do you think about the allegations against Moonshot and the broader debate over AI distillation? Share this article and join the conversation.

Source: TechCrunch