Connect with us

Hi, what are you looking for?

Blog

The Software End-Run: How DeepSeek, Qwen, and Open-Weight Models Became the Sanctions Workaround

How DeepSeek, Qwen and open-weight AI models became a state tool for sanctions circumvention — and why chip export controls can’t stop software that spreads as bits.

How DeepSeek, Qwen, and Open-Weight Models Became the Sanctions Workaround
How DeepSeek, Qwen, and Open-Weight Models Became the Sanctions Workaround

Exports controls were designed to choke China’s AI at the hardware bottleneck. The weights slipped through instead.

Since October 2022, the United States has tried to slow China’s AI development by cutting off the physical inputs: advanced semiconductors. The logic was simple and elegant no advanced chips, no frontier models. It worked for a while. Then DeepSeek released R1 in January 2025, and the theory cracked. The model matched or beat American reasoning systems while being trained, by the lab’s own accounting, on export-controlled chips that were a generation old, at a fraction of the cost.

The deeper story is not about DeepSeek’s algorithm. It is about what sanctions can and cannot intercept. Chips are physical objects that cross borders and can be counted, banned, and seized. Model weights are patterns of numbers that can be copied infinitely, hosted on any server, and downloaded by anyone. When capability migrates from the first category to the second, the entire architecture of export control — which presumes a chokepoint — loses its target.

This article examines how open-weight Chinese models became a state-scale tool of technological influence and sanctions circumvention, which mechanisms made it work, who actually adopted them, and where the strategy hits its own ceiling.

The control regime developed in three stages:

  • October 2022: BIS restricted advanced computing chips (A100-class) and manufacturing equipment to China.
  • October 2023: Controls were tightened to catch workaround chips (A800, H800) and to require licenses for a wider category of advanced processors.
  • January 2025: The Biden administration’s AI Diffusion Rule tried something new — for the first time it controlled closed AI model weights under ECCN 4E091 and split the world into three access tiers. Notably, it controlled closed weights only; open weights were exempt.

The premise behind all three was the same: deny the chips, deny the model. What the October 2023 rule could not foresee was that a startup would demonstrate that frontier-adjacent performance was achievable without the newest chips, using efficiency engineering — and that the resulting models would then be distributed as open weights that no physical border could stop.

Export controls regulate things: integrated circuits, servers, lithography machines. Each unit is traceable. Model weights are not a thing in that sense. A trained neural network is a file. Once released under a permissive license, it can be:

  • copied without limit,
  • downloaded from dozens of mirrors,
  • modified and redistributed by third parties,
  • and run on any hardware with sufficient (not necessarily sanctioned) compute.

This is the structural gap at the heart of the sanctions regime. The AI Diffusion Rule explicitly exempted open-weight models from ECCN 4E091 — the administration chose not to control what it could not practically control. That exemption effectively conceded that the chokepoint strategy only works for the closed frontier, and that the moment capability is published as open weights, it is outside the reach of export law.

The adoption numbers show how completely the strategy succeeded on the distribution side. The most rigorous measurement, the ATOM Report (arXiv 2604.07190), tracks downloads on Hugging Face:

MetricChinaUnited States
Cumulative downloads (March 2025)97M177M
Cumulative downloads (March 2026)1.15B723M
Year-over-year growth11.9×4.1×
Date of overtakingLate July 2025

Chinese open models overtook US models in cumulative downloads in late July 2025 and have widened the gap since. Total tracked open-model downloads reached 2.04 billion by March 2026 — a 6× year-over-year jump.

Within that ecosystem, one family dominates. Qwen (Alibaba) surpassed Meta’s Llama as the most downloaded model family on Hugging Face in September 2025, passed 700 million cumulative downloads in January 2026, and crossed 1 billion by April 2026. In February 2026, Qwen generated 153.6 million downloads — more than the next eight model families combined — and 69% of all new fine-tunes on Hugging Face that month were Qwen-derived. The family has spawned more than 113,000 direct derivative models (roughly 180,000 counting all tags) — more than Google and Meta combined.

On OpenRouter, the neutral routing layer developers use to compare models, Chinese-origin models peaked at roughly 46% of enterprise token volume by mid-2026 against 35.7% for all US models combined, and Qwen alone handled 13.9% of all routed tokens. Chinese models have been priced 60–90% below US equivalents.

Openness is not only technical; it is contractual. DeepSeek’s models ship under the MIT license. Qwen’s open releases ship under Apache 2.0 — the same license family that powered Linux and Kubernetes. Apache 2.0 permits commercial use, modification, and redistribution with no revenue caps and an explicit patent grant.

Meta’s Llama, by contrast, ships under a custom “community license” with usage thresholds and conditions that require legal review. In practice, this legal asymmetry has driven enterprise adoption: companies that would spend weeks getting Llama approved by counsel can deploy Qwen the same day. Alibaba has also covered the full size range (0.5B to 235B parameters, and a 2.4-trillion-parameter Qwen3.8-Max previewed in July 2026), meaning a developer can find a Qwen that fits their hardware from a single laptop to a data center.

The strategic insight is that China does not need to win every benchmark to win influence. It needs to be the default foundation on which the rest of the world builds. Licenses are the mechanism that makes default adoption legal and frictionless.

DeepSeek’s R1 demonstrated something that changes the sanctions calculus: you do not need the newest chips to build frontier-adjacent models. The lab trained on older Hopper-generation hardware (H800-class), using mixture-of-experts architecture, FP8 precision, and multi-head latent attention to squeeze training compute dramatically. The frequently cited figure — a reported ~$5.6 million for the final training run of V3 — should be treated with care: it refers to the final run only, excludes the extensive research and development that produced the technique, and the figure itself is contested. But the qualitative point stands. R1’s release triggered the largest single-day market-capitalization loss in the history of any US company (Nvidia, ~$589 billion in early 2025) — a market verdict that the efficiency lever was real.

The efficiency lesson compounds: if capability is partly a function of cleverness rather than raw flops, then denying the latest silicon slows an adversary less than intended. The US responded not by relaxing controls but by adding restrictions on the techniques themselves (see Section 12).

DeepSeek’s R1 also made distillation — training a smaller “student” model on the outputs of a larger “teacher” — a geopolitical flashpoint. OpenAI publicly claimed DeepSeek distilled from its o1 models; DeepSeek denied it. The claim is contested and unproven. What is not contested is that open-weight models are built for distillation: open architecture and accessible weights make knowledge transfer straightforward, and the technique is standard practice across the industry, including at US labs.

The policy significance is that distillation converts closed, controlled capability into open, uncontrolled capability. A state that cannot buy the frontier model can sometimes obtain its behavior through an intermediary that can. This is why the July 2026 industry open letter (Section 13) explicitly defended distillation as a legitimate technique that should not be restricted wholesale.

Open weights are effectively un-ban-able at the distribution level. They sit on Hugging Face, ModelScope, GitHub, and countless mirrors. A country that prohibits them finds them re-hosted the next day. The US government has acknowledged this directly: any prohibition broad enough to be effective would sweep in thousands of derivative models created by developers with no connection to Chinese state security, and enforcement would be practically impossible.

This property matters most for the sanctions-circumvention use case. A state under US embargo cannot be cut off from open Chinese models, because the distribution network does not respect the embargo.

The newest front in the control regime is not the chip itself but the cloud it runs in. The January 2026 H200 rule and the House-passed Remote Access Security Act (January 2026) target infrastructure-as-a-service: who can rent compute, and who can access models remotely. The US is effectively extending controls from hardware to the services that deliver hardware capability remotely.

But this mechanism is weaker against open weights, because open-weight models can be run on domestic hardware in the adopting country. Russia does not need US clouds to run Qwen; it runs it on its own infrastructure. This is the essential difference from the closed-model era, where access to OpenAI or Anthropic APIs was the only path and could be switched off.

The distribution of Chinese models is not purely organic. It is actively assisted by state-adjacent channels:

  • Russia: T-Bank (formerly Tinkoff) built its Gen-T model family on Qwen 2.5 and Qwen 3, reporting that basing on Qwen cut model-building costs by 80–90% versus training from scratch. Its T-Pro 2.0 reasoning model was released openly in 2025 as a Russian-language alternative to Western and Chinese models. The Russian AI stack has effectively been scaffolded on Chinese open weights.
  • Africa: Microsoft’s AI Economy Institute found DeepSeek usage in Africa is estimated at 2–4× higher than in other regions, aided by strategic promotion and partnerships with Huawei, which has integrated DeepSeek for African markets.
  • Singapore: The national AI program selected Qwen as its foundation model.
  • China’s domestic base: DeepSeek reportedly reached an estimated 89% market share of generative-AI usage in China; at least 72 local government agencies had integrated localized DeepSeek deployments by early 2025.

Microsoft’s January 2026 AI Economy Institute report measured DeepSeek’s market share across countries. The pattern is unambiguous:

CountryEstimated DeepSeek market share
China89%
Belarus56%
Cuba49%
Russia~43%
Ethiopia / Uganda11–14%

Adoption in North America and Western Europe remained below 5% — not because of capability, but because of restrictions and entrenched alternatives. The report’s own framing is explicit: DeepSeek gained traction “in places where U.S. services face restrictions or where foreign tech access is limited.” The models are, in effect, an alternative import channel that sanctions could not close.

The circumvention framing must be kept precise. DeepSeek and Qwen were not designed as state sanctions-engineering projects; DeepSeek spun out of a private hedge fund, and researchers have noted limited direct state support for its early work. The circumvention outcome is structural, not conspiratorial: when Western capability is denied and Chinese capability is free, open, and multilingual, the sanctioned economy imports the Chinese alternative by default. Stanford HAI’s finding that Chinese labs benefited from “an enabling environment” rather than direct subsidies supports this reading.

Here the story splits into two different kinds of leverage.

For sanctioned states, open Chinese models are a lifeline — the first time they can run frontier-class AI without Western permission.

For China, the same distribution is a soft-power asset. Weights are not neutral artifacts; they carry the norms of their origin. Chinese models must legally comply with Beijing’s content and censorship requirements at the training stage, and those guardrails travel:

  • NewsGuard found leading Chinese systems repeat or fail to correct pro-Chinese false claims 60% of the time in testing.
  • US government testing (CAISI/NIST) found DeepSeek models roughly 12× more susceptible to jailbreaking than comparable US models, and compliant with 94% of overtly malicious jailbroken requests versus 8% for US frontier models. Cisco independently reported a 100% attack success rate against R1.
  • US analysts note Chinese companies operate under China’s National Intelligence Law, which obliges them to “support, assist, and cooperate” with state intelligence work — raising data-security questions for API users who transmit proprietary information to Chinese servers.

This is the uncomfortable asymmetry: the same properties that make the models attractive to sanctions-targeted states — permissive, ungoverned, cheap — also make them risky and normatively loaded. The state adopting the model for autonomy also inherits the model’s training politics.

US policy has swung sharply, and the swings reveal the structural problem:

DateAction
Oct 2022 / Oct 2023Chip export controls to China
Jan 2025AI Diffusion Rule — controls closed weights (ECCN 4E091), tiers the world
May 13, 2025BIS announces rescission; rule never enforced; ECCN 4E091 effectively removed
July 2025America’s AI Action Plan elevates open weights as a US strategic asset; OpenAI releases its first open-weight models since GPT-2
Dec 2025Trump announces H200 exports to China will be allowed, with a 25% revenue share
Jan 2026BIS final rule: case-by-case licensing for H200/MI325X-class chips; 25% tariff on covered chips; Beijing’s customs halts H200 entry
Jul 202625 US tech companies — Nvidia, Microsoft, Meta, Palantir — sign an open letter against “premature restrictions” on open weights; OpenAI and Anthropic decline to sign

The January 2026 H200 rule is the most telling signal. It represents the US selling the chips it once banned, under conditions (50% cap on China-bound shipments, third-party testing, KYC certification) that read as an attempt to tax the leakage rather than stop it. Washington has effectively accepted that it cannot prevent Chinese models from running on some version of advanced silicon — so it has moved to monetize the hardware and control the most advanced tier.

In parallel, reported deliberations in Beijing about restricting overseas access to Qwen’s open weights show that openness is a policy choice China could reverse — the “double-edged sword” of government attention that Stanford HAI highlights, with the 2020 tech crackdown as precedent. If Beijing were to close the tap, the sanctions-circumvention channel would narrow at its source.

The circumvention is real but not total. Three limits matter:

  1. The frontier still runs on chips. DeepSeek’s delayed V4 illustrates the bind: US controls pushed it toward Huawei’s Ascend ecosystem, and the migration forced a full rewrite of its Nvidia-based stack — a cost in time and engineering that controls did impose. Anthropic’s top model still leads the field by about 2.7% as of March 2026 (Stanford AI Index). Open weights redistribute frontier-adjacent capability; they have not yet equalized frontier training.
  2. China may restrict its own models. If Beijing’s deliberations on Qwen’s overseas availability conclude with controls, the source of the open ecosystem narrows — and the states that depended on it lose their alternative import channel.
  3. Adoption is not influence by default. Download counts measure reach, not compliance. The dual-use pattern — states take the capability but also the censorship — means Chinese influence is co-produced, negotiated, and often locally modified, not simply projected.

Right: The most advanced models remain American, the US retains the compute lead, and open-weight diffusion does not automatically translate into geopolitical alignment.

Wrong: The claim that export controls successfully “contained” Chinese AI. By the measures that matter for influence — downloads, derivatives, default adoption in sanctions-targeted and Global South markets — the open-weight channel has outpaced anything the controls were designed to prevent. As the American Action Forum concluded, DeepSeek’s success “revealed that American efforts to sustain leadership in the field of AI have fallen short of their objectives.”

Open-weight models did not defeat export controls in the sense of overturning them. They circumvented them structurally, by moving the locus of capability from a controllable physical input to an uncontrollable digital one. Two consequences follow.

For the United States, the controls’ original premise — deny chips, deny capability — no longer holds alone. The response has shifted, haltingly, from pure denial toward a mixed strategy: sell the previous-generation silicon, tax the leakage, keep the newest tier closed, and try to win the open-ecosystem competition on merit. The July 2026 open letter, signed by most of the US AI establishment, is an admission that the open-weight race is now the arena — and that the US intends to compete there rather than ban its way out.

For sanctioned and Global South states, the immediate effect is a historic levelling of access. A Russian bank, a Cuban researcher, an Ethiopian startup can now run frontier-class AI on open weights, free, in their own language, on their own hardware — options that did not exist two years ago. The cost is that the capability arrives pre-packaged with the norms, guardrails, and legal obligations of its country of origin. Circumvention was never free; it is merely cheaper than the alternative.

Are DeepSeek’s models actually open source? They are open-weight under the MIT license — freely usable and modifiable — though the training data is not fully disclosed. “Open weights” and “open source” are often conflated; the weights are open, the full data pipeline is not.

Did DeepSeek break US law? The reported $5.6M training figure refers to a final run on chips that were legal to possess at the time of training. Contested allegations include distillation from OpenAI’s o1 (denied by DeepSeek, unproven) and, in February 2026, an anonymous official’s claim that DeepSeek trained a next-generation model on smuggled Blackwell chips — unconfirmed and denied by Beijing.

If Chinese models are easier to jailbreak, isn’t that a US security concern? Yes, and it is a documented one. CAISI and Cisco testing found DeepSeek substantially more vulnerable to jailbreaks than US frontier models, which is why several US agencies and states restrict their use on official systems.

Could Beijing shut this down? Reported deliberations about restricting overseas access to Qwen’s open weights suggest the possibility. Government attention is a double-edged sword for Chinese labs, and the 2020 tech crackdown is the cautionary precedent.

What should a business do? Treat open Chinese models as high-capability but politically loaded inputs. Test them in sandboxes, evaluate against task-specific benchmarks, implement a model-abstraction layer so you can switch providers, and involve legal counsel before deployment in regulated industries.

  1. Assume distribution cannot be banned. Resources spent on prohibition should move toward supply-side competition: open US models that are better, cheaper, and more multilingual.
  2. Regulate the risky surfaces, not the weights. Target the documented harms — jailbreak insecurity, data transmission to intelligence-law jurisdictions, malicious misuse — via safety standards, procurement rules, and liability, as the July 2026 industry letter itself recommends.
  3. For sanctioned states: build sovereignty buffers — domestic fine-tuning capacity, hardware independence, and license diversification — so that dependence on any single open-weight source is a choice, not a necessity.
  4. Monitor the two closures: whether the US tightens closed-frontier controls further, and whether Beijing restricts Qwen’s overseas release. Either event changes the calculus within a quarter.
  5. Measure influence by adoption, not announcements. Track derivative counts, token volume, and enterprise deployment — the metrics that decided this race.

Key sources: ATOM Report (arXiv 2604.07190); Stanford HAI “Beyond DeepSeek” and 2026 AI Index; Microsoft AI Economy Institute (Jan 2026); BIS press releases (May 2025, Jan 2026); the July 24, 2026 industry open letter; CNBC; War on the Rocks; CAISI/NIST evaluations; Interconnects AI / Hugging Face data; SCMP; American Action Forum.

You May Also Like