Deep Research Global

Deep Research Global

Will Nvidia’s CUDA Moat Crack in a Year or So?

Deep Research Global's avatar
Deep Research Global
Jul 28, 2026
∙ Paid

Dear Investor, Welcome to Deep Research Global.


A partial transcript of a private May 2026 investor call attributed to DeepSeek founder Liang Wenfeng leaked in late July, then vanished from Chinese platforms within hours.

The claim getting the most airtime on Wall Street trading desks was blunt. He allegedly told the room that Nvidia’s CUDA moat is “rapidly disintegrating,” and that Chinese labs could pair Huawei silicon with AI-driven code generation and tools like TileLang to bypass Nvidia’s software lock-in inside roughly a year.

For the US investor watching an $81.6 billion quarterly print at Nvidia and wondering how durable it is, that comment is either a hollow flex from a company that often delays its own model launch, or the loudest early warning bell so far.

Either way it deserves a careful analysis.

Get Stock / Company Analysis Reports & Investment Insights Direct to Your Inbox. Read by 6,000+ Investors & VCs. Don’t Miss Out.


Recommended - Read Full Reports

DeepSeek - Fundamental Analysis Report 2026 (Updated)

DeepSeek - Fundamental Analysis Report 2026 (Updated)

Deep Research Global
·
Jul 26
Read full story
Nvidia (NVDA) - Fundamental Analysis Report 2026 (Updated)

Nvidia (NVDA) - Fundamental Analysis Report 2026 (Updated)

Deep Research Global
·
May 26
Read full story

Read All Reports


Disclaimer: This analysis is for informational & educational purposes only and should not be construed as investment advice. Investors should conduct their own due diligence before making investment decisions. Past performance does not guarantee future results.


What Liang Reportedly Said

The transcript comes from a nearly four-hour closed-door meeting with new shareholders, and multiple reconstructions have surfaced.

The reason it matters is the pace of progress. DeepSeek had earlier disclosed the UE8M0 FP8 numeric format in V3.1, a data format explicitly co-designed with China’s “next-generation domestically produced chips.”

Liang’s own framing was that the US-China gap now sits mostly in compute access, not in algorithms or talent. Which is fine. It’s also, notably, the one gap the CHIPS Act and export controls were specifically designed to widen.

LIANG'S CORE CLAIMS (per the leaked notes)
- 4 x Huawei Ascend 950 ≈ 1 x Nvidia GB300
- Paying 200% more for Ascend clusters "doesn't matter"
- DeepSeek replaced CUDA internally during V3 training
- TileLang + AI-written kernels shrink the CUDA moat

Why CUDA Was A Moat In The First Place

CUDA is not one thing. It’s a parallel programming model, a set of compilers, a decade of library work (cuDNN, NCCL, TensorRT), and a global developer base that already knows the API.

Nvidia has been reinvesting into that stack since 2006, longer than most of today’s AI startups have existed.

Nvidia CUDA platform
Image source: NVIDIA Developer

The practical result is that most AI code paths in production quietly assume CUDA underneath.

Even PyTorch, technically hardware-agnostic, has its most optimized kernels written for Nvidia hardware first and everyone else later. Migrating a serious training workload to a new stack used to take a full engineering year.

The catch is that “used to.”

Compiler-driven approaches like OpenAI’s Triton already let developers write GPU kernels in Python without touching low-level CUDA. And Huawei’s answer, CANN, has been open-sourced exactly to shorten the porting distance.

THE STACK, SIMPLIFIED

NVIDIA:   Model  ->  PyTorch  ->  CUDA/cuDNN  ->  GPU
HUAWEI:   Model  ->  MindSpore/PyTorch  ->  CANN  ->  Ascend NPU
DEEPSEEK: Model  ->  In-house compiler + TileLang  ->  Any accelerator

Notice the third row. It’s the interesting one, because it treats the underlying silicon as swappable.

If a model house can compile down to whatever chip it happens to have, the moat is no longer the software layer, it’s just supply.

The Huawei Side Of The Trade

Huawei laid out a rare three-year roadmap at its September 2025 Connect event.

The Ascend 950PR shipped in Q1 2026, followed by the 950DT, then the 960 planned in 2027 and the 970 in 2028, with each generation roughly doubling compute per chip.

Huawei CloudMatrix 384 AI cluster
Image source: Huawei Central

Huawei’s headline system, the CloudMatrix 384, is worth understanding on its own. It packs 384 Ascend 910C chips into a single supernode and, at the rack level, delivers roughly 70% more FLOPS than Nvidia’s flagship GB200 NVL72.

The trade-off is power.

A single CM384 pulls about 560 kW against roughly 145 kW for the NVL72, so on a performance-per-watt basis Nvidia is still the cleaner design.

But, in China, where power is heavily subsidized in strategic corridors and grid buildout is state-directed, that trade-off is more tolerable than it would be in US or Europe.

CLOUDMATRIX 384 vs GB200 NVL72 (system level)

Aggregate FLOPS: CM384 ~1.7x  vs  NVL72 baseline
Memory capacity: CM384 ~3.6x  vs  NVL72
Power draw:      CM384 ~560 kW vs NVL72 ~145 kW
Efficiency:      Nvidia wins per watt by a wide margin

Huawei has also revealed that its next roadmap generation will move to in-house HBM, which is arguably the harder problem than the logic die itself.

Whether SK hynix and Samsung continue to sell HBM into China at scale is the real question, and it’s more of a US State Department question than a technical one.

DeepSeek’s Play, And Why It’s Different

Most of the “China is catching up” coverage focuses on the hardware. The leaked call flips that.

Keep reading with a 7-day free trial

Subscribe to Deep Research Global to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Deep Research Global · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture