Will Nvidia’s CUDA Moat Crack in a Year or So?
Dear Investor, Welcome to Deep Research Global.
A partial transcript of a private May 2026 investor call attributed to DeepSeek founder Liang Wenfeng leaked in late July, then vanished from Chinese platforms within hours.
The claim getting the most airtime on Wall Street trading desks was blunt. He allegedly told the room that Nvidia’s CUDA moat is “rapidly disintegrating,” and that Chinese labs could pair Huawei silicon with AI-driven code generation and tools like TileLang to bypass Nvidia’s software lock-in inside roughly a year.
For the US investor watching an $81.6 billion quarterly print at Nvidia and wondering how durable it is, that comment is either a hollow flex from a company that often delays its own model launch, or the loudest early warning bell so far.
Either way it deserves a careful analysis.
Recommended - Read Full Reports
Read All Reports
Disclaimer: This analysis is for informational & educational purposes only and should not be construed as investment advice. Investors should conduct their own due diligence before making investment decisions. Past performance does not guarantee future results.
What Liang Reportedly Said
The transcript comes from a nearly four-hour closed-door meeting with new shareholders, and multiple reconstructions have surfaced.
The reason it matters is the pace of progress. DeepSeek had earlier disclosed the UE8M0 FP8 numeric format in V3.1, a data format explicitly co-designed with China’s “next-generation domestically produced chips.”
Liang’s own framing was that the US-China gap now sits mostly in compute access, not in algorithms or talent. Which is fine. It’s also, notably, the one gap the CHIPS Act and export controls were specifically designed to widen.
LIANG'S CORE CLAIMS (per the leaked notes)
- 4 x Huawei Ascend 950 ≈ 1 x Nvidia GB300
- Paying 200% more for Ascend clusters "doesn't matter"
- DeepSeek replaced CUDA internally during V3 training
- TileLang + AI-written kernels shrink the CUDA moatWhy CUDA Was A Moat In The First Place
CUDA is not one thing. It’s a parallel programming model, a set of compilers, a decade of library work (cuDNN, NCCL, TensorRT), and a global developer base that already knows the API.
Nvidia has been reinvesting into that stack since 2006, longer than most of today’s AI startups have existed.
The practical result is that most AI code paths in production quietly assume CUDA underneath.
Even PyTorch, technically hardware-agnostic, has its most optimized kernels written for Nvidia hardware first and everyone else later. Migrating a serious training workload to a new stack used to take a full engineering year.
The catch is that “used to.”
Compiler-driven approaches like OpenAI’s Triton already let developers write GPU kernels in Python without touching low-level CUDA. And Huawei’s answer, CANN, has been open-sourced exactly to shorten the porting distance.
THE STACK, SIMPLIFIED
NVIDIA: Model -> PyTorch -> CUDA/cuDNN -> GPU
HUAWEI: Model -> MindSpore/PyTorch -> CANN -> Ascend NPU
DEEPSEEK: Model -> In-house compiler + TileLang -> Any accelerator
Notice the third row. It’s the interesting one, because it treats the underlying silicon as swappable.
If a model house can compile down to whatever chip it happens to have, the moat is no longer the software layer, it’s just supply.
The Huawei Side Of The Trade
Huawei laid out a rare three-year roadmap at its September 2025 Connect event.
The Ascend 950PR shipped in Q1 2026, followed by the 950DT, then the 960 planned in 2027 and the 970 in 2028, with each generation roughly doubling compute per chip.
Huawei’s headline system, the CloudMatrix 384, is worth understanding on its own. It packs 384 Ascend 910C chips into a single supernode and, at the rack level, delivers roughly 70% more FLOPS than Nvidia’s flagship GB200 NVL72.
The trade-off is power.
A single CM384 pulls about 560 kW against roughly 145 kW for the NVL72, so on a performance-per-watt basis Nvidia is still the cleaner design.
But, in China, where power is heavily subsidized in strategic corridors and grid buildout is state-directed, that trade-off is more tolerable than it would be in US or Europe.
CLOUDMATRIX 384 vs GB200 NVL72 (system level)
Aggregate FLOPS: CM384 ~1.7x vs NVL72 baseline
Memory capacity: CM384 ~3.6x vs NVL72
Power draw: CM384 ~560 kW vs NVL72 ~145 kW
Efficiency: Nvidia wins per watt by a wide margin
Huawei has also revealed that its next roadmap generation will move to in-house HBM, which is arguably the harder problem than the logic die itself.
Whether SK hynix and Samsung continue to sell HBM into China at scale is the real question, and it’s more of a US State Department question than a technical one.
DeepSeek’s Play, And Why It’s Different
Most of the “China is catching up” coverage focuses on the hardware. The leaked call flips that.
Keep reading with a 7-day free trial
Subscribe to Deep Research Global to keep reading this post and get 7 days of free access to the full post archives.




