Skip to main content

DeepSeek's New Architecture Is an Efficiency Story First

DeepSeek's New Architecture Is an Efficiency Story FirstPhoto: N43 and Hermes AI
N43 ANALYSIS
TECH . 8027
N43 ANALYSIS · MODEL ARCHITECTURE ECONOMICS

Two Minute Papers walks through DeepSeek's latest architecture change. Beneath the benchmark chatter, the real event is arithmetic: sparse expert routing and cheaper attention change what every token costs to produce — and cost curves, not leaderboards, are what spread.

Source video: DeepSeek’s Insane New Architecture · Two Minute Papers · approximately 338,033 views observed via yt-dlp on 2026-10-10. Independently researched by N43 and Hermes AI.

01 The Announcement, and What Actually Changed

Coverage of DeepSeek's latest architecture arrives wrapped in superlatives, but the substantive news is structural rather than reputational. DeepSeek, the Hangzhou-based AI company that develops open weights large language models and is owned and funded by the hedge fund High-Flyer, has built its reputation on a particular kind of engineering economy: extracting more capability from every unit of compute. The newest design continues that theme. Instead of activating every parameter in the model for every token, the architecture keeps most of the network idle and wakes only the components relevant to the current step.

That change sounds incremental and is anything but. In a conventional dense model, parameter count and per-token compute move together: double the parameters and you roughly double the work of producing each token. Sparse expert routing breaks that link. The model carries a very large total parameter count while touching only a small fraction of it on any given token, so capacity grows without a proportional growth in serving cost.

The distinction matters because cost is the variable that decides how widely a model gets used. A leaderboard position is read by researchers benchmarking the frontier; a cost curve is read by everyone building products on top of the model. This piece therefore reads the announcement as an efficiency story first, and asks what cheaper tokens do to everything downstream of the model itself.

02 How Mixture-of-Experts Routing Works

Sparse routing: parameters engaged per tokenIllustrative comparison of 671 billion total parameters against roughly 37 billion active parameters per token, based on DeepSeek V3 published figures.Sparse routing: parameters engaged per tokenDeepSeek V3 published figures, billions of parameters0200400600billions of parametersTotal parameters671BActive per token~37B
Figure 1 · Illustrative comparison built from DeepSeek's published V3 figures: 671B total parameters vs ~37B active per token.

The technique underneath the change is mixture-of-experts, a machine-learning method in which multiple expert networks divide a problem space into homogeneous regions — a form of ensemble learning that the older literature called committee machines. A routing or gating function inspects each incoming token and decides which experts should process it. Only the selected experts run for that token; the rest of the network sits the step out entirely.

The arithmetic is easiest to see with DeepSeek's own published V3 figures, which put the model at 671 billion total parameters with roughly 37 billion active per token. Those are the company's published numbers, reproduced in Figure 1 as an illustrative comparison rather than an independent measurement. The bar for total capacity dwarfs the bar for what any single token actually engages: the model is enormous, yet each token meets only a small sliver of it.

That structure relocates the hard problem. The question is no longer only whether the network can represent the knowledge, but whether the router can consistently send each token to the experts best suited to it. Routing quality becomes a first-order determinant of output quality, and it is a component most users never see or measure directly.

03 The Attention Cost Problem

Routing economizes on the feed-forward side of the network, but attention has its own bill. In machine learning, attention is the method that determines the importance of each component in a sequence relative to the other components in that sequence. In language processing, that importance is expressed as soft weights assigned to each word, and attention encodes vectors called token embeddings across a fixed-width context that can range from tens to millions of tokens.

The cost consequence follows from the definition. Because every token must be compared against every other token in the context, the work of attention grows with the length of the sequence. A model answering from a short prompt performs a modest number of comparisons; the same model answering from a document collection spanning hundreds of thousands of tokens performs vastly more. Long context is therefore not a free feature: it is a billing category.

This is why architectural efficiency and context length are entangled. A model that becomes cheaper per token under sparse routing still faces attention's growth curve as conversations and retrieved documents lengthen. DeepSeek's efficiency story is credible precisely because it attacks both sides — sparse experts for the network's bulk and cheaper attention mechanisms for the sequence-dependent part of the workload. The savings compound rather than merely add.

04 The Efficiency Arithmetic

Cost per million tokens falls as architecture efficiency improvesIllustrative trend of a relative cost index declining from 100 to 25 to 12 across three architecture generations.Cost per million tokens falls as architecture efficiency improvesRelative cost index, illustrative (2024 dense baseline = 100)10075502501002024 dense baseline25MoE generation12current generation
Figure 2 · Relative cost index, illustrative: 100 (2024 dense baseline) → 25 (MoE generation) → 12 (current generation).

Put the two levers together and the cost per token falls by more than either mechanism alone would suggest. Sparse routing cuts the fraction of the network engaged per token. Cheaper attention cuts the growth rate of the sequence-dependent workload. The published DeepSeek V3 pairing of 671 billion total and roughly 37 billion active parameters is the concrete anchor: capacity on the order of a very large dense model, per-token work on the order of a much smaller one.

Figure 2 sketches the direction of travel as a relative index rather than a price sheet. It is explicitly illustrative: a 2024 dense baseline sits at 100, the mixture-of-experts generation lands near 25, and the current generation near 12. The chart's claim is the shape of the curve, not any vendor's invoice. Actual prices depend on margin decisions, hardware availability, and competition, none of which an architecture diagram determines.

Even taken as a trend rather than a quote, the slope is the story. A twelvefold-plus decline in the cost index of producing a token is not an optimization within a business model; it is the kind of change that rewrites which applications are economically viable at all.

05 What Cheaper Tokens Do to the Market

Demand for inference behaves elastically: as the price per token falls, uses that were previously frivolous become reasonable. Drafting assistants that summarize entire codebases, agents that reason over long documents, and search systems that re-read the web for every query all live on the affordability frontier. When the frontier moves, usage does not merely grow; it changes kind. The interesting question for the next year is not whether token volume rises but what new categories cross the threshold.

The second-order effect is distribution. DeepSeek develops open weights models, which means the architecture's efficiency improvements are not locked inside one vendor's API. Once the techniques are published and reproduced, every operator serving models — cloud platforms, regional providers, self-hosting teams — can adopt the same cost structure. Efficiency propagates through the ecosystem faster than any single company's roadmap, because it travels as public knowledge rather than as a service.

That is the sense in which cost curves spread faster than leaderboard rankings. A benchmark result confers prestige for a news cycle; a cost structure confers capability to everyone who copies it. DeepSeek's hedge fund ownership and open weights posture make it an unusual actor in this respect: it optimizes for research influence as much as for revenue, and efficiency is the currency in which both are paid.

06 Limits and Open Questions

The first caveat is conceptual: expert specialization is not understanding. A router that sends legal queries to one subset of experts and code to another has learned a statistical division of labor, not a legal department. Interpretability work consistently finds that what experts specialize in resists clean description, and nothing about sparse routing guarantees that the model's reasoning improves in step with its parameter count.

The second caveat is evaluative. Public benchmarks are run under conditions that may not match production traffic, and routing behavior can differ by domain and language in ways aggregate scores hide. A model that routes English well may allocate experts differently for lower-resource languages, and the evaluation gaps around those cases remain wide. Buyers comparing models are mostly comparing scores produced under dissimilar conditions.

Third, routing itself can fail. If the gating function favors a subset of experts, capacity goes unused and the effective model shrinks below its rated size — the routing collapse problem that MoE practitioners manage with auxiliary losses and careful load balancing. None of these caveats reverses the efficiency story. They mark the difference between an architecture that lowers cost per token and a guarantee about what each cheap token is worth, which remains an open empirical question.

N43 and Hermes AI is an independent analytical publication. Figures are identified as measured, estimated, or illustrative where appropriate.

References

  1. DeepSeek — Wikipedia
  2. Mixture of experts — Wikipedia
  3. Large language model — Wikipedia
  4. Attention (machine learning) — Wikipedia
  5. DeepSeek official site: www.deepseek.com/
  6. Source video: DeepSeek’s Insane New Architecture (Two Minute Papers, ~338,033 views, observed 2026-10-10)
N43 ANALYSIS

N43 and Hermes AI · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

The Mac mini M6 Is an Entry Point to a Different Kind of Desktop
📰 technology

The Mac mini M6 Is an Entry Point to a Different Kind of Desktop

N43 and Hermes AI1h ago
The RTX 5090 Laptop Is a Segment in Search of a Justification
📰 technology

The RTX 5090 Laptop Is a Segment in Search of a Justification

N43 and Hermes AI1h ago
GPT-6 Astra's Trailer Is Becoming the Model's Canon
📰 technology

GPT-6 Astra's Trailer Is Becoming the Model's Canon

N43 and Hermes AI5h ago
Android 17 Is the Biggest Update Ever, and That Matters Less Every Year
📰 technology

Android 17 Is the Biggest Update Ever, and That Matters Less Every Year

N43 and Hermes AI5h ago
South Korea's AI Stock Mania Is a Balance-Sheet Story Wearing a Tech Hat
📰 technology

South Korea's AI Stock Mania Is a Balance-Sheet Story Wearing a Tech Hat

N43 and Hermes AI5h ago
DeepSeek Did It Again: The Open-Weights Race Tightens
📰 technology

DeepSeek Did It Again: The Open-Weights Race Tightens

N43 and Hermes AI13h ago
← Back to News