DeepSeek V4.1 Flash Slashes AI Agent Costs: Cache Memory Cut to a Quarter

The short version

On 10 September 2026, DeepSeek released V4.1 Flash — and then did something unusual two days later: it published a note saying it had changed its mind about retiring its flagship.

Three things about this release matter to anyone in Malaysia paying for AI, and only one of them is about the model being smarter.

1. The memory cut — and why it decides what an agent costs

V4.1 Flash is a 552-billion-parameter mixture-of-experts model built on a new causal encoder–decoder architecture. Only 8 billion parameters are active for input and 16 billion for output — a deliberately asymmetric design. It also adds native visual understanding to the mainline model, where vision had previously been a separate API.

Those are specification details. The commercially interesting number is this one.

DeepSeek reports that V4.1 Flash’s key-value cache now needs one quarter of the HBM memory and one eighth of the SSD storage compared with the previous generation.

That sounds like an infrastructure footnote. It is not, and DeepSeek says so plainly: “Cache-hit charges often account for a large share of agent costs.”

The reason is structural. A chatbot answers a question and forgets it. An agent — something that runs a multi-step task, calls tools, and keeps working for hundreds or thousands of turns — carries its accumulated context with it. Every turn re-reads that context from cache. On a long-running job, cache reads can outweigh the actual thinking. Cut the cache to a quarter and you have attacked the largest line on the bill, not the smallest.

DeepSeek announcement page for DeepSeek-V4.1-Flash
DeepSeek release note for V4.1 Flash, 10 September 2026. Source: deepseek.com

2. The price, and the hours that halve it

DeepSeek runs peak and off-peak pricing, and the off-peak rate is exactly half the peak rate. New pricing took effect at 04:00 UTC on 10 September 2026.

Per 1M tokens V4.1 Flash
off-peak
V4.1 Flash
peak
V4-Pro
off-peak
V4-Pro
peak
Input — cache hit US$0.003 US$0.006 US$0.022 US$0.044
Input — cache miss US$0.15 US$0.30 US$0.66 US$1.32
Output US$0.60 US$1.20 US$1.98 US$3.96
DeepSeek API pricing page showing V4.1 Flash and V4-Pro rates per million tokens
DeepSeek published rate card, including the peak and off-peak split. Source: DeepSeek API docs.

Context window is 1 million tokens, maximum output 384K, and the concurrency limit is 2,500 against V4-Pro’s 500.

When “off-peak” actually is, in Malaysian time

This is the part almost no coverage translates. DeepSeek defines peak hours as 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday. Everything else is off-peak. Converted to Malaysian time (UTC+8):

Malaysian time Rate
09:00 – 12:00 (Mon–Fri) peak — full price
12:00 – 14:00 off-peak — half price
14:00 – 18:00 (Mon–Fri) peak — full price
18:00 – 09:00 next day off-peak — half price
All day Saturday and Sunday off-peak — half price

In practical terms: a Malaysian business that batches its AI work overnight, over lunch, or at the weekend pays half of what it pays at 10am on a Tuesday. Document processing, bulk content generation, nightly reporting, data extraction, test runs — none of those are time-critical, and all of them are 50% cheaper pushed outside the two weekday windows.

3. The U-turn on V4-Pro

The original release said V4.1 Flash had beaten V4-Pro “on performance, cost, speed and total runtime”, that V4-Pro would be phased out, and that from 04:00 UTC on 14 September all deepseek-v4-pro requests would be routed to V4.1 Flash at V4.1 Flash rates.

Then DeepSeek’s own pricing page changed. It now reads:

“In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes.”

So V4-Pro is not being retired — it stays, at its old prices. The two older names, deepseek-v4-flash and deepseek-v4-flash-vision-exp, are still accepted but now route to V4.1 Flash and bill at Flash rates. The model name to use is deepseek-flash.

Legacy integrations do not break. But anyone who planned a migration around a forced cutover should know it is no longer happening on that schedule.

Why this lands in Malaysia specifically

DeepSeek named its official partners in the release: WorkBuddy (including CodeBuddy) and OpenCode are both fully supporting V4.1 Flash. For a Malaysian business already running WorkBuddy — and there are a growing number, given the Tencent Cloud partnership activity in this market through 2026 — the model underneath improved and got cheaper without a migration step.

The wider point is the direction of travel. The cost of running an agent is now falling faster than the cost of the intelligence inside it. Cache memory cut to a quarter; off-peak rates at half; concurrency limits up fivefold. For a Malaysian SME deciding whether an always-on AI workflow is affordable, those three numbers change the answer more than any benchmark score.

ClinePass is the flat-rate alternative if you would rather not meter tokens at all: a US$9.99/month subscription that bundles DeepSeek V4 Pro and V4 Flash alongside GLM, Kimi, MiniMax, MiMo and Qwen models inside the Cline agent, with quotas Cline describes as two to five times standard API rate limits. See ClinePass →

What we would do with this

  • Move anything schedulable to off-peak. 18:00–09:00 and 12:00–14:00 on weekdays, and all weekend. Same work, half the bill.
  • Audit your cache-hit share. If you are running agents, cache reads are likely your biggest cost line — and it is the one this release attacked.
  • Do not plan a forced migration. V4-Pro survives. Test V4.1 Flash against your own workload rather than assuming the newer model wins on your task.

Sources

Prices are US dollars per 1 million tokens as published by DeepSeek. Peak hours are stated by DeepSeek in UTC (01:00–04:00 and 06:00–10:00, Monday–Friday); Malaysian times above are converted at UTC+8. Published rates can change — check DeepSeek’s pricing page before committing a budget.

Cheaper per token is not the point. Cheaper per agent run is.

The memory cut matters more to a business than the headline price. A lower rate per token is pleasant; a lower cost per agent run is what decides whether an AI workflow is worth running at all. If an agent reads your documents, calls your tools and runs unattended, the number that governs the bill is the cache – and V4.1 Flash has just cut it to a quarter.

Big Domain is a Tencent official agentic cloud partner. We build and host agentic AI for Malaysian businesses, from the stack underneath to the agents on top.

See how Big Domain builds agentic AI →
Want to know what it would cost on your own workflows? Talk to Big Domain →

Message us on WhatsApp →

Big Domain is a one-stop technology partner for Malaysian businesses – hosting, domains, SEO, app development and agentic AI. bigdomain.my

Some links in this article are affiliate links. If you sign up through one of them, Big Domain may earn a commission, at no extra cost to you. It does not influence what we cover, what we recommend, or where we rank anything.