The short answer

On the providers' published list prices used in this comparison, DeepSeek V4 Flash produces the lowest bill for the reference workload. Qwen offers a broader ladder of Plus models and multimodal capabilities, while its Singapore prices differ from prices shown for other regions.

Do not mix regions.

Alibaba Cloud publishes separate Singapore, China, Germany, US, Japan, and Hong Kong tables. This article uses the Singapore international table for Qwen and DeepSeek's first-party API for DeepSeek.

Verified list prices

ModelInput / 1MOutput / 1MScope
DeepSeek V4 Flashdeepseek-v4-flash$0.14$0.28First-party API
DeepSeek V4 Prodeepseek-v4-pro$0.435$0.87First-party API
Qwen 3.7 Plusqwen3.7-plus · ≤256K tier$0.40$1.60Singapore international
Qwen 3.6 Plusqwen3.6-plus · ≤256K tier$0.50$3.00Singapore international
Qwen 3.5 Plusqwen3.5-plus · ≤256K tier$0.40$2.40Singapore international
Qwen Plusqwen-plus · non-thinking output$0.40$1.20Singapore international

DeepSeek also publishes cache-hit input prices of $0.0028/M for Flash and $0.003625/M for Pro. The table above uses cache-miss input prices so every model starts from a conservative, comparable assumption.

Monthly cost at 10M input and 2M output

The calculation is input volume multiplied by input price, plus output volume multiplied by output price.

ModelInput costOutput costTotal
DeepSeek V4 Flash$1.40$0.56$1.96
DeepSeek V4 Pro$4.35$1.74$6.09
Qwen Plus$4.00$2.40$6.40
Qwen 3.7 Plus$4.00$3.20$7.20
Qwen 3.5 Plus$4.00$4.80$8.80
Qwen 3.6 Plus$5.00$6.00$11.00

What changes the answer

Context tiers

Qwen pricing increases when a request crosses the published context threshold. A monthly average cannot reveal this: one very long request may enter a higher tier even when total monthly usage is modest.

Caching

Repeated system prompts and stable prefixes can lower effective input cost. Treat caching as a measured optimization, not an assumed discount: verify your provider dashboard shows cache hits before budgeting around them.

Capability requirements

The cheapest row is not automatically the right model. Vision input, tool use, latency, rate limits, and the quality of your own evaluation set should decide which rows remain eligible before price breaks the tie.

Decision rule

Start with the least expensive model that passes your real evaluation set. Keep the region, context tier, thinking mode, and cache assumptions beside every cost number. If any of those labels disappear, the comparison is no longer auditable.

Next / 02 Cheapest LLM APIs for coding