The short answer
On the providers' published list prices used in this comparison, DeepSeek V4 Flash produces the lowest bill for the reference workload. Qwen offers a broader ladder of Plus models and multimodal capabilities, while its Singapore prices differ from prices shown for other regions.
Alibaba Cloud publishes separate Singapore, China, Germany, US, Japan, and Hong Kong tables. This article uses the Singapore international table for Qwen and DeepSeek's first-party API for DeepSeek.
Verified list prices
| Model | Input / 1M | Output / 1M | Scope |
|---|---|---|---|
| DeepSeek V4 Flashdeepseek-v4-flash | $0.14 | $0.28 | First-party API |
| DeepSeek V4 Prodeepseek-v4-pro | $0.435 | $0.87 | First-party API |
| Qwen 3.7 Plusqwen3.7-plus · ≤256K tier | $0.40 | $1.60 | Singapore international |
| Qwen 3.6 Plusqwen3.6-plus · ≤256K tier | $0.50 | $3.00 | Singapore international |
| Qwen 3.5 Plusqwen3.5-plus · ≤256K tier | $0.40 | $2.40 | Singapore international |
| Qwen Plusqwen-plus · non-thinking output | $0.40 | $1.20 | Singapore international |
DeepSeek also publishes cache-hit input prices of $0.0028/M for Flash and $0.003625/M for Pro. The table above uses cache-miss input prices so every model starts from a conservative, comparable assumption.
Monthly cost at 10M input and 2M output
The calculation is input volume multiplied by input price, plus output volume multiplied by output price.
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| DeepSeek V4 Flash | $1.40 | $0.56 | $1.96 |
| DeepSeek V4 Pro | $4.35 | $1.74 | $6.09 |
| Qwen Plus | $4.00 | $2.40 | $6.40 |
| Qwen 3.7 Plus | $4.00 | $3.20 | $7.20 |
| Qwen 3.5 Plus | $4.00 | $4.80 | $8.80 |
| Qwen 3.6 Plus | $5.00 | $6.00 | $11.00 |
What changes the answer
Context tiers
Qwen pricing increases when a request crosses the published context threshold. A monthly average cannot reveal this: one very long request may enter a higher tier even when total monthly usage is modest.
Caching
Repeated system prompts and stable prefixes can lower effective input cost. Treat caching as a measured optimization, not an assumed discount: verify your provider dashboard shows cache hits before budgeting around them.
Capability requirements
The cheapest row is not automatically the right model. Vision input, tool use, latency, rate limits, and the quality of your own evaluation set should decide which rows remain eligible before price breaks the tie.
Decision rule
Start with the least expensive model that passes your real evaluation set. Keep the region, context tier, thinking mode, and cache assumptions beside every cost number. If any of those labels disappear, the comparison is no longer auditable.
Next / 02 Cheapest LLM APIs for coding →