Qwen3.8-Max vs Kimi K3

Quick Answer

Qwen3.8-Max and Kimi K3 are both very large, very new Mixture-of-Experts models from Chinese AI labs, positioned as frontier-capability alternatives to closed models. Kimi K3 is already open-weight and larger by total parameter count (2.8T vs 2.4T). Qwen3.8-Max is API-only for now, with open weights committed for the week of August 10, 2026. Neither company has published independent, head-to-head benchmark results against the other, so treat any capability claim from either vendor as a starting point, not a verdict.

At a Glance

Qwen3.8-MaxKimi K3
CompanyAlibaba CloudMoonshot AI
Total parameters2.4 trillion2.8 trillion
Active parametersNot disclosedNot disclosed
ArchitectureSparse Mixture-of-ExpertsMixture-of-Experts with hybrid linear attention (Kimi Delta Attention + Attention Residuals)
Context window1M tokens1M tokens
ModalitiesText, image, video, documentsText, native vision
ReleasedPreview July 19, 2026; GA August 3, 2026July 17, 2026 (full weights by July 27)
Open weightsNot yet; committed for week of August 10, 2026Yes, released
API pricing$2.00 input / $6.00 output per million tokensUsage-based via Kimi.com or Moonshot Open Platform; check current rates
Self-hostingNot possible yetPossible, needs serious multi-GPU hardware
Independent benchmarksNot yet publishedNot yet published

What Is Qwen3.8-Max?

Qwen3.8-Max is Alibaba’s current flagship model, a 2.4-trillion-parameter sparse Mixture-of-Experts system with a 1M-token context window and multimodal input (text, image, video, documents). It reached general availability on August 3, 2026, with published pricing and a benchmark table, though its active-parameter count and a full model card remain undisclosed at the time of writing.

What Is Kimi K3?

Kimi K3 is Moonshot AI’s open-weight frontier model, released July 17, 2026 with the full weight rollout completed by July 27, 2026. At 2.8 trillion total parameters, it’s reported to be the largest open-weight model released to date, built on a hybrid linear attention architecture with native vision support and a 1M-token context window.

Architecture and Scale

Both are enormous sparse Mixture-of-Experts models where only a fraction of total parameters activate per token, but neither company has disclosed the active-parameter count, which makes a true efficiency comparison impossible right now. This matters because total parameters don’t determine intelligence or cost on their own; active parameters are what actually drive inference speed and API pricing. Kimi K3’s architecture also differs structurally, using a hybrid linear attention design (Kimi Delta Attention plus Attention Residuals) rather than the more standard transformer attention Qwen3.8-Max is presumed to use, though Alibaba hasn’t detailed its exact attention mechanism publicly.

Open Weights: Available Now vs. Not Yet Released

This is the clearest practical difference between the two today. Kimi K3’s weights are already out, so if self-hosting or fine-tuning matters to you, it’s usable right now, assuming you have the multi-GPU infrastructure a 2.8-trillion-parameter model demands. Qwen3.8-Max is API-only as of this writing; Alibaba has committed to an open-weight release the week of August 10, 2026, alongside a smaller, far more practically self-hostable Qwen3.8-27B checkpoint, but that release hadn’t happened at the time this page was last updated. If you need open weights today, Kimi K3 is the only one of the two that currently delivers them.

Benchmarks: Both Vendor-Reported, Neither Independently Verified

Alibaba’s GA launch included a benchmark table with mixed results, strong on Terminal-Bench 2.1 and PaperBench, weaker on SWE-bench Pro. Moonshot reports Kimi K3 outperforming most rivals except Claude Fable 5 and GPT-5.6 on general capability, and beating Claude Opus 4.8 and GPT-5.5 on coding and agent benchmarks specifically. Neither set of numbers has been reproduced by an independent evaluator like Artificial Analysis or LMArena as of this writing, and there’s no direct Qwen3.8-Max-versus-Kimi-K3 benchmark from either company. Treat both as marketing until third-party verification catches up, and test on your own workload before making a decision either way.

Pricing and Access

Qwen3.8-Max has published, straightforward API pricing: $2.00 per million input tokens, $6.00 per million output tokens through Alibaba Cloud’s Model Studio. Kimi K3 is accessible through Kimi.com for casual use or the Moonshot AI Open Platform API for developers, with usage-based pricing that should be checked directly since specific per-token rates weren’t fully broken out in Moonshot’s own materials at launch.

Which Should You Choose?

Choose Kimi K3 if open weights and self-hosting availability matter to you right now, or if you want the larger of the two models by total parameter count. Choose Qwen3.8-Max if you’re comfortable with API-only access today and want to wait for Alibaba’s committed open-weight release, or if its specific multimodal input range (image, video, documents) fits your use case better than Kimi K3’s text-and-vision focus. For most production decisions, the honest answer is to test both against your actual tasks, since neither vendor’s benchmark claims have been independently confirmed against the other yet.

Keep Exploring

Continue learning

Explore related guides, tools, workflows, and prompts that help you go deeper into this topic.

See more AI tool comparisons

Browse all side-by-side AI tool comparisons on Ainanza.

Frequently Asked Questions

Is Qwen3.8-Max open source?

Not yet, as of this writing. It's currently API-only through Alibaba Cloud's Model Studio. Alibaba has committed to an open-weight release the week of August 10, 2026, alongside a smaller Qwen3.8-27B checkpoint, but the weights weren't out at the time this comparison was last updated.

Can Kimi K3 be self-hosted?

Yes, in principle. Moonshot AI released Kimi K3 as open-weight on July 17, 2026, with the full weight rollout completing by July 27, 2026. At 2.8 trillion total parameters, described as the largest open-weight model released to date, self-hosting requires serious multi-GPU infrastructure, not something a small team runs casually.

Which model is bigger, Qwen3.8-Max or Kimi K3?

By total parameters, Kimi K3 is larger at 2.8 trillion versus Qwen3.8-Max's 2.4 trillion. Total parameter count isn't a reliable proxy for real-world quality, though, and neither company has published a full independent benchmark comparison between the two models specifically.

Which is better for coding?

Both report strong coding and agent benchmark results, but they're vendor-reported by different companies using different test conditions, so a direct, apples-to-apples comparison isn't publicly available yet. Alibaba's own table shows mixed results across benchmarks (strong on Terminal-Bench, weaker on SWE-bench Pro); Moonshot reports Kimi K3 beating several closed rivals on coding and agent tasks. Test both on your own repository before choosing.

Last updated: