On August 28, 2026, Tencent's Hunyuan team released and open-sourced Hy4 preview, its next-generation large language model: 770 billion total parameters, 49 billion activated per token, and a context window exceeding 1 million tokens. The positioning is explicit — "built for productivity." In Tencent's own blind evaluation (163 internal experts, 203 engineering tasks), Hy4 preview scored 2.99/4, narrowly beating Zhipu's GLM 5.3 (2.92) and Moonshot's Kimi K3 (2.94). The release lands days after ByteDance's Doubao Work launch and Alibaba's Qwen3.8-Flash, completing a week in which all three giants shipped both a model and an office-agent product. The AI office war now has a full model-layer front.
1. What Happened: Release Plus Open Weights
Hy4 preview is Hunyuan's biggest scale-up since it rebuilt its infrastructure in February 2026. Key specifications:
| Spec | Value |
|---|---|
| Total parameters | 770B |
| Active parameters | 49B per token |
| Context window | 1M+ tokens |
| Positioning | "Built for productivity" — coding, office, science |
| Status | Released and open-sourced (top tier of open models, per Tencent) |
The model ships simultaneously across Tencent products: WorkBuddy and CodeBuddy (Chinese and international versions), Yuanbao, and ima. Developers can call the API via Tencent Cloud TokenHub or OpenRouter.
2. The Numbers: A Blind Test Against GLM 5.3 and Kimi K3
The most quotable evaluation is Tencent's internal blind test: 163 internal experts scored 203 engineering tasks, with Hy4 preview averaging 2.99/4 versus GLM 5.3 at 2.92 and Kimi K3 at 2.94. The margins are narrow — and the evaluation is Tencent's own, so treat it as a directional signal rather than an independent benchmark. Public benchmark tables released alongside the model position Hy4 preview against Qwen 3.8 Max, DeepSeek V4 Pro, GPT 5.6 Sol, GLM 5.3, Kimi K3, and Claude Opus 5 across agentic-coding suites (Terminal Bench 2.1, DeepSWE, ALE-CLI, Toolathlon) and reasoning sets (Humanity's Last Exam, HorizonMath).
Two caveats worth stating plainly: internal blind tests favor the evaluator's own task distribution, and "preview" means the final Hy4 may shift. The credible core claim is narrower but still significant: Hy4 preview is competitive at the top of the open-source tier, and it is differentiated specifically on productivity workloads.
3. Built for Productivity: Four Scenario Upgrades
Hy4 was co-designed with WorkBuddy and CodeBuddy, with training data co-developed alongside Tencent's senior experts in software engineering, gaming, finance, and security:
| Scenario | What's New |
|---|---|
| Software engineering | Stronger long-horizon understanding, planning, debugging, and validation; better front-end visual quality — demoed generating a native-WebGL multi-chapter interactive comic page that maintained character art style and narrative continuity |
| Office & analysis | Better complex-workspace understanding and financial analysis; optimized data analysis and cross-file collaboration — demoed checking invoice compliance across 72 files against 3 policy documents and delivering documents, spreadsheets, and presentations |
| Game development | One-sentence requests produce playable prototypes; proficient with game engines — demoed building a shooting-game demo from scratch in Unreal 5 via MCP, purely through conversation |
| Scientific research | Improved reasoning on complex research problems across AI R&D, molecular dynamics simulation, condensed matter physics, and foundational mathematics |
The office scenario deserves emphasis for marketing readers: a model that can ingest dozens of files, apply policy rules, and deliver finished documents and decks is precisely the "digital employee" workflow that Doubao Work and Qianwen Office are also racing to own.
4. Recursive Self-Improvement: The Model Helped Build Itself
The most technically notable claim: Hy4 preview participated in its own development. It proposed approaches for training methods, data strategies, evaluation systems, and low-level operator optimization; ran experiments; and fed results back into subsequent iterations — an early closed loop of recursive self-improvement. It also analyzed its own inference infrastructure bottlenecks, optimizing operator fusion and communication for a 31.8% end-to-end throughput gain over baseline, stable across context lengths and concurrency levels. Self-directed infrastructure tuning at this scale is an early but real marker of where model development is heading: models as agents in their own training pipelines.
5. Pricing and Availability
| Item | Price |
|---|---|
| Input | ¥6 / million tokens ($0.834) |
| Output | ¥18 / million tokens ($2.501) |
| Cache hit | ¥0.3 / million tokens ($0.042) |
Premium relative to Qwen3.8-Flash (¥1/¥3) but priced as a flagship-tier model, not a commodity. To accelerate feedback collection, WorkBuddy and CodeBuddy are free for a limited two-week period, and Hy3 access has been extended free until September 30. This is a land-grab move: get enterprise users into WorkBuddy on Hy4 before ByteDance's 30-day Doubao Work trial window converts them elsewhere.
6. The Preview-First Strategy and the Two-Month Cadence
Since rebuilding its infrastructure in February 2026, Hunyuan has shipped a major model iteration roughly every two months, using a preview-first, GA-second rhythm that folds real-world feedback into development. Tencent says the next Hy4 version will begin rolling out shortly. The pattern mirrors Alibaba's "architecture canary" with Qwen3.8-Flash-Next: ship the frontier openly, let the ecosystem adapt, then formalize. Between Qwen's preview strategy and Hunyuan's preview cadence, China's top labs have converged on continuous-open-release as the default competitive posture — very different from the closed, wait-for-the-keynote approach of their US counterparts.
7. What This Means for Marketers and Brands
- The model layer and the office-agent layer are now one battlefield. Tencent pairs Hy4 with WorkBuddy; Alibaba pairs Qwen with Qianwen Office; ByteDance pairs Doubao models with Doubao Work. For B2B brands selling into China, the practical question is no longer "which model is best" but "which agent ecosystem will host your customers' daily work" — and each ecosystem is now locking in users with free windows (WorkBuddy 2 weeks, Doubao Work 30 days).
- Open weights lower the barrier to custom AI. Hy4's open release means Chinese enterprises and agencies can self-host a top-tier productivity model for compliance-sensitive workflows — relevant for finance, healthcare, and government-adjacent marketing programs that cannot send data to foreign APIs.
- Productivity benchmarks are becoming procurement criteria. Blind expert panels on real engineering and office tasks (Tencent's 203-task test) are a better proxy for enterprise value than chat leaderboards. When evaluating models for client work, replicate the pattern: score candidates on your client's actual deliverables.
Takeaway: In one week, ByteDance, Alibaba, and Tencent each shipped both a frontier model and an office agent. The Chinese AI stack is consolidating into three vertically integrated ecosystems, each with its own model, agent, and distribution surface. Brands should map which ecosystem their Chinese audience works in — and prepare content and services that operate inside those agents.
Sources: 腾讯混元官方发布 via 腾讯新闻 (2026-08-28) https://new.qq.com/rain/a/20260828A07JQX00 ;智通财经 (2026-08-28) https://view.inews.qq.com/a/20260828A07PM500 ;Pandaily (2026-08-28);The Edge Malaysia (2026-08-28).