Tencent released and open-sourced Hy4 preview on August 28, 2026, a 770-billion-parameter Mixture-of-Experts model that marks two notable firsts in the open-weight landscape: a top-tier productivity benchmark score and, more significantly, a documented recursive self-improvement loop in which the model helped design its own training process.

For enterprise AI teams evaluating open-source alternatives to closed frontier models, Hy4 preview is the strongest productivity-focused open-weight release since Kimi K3 in July 2026, with meaningful differences in design philosophy and an unprecedented self-optimization story.

What Tencent Released

Hy4 preview ships as a Mixture-of-Experts model with 770 billion total parameters and 49 billion active parameters per token. The context window exceeds 1 million tokens. According to Tencent’s official announcement, the model is available globally through:

  • WorkBuddy and CodeBuddy (free for two weeks from launch)
  • Tencent Cloud TokenHub API at $0.834 per million input tokens, $2.501 per million output tokens, $0.042 per million cache hits
  • OpenRouter for API access
  • Hugging Face for self-hosted weights

The model is positioned around five use cases: software engineering, office productivity and data analysis, game development, scientific research, and financial analysis. This is a deliberate departure from the benchmark-optimized approach of most frontier releases. Tencent built Hy4 preview using high-quality training data co-created with internal domain experts across software engineering, gaming, finance, and security rather than relying on generic web crawl data for the critical domains.

TechNode confirmed the model is available through both the Chinese and international versions of WorkBuddy and CodeBuddy, making it accessible to global enterprise teams without requiring a Tencent Cloud account for initial evaluation.

The Recursive Self-Improvement Story

The most consequential detail in the Hy4 preview announcement is not the parameter count or the benchmark score. It is the role the model played in its own development.

According to Tencent’s official documentation, Hy4 preview participated for the first time in the automated optimization of its own training methods, data strategies, evaluation frameworks, and low-level operators. The model proposed approaches, ran experiments, and fed the resulting code, logs, and feedback back into subsequent rounds of exploration. Tencent describes this as an early-stage recursive self-improvement loop.

Separately, the model autonomously identified performance bottlenecks in Tencent’s inference system and carried out multiple rounds of optimization on operator fusion and communication optimization. The result was a 31.8% increase in end-to-end inference throughput compared to the baseline, with consistent gains across different context lengths and concurrency levels.

This matters for enterprise AI teams for a practical reason beyond the capabilities of the model itself. A system that can meaningfully contribute to its own infrastructure optimization signals a new category of utility: not just a model you run, but one that can analyze and improve the systems it runs on. For engineering teams building large-scale AI infrastructure, this capability profile has direct implications for how to structure human-AI collaboration on optimization work.

For broader context on how AI systems are being used to secure and monitor their own environments, see the earlier coverage of Anthropic’s multiagent research findings and governance considerations for AI agents that can affect their own operating context.

Benchmark Position Against Open-Weight Peers

Tencent conducted a blind evaluation with 163 experts across 203 engineering tasks. The results position Hy4 preview at the top of the current open-weight productivity tier:

ModelInternal Eval Score (out of 4.00)Active Parameters
Tencent Hy4 preview2.9949B
Kimi K32.94(disclosed at launch)
GLM-5.32.9218B

Source: Tencent official announcement, August 28, 2026. Evaluation conducted internally by Tencent; independent third-party verification pending.

The Kimi K3 and GLM-5.3 comparisons are important reference points. Kimi K3 was the previous open-weight benchmark leader for enterprise workloads when it launched in July 2026 with 2.8 trillion total parameters. GLM-5.3, the open-weight model from Z.ai, established strong coding benchmarks, as covered in earlier analysis of GLM-5.3’s enterprise security capabilities.

Hy4 preview edges both on Tencent’s composite productivity evaluation. The important caveat: all competitor numbers in this comparison come from Tencent’s internal evaluation runs. Until independent benchmarking appears on indices like Artificial Analysis or Hugging Face’s public leaderboards, treat these figures as directionally useful rather than definitive.

Hy4 preview does not publish granular benchmark breakdowns across categories like SWE-bench, BrowseComp, or MCP-Atlas in the initial announcement. Tencent’s benchmark narrative focuses on the composite expert evaluation and the productivity domain coverage. Teams running specific workloads should test on their own data rather than waiting for complete third-party benchmarks.

Productivity Design Over Benchmark Optimization

Where Hy3, Kimi K3, and GLM-5.3 each emphasized certain technical benchmarks in their launch narratives, Hy4 preview leads with domain coverage and product co-design. The model was developed through deep collaboration with Tencent’s internal product teams, including WorkBuddy and CodeBuddy, across a sustained cycle of real-world feedback rather than pure benchmark optimization.

In software engineering, Tencent reports improvements in understanding, planning, debugging, and validation for long-context development tasks, along with stronger frontend generation quality. In office productivity, the model demonstrates better handling of financial analysis workflows and cross-document synthesis. In game development, the model can generate a playable prototype from a single natural-language request and iterate on complex game projects through multi-turn interactions. In scientific research, it improves on complex reasoning across AI research and development, molecular dynamics simulation, condensed-matter physics, and fundamental mathematics.

These are capability claims from Tencent’s own evaluation. But the design emphasis on real-world productivity rather than academic benchmarks is consistent with the broader shift in how enterprise teams should evaluate models: the question is not which model leads on MMLU but which model completes the actual workflows your team runs.

Reuters reporting from August 28 notes that Hy4 preview, as an early release, can sometimes take longer than necessary to work through complex questions and may over-verify its own answers. This is the honest disclosure expected from a preview model and should factor into evaluation planning.

Pricing and the Competitive Context

At $0.834 per million input tokens and $2.501 per million output tokens, Hy4 preview sits in a competitive range for open-weight models served via API. For comparison, Claude Opus 4.8 (the current Anthropic production workhorse) runs at significantly higher cost per token for equivalent capability tiers.

The two-week free period on WorkBuddy and CodeBuddy makes this the lowest-friction point for enterprise teams to evaluate Hy4 preview on real workflows before any API spend. For teams that have used the Hy3 free period for evaluation, Tencent has extended free access to Hy3 on both platforms until September 30.

The Hy4 series is not complete. Tencent notes that the next batch of models in the Hy4 series is expected soon. The preview-first approach means enterprise teams should expect a production-ready official release with improved reliability within weeks.

What Enterprise AI Teams Should Do

The Hy4 preview release presents a practical evaluation opportunity this week:

  1. Test on your actual productivity workloads via WorkBuddy, CodeBuddy, or OpenRouter during the free window. Focus on office document workflows, multi-step data analysis, and any agentic research pipelines your team runs.

  2. Do not use Hy4 preview for production coding agents yet. The preview label and Reuters’ observation about over-verification suggest it needs further stabilization for high-stakes autonomous code execution. GLM-5.3 remains the current open-weight coding benchmark leader.

  3. Watch for the official Hy4 release. A preview-first approach with this level of internal evaluation data suggests the official Hy4 is already in production validation. When it ships with reliability improvements, it becomes a serious evaluation target for production deployment.

  4. Track the self-improvement capability trajectory. The documented recursive loop in Hy4’s training is the most forward-looking signal in this release. Teams building AI infrastructure should monitor how Tencent develops this capability in subsequent Hy4 series releases.

The open-weight productivity tier is now competitive with the lower end of closed frontier models for real-world enterprise workloads. Hy4 preview is the clearest signal yet that the gap is closing on productivity-specific tasks, not just benchmark leaderboards.

If your team is working through model selection for enterprise agentic deployments or wants to understand how to build evaluation frameworks that reflect actual workload performance, connect with the Enera team.