BLOG
GLM Reveals Self-Built Inference Cluster of 100,000 Domestic Accelerator Cards: Examining Engineering Quality and Narrative Quality Separately
Trend Tracking: Trend Release × Technical Judgment × Practical Advice. Author: Yongliang
On September 17, the z.ai official blog published “Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure”, which hit the Hacker News front page the same day with 366 points. The blog claims that Zhipu built a complete production-grade inference service from scratch on a cluster of over 100,000 domestic AI accelerator cards, and that all production inference for GLM-5.3-Flash runs on this system. It also claims that no one has deployed a domestic accelerator card cluster at this scale before. The blog frames this within the context of “Recursive Self-Improvement”: AI helps humans build AI infrastructure, and the next generation of models will be trained on facilities built with the participation of this generation of AI. On the same day, however, the hottest topic in the HN comment section was another matter—price hikes. The story of engineering self-reliance and the feeling in users’ wallets collided on the same front page; in this issue, we’ll read about both things together.
What Happened
Presenting the facts chronologically.
On September 17, the blog post was published and hit the HN front page with 366 points on the same day. To summarize the core content in one sentence: The official statement claims that Zhipu built a production-grade inference service from scratch on over 100,000 domestic AI accelerator cards, and all production inference for GLM-5.3-Flash runs on this system.
The blog also recounted a story about an anonymous launch: According to the official statement, GLM-5.3-Flash was launched on OpenCode and OpenRouter under the anonymous codename Ox-Alpha, becoming the number one in usage on both platforms within a week of launch, processing over 62 trillion tokens in six days. That is to say, without even discussing the cluster scale, this model itself has already undergone a round of market verification on two third-party platforms under an anonymous identity—according to the official account.
Official Stance: Four Points Consolidated
Consolidating the self-reported official content into four points; all below carry the attribution “the blog states”.
First point, scale and substance. Over 100,000 domestic AI accelerator cards, building production-grade inference from scratch, with all production inference of GLM-5.3-Flash running on them. The blog self-reports four categories of difficulties: limited chip video memory and bandwidth; new architecture combined with 1M context and multimodality, pushing up complexity simultaneously; immature software ecosystem; incomplete kernel support and missing documentation, with some parts relying on guesswork. The self-reported tech stack includes: intra-node tensor parallelism (linear attention and LM Head), ReplaySSM, W8A8 quantization, INT8/FP8/BF16 mixed precision cache quantization, Layer Split, EPD (Encode-Prefill-Decode) decoupled architecture. This list itself is informative: it roughly outlines the challenges to tackle when doing inference on domestic cards.
Second point, numbers. The blog states that end-to-end performance improved by approximately 3× compared to the initial baseline; hardware utilization and single token cost “reached a level comparable to mainstream NVIDIA GPUs” (official quote).
Third point, Agent narrative. The blog states that a large amount of work was completed by the Infra Agent driven by GLM-5.3, taking less than two weeks from the first successful run to production readiness; security partners discovered thousands of vulnerabilities in real codebases using GLM; GLM-5.3 is the team’s daily coding partner, “steadily moving towards replacing us” (official quote).
Fourth point, RSI framework. The blog’s title is “Towards Recursive Self-Improvement”: GLM helping humans build AI infrastructure will change the training method of the next generation of models; accompanied by a trusted access program, opening capabilities to security research partners.
The Calm Other Half
First Point: 3× is relative to one’s own initial baseline. “Approximately 3× improvement over the initial baseline”—the baseline is the first version built by the team, and 3× is a relative figure against themselves. It indicates that the system has been optimized diligently this year, but it does not indicate its position in any industry coordinate system. Publishing relative figures as a report card isn’t wrong in itself, but readers need to automatically perform a conversion: the lower the starting point, the different the value of the same 3×. The official blog won’t do this calculation for you.
Second Point: “Comparable to mainstream NVIDIA GPUs” lacks a benchmark target. Which specific model, which generation, what workload, and under what precision is the data for “mainstream NVIDIA GPU”? The blog doesn’t say. This isn’t textual nitpicking: “comparable” is a benchmarking conclusion, and a benchmarking conclusion must include a benchmark target; otherwise, it is just an adjective that sounds like a conclusion. The part readers can verify is zero; the part they can feel is entirely tone.
Third Point: RSI is a framework, not a conclusion. Recursive Self-Improvement (RSI) has a strict technical meaning: AI improves itself, and the rate of improvement accelerates on its own. What actually happened in the blog is: inference engineering optimization (quantization, parallelism, caching strategies), plus Agent-assisted infrastructure development. This is solid engineering, but it is not “AI building AI”. Packaging standard inference engineering into a grand narrative is a standard move for vendor blogs nowadays—the weight of the terminology determines how far the story travels, but readers receive it based on the story’s weight and realize it based on the engineering’s weight, creating a gap in between. Recognizing this gap is more useful than arguing whether RSI holds true.
Fourth Point: The temperature difference on the same day. The blog tells a story of inclusivity and computing power self-reliance; on the same day, the user sentiment in the HN comment section was price hikes and tightened quotas—according to the comment section, the mid-tier plan rose from about $20/month to about $80/month, the old Max plan ($360/year) was delisted, the current Max is about $168/month, and according to official documentation, this equates to about $1100 worth of GLM-5.3 quota, with multiple users complaining about tight quota limits. It needs to be made clear: these are currently based on comment section and documentation figures, and no official pricing announcement has been seen yet. But the perceived temperature difference is real—the grander the infrastructure narrative, the more users will check the price list against the promise of inclusivity. Two pieces of information appearing on the same day: reading either one alone is biased; reading them together gives the complete meaning.
Fifth Point: Stance. The 100,000-card cluster currently relies solely on official self-reporting, with no confirmation from any third party or client side. Let me state my stance clearly: if true, this is a factual-level engineering achievement; the production and deployment of domestic cards at this magnitude and difficulty is a real tough battle and shouldn’t be dismissed with a single word like “hype”; but the other side of the audit also holds true—100,000 cards, 3×, two weeks, 62 trillion tokens, thousands of vulnerabilities, all key figures are currently self-reported by the official side, and not a single one has third-party verification. Don’t talk it down, and don’t accept it all blindly. This applies not just to this one company, but is the default posture for reading any vendor blog.
Worth Watching
First, independent verification of cluster scale. Whether third-party corroboration for the 100,000 cards will appear—follow-up by foreign media, evidence from the client side, independent review by the technical community. Currently, it is zero; this is the most significant number in the entire piece, yet also the one most lacking in verification.
Second, official announcements on pricing. Whether the price increases circulating in the comments section materialize into updates on the official pricing page. If materialized, the “temperature difference” mentioned earlier would upgrade from a comment-level stance to an official stance, making the foundation for discussion more solid.
Third, completing the “comparable”. Whether the official source publicly reveals the benchmarking target and load conditions. Even if it is just a sentence like “under a certain model and typical load”, this conclusion would only then begin to enter the realm of discussable topics.
Fourth, follow-up by foreign media. Currently, the sources for the entire piece are limited to the official blog and HN. If media outlets like The Information, Bloomberg, and SCMP follow up, especially if independent technical reviews appear, the certainty of the corresponding facts can be upgraded.
Conclusion
If the 100,000 domestic accelerator cards can be independently verified, the engineering significance of this matter requires no narrative embellishment; it is worth recording on its own merits. What needs to be discounted is the narrative part: 3× is a relative figure, “comparable” has no benchmark, and RSI is framework packaging. Separate the engineering substance from the narrative substance—the former awaits third-party verification, while the latter can be read as an audit right now. Vendor blogs are both advertising spaces and technical archives; the difference lies not with the publisher, but in whether the reader performs that step of division.
Sources
- z.ai official blog 2026-09-17 “Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure” (key figures are all officially self-reported)
- Hacker News discussion on the same day (366 points): pricing details in the comments, quota experience, and other viewpoints such as “a minor planet-level event for Western AI labs” (commentary viewpoints, not facts)
- z.ai official documentation: Max plan quota specifications (via HN comments)