From our other sites
The story
On OpenRouter, a platform that lets users access multiple AI models, an unidentified high-performance AI called “Ox Alpha” appeared and drew attention — but the Chinese company Z.ai has now revealed itself as the developer, confirming it as “GLM-5.3-Flash.” The model’s weights exceed 300GB, and while it can be downloaded and run for free, actually running it reportedly requires substantial computing resources. On the 5ch Science News+ thread, some argued that Chinese firms are exploiting a gap left by Anthropic and other companies skimping on their cheaper model tiers, while others raised doubts about whether it’s merely a distillation of an existing model, alongside technical questions about whether it can really run on an ordinary PC.
Mystery high-performance AI revealed to be China’s “GLM-5.3-Flash” — a rival to Anthropic and OpenAI
Last week, an anonymous AI model called “Ox Alpha” appeared on OpenRouter, a platform offering access to multiple AI models.
The Chinese company Z.ai has now revealed that it is the developer behind the model.
Source: japan.cnet.com / Original article here
What people said
Chinese AI companies are throwing their own frontier models right into that price bracket.
>> Since GLM-5.3-Flash is open-weight, you can also download it and run it on your own hardware. In that case, you don't need to pay any subscription or API fees. That said, it likely requires serious computing power, since the model weights alone exceed 300GB.
If you optimize Re: #1 to work like the AI program in Re: #10, it should become usable on an ordinary PC too.
"Colibrì," an inference engine that runs the massive 744-billion-parameter AI "GLM-5.2" on an ordinary PC with just 25GB of memory, has arrived
July 10, 2026, 4:20 PM
https://gigazine.net/news/20260710-colibri-glm/
>> Runs on a 12-core CPU with 25GB of memory
Tencent releases its "Hy3" AI model as an open model — at 295B parameters it rivals GLM-5.2 and DeepSeek-V4, and even beats GPT-5.5 on science tasks
July 7, 2026, 10:51 AM
https://gigazine.net/news/20260707-tencent-ai-hy3/
US-made open model "Laguna XS 2.1" arrives — runs locally and beats Claude Haiku 4.5 on some benchmarks
July 3, 2026, 11:43 AM
https://gigazine.net/news/20260703-laguna-xs-2-1/
When did they develop chips like that?
If they can build a whole AI stack without Nvidia, nobody's going to be able to touch China in AI anymore.
I asked it about some obscure Japanese topics and it was completely useless.
ChatGPT and Claude seem to have trained on Japanese web content to some degree, at least.
Background and key points of the discussion
GLM-5.3-Flash is a large language model built by the Chinese AI company Z.ai, and it took an unusual path to the spotlight — quietly deployed on OpenRouter under an anonymous name before its identity was later revealed. In the thread, two separate facts tend to get conflated: that the model’s weights are a massive 300GB+, making it hard to run on a typical PC, and that it’s “open-weight,” meaning there’s no usage fee once you’ve downloaded it. Actually running it locally would require additional compression and optimization tools — something like the “Colibrì” inference engine, which lets GLM-5.2 run on a PC with just 25GB of memory. With large Chinese-made models, the recurring question is always whether they’re genuinely original or just a “distillation” trained on existing Western models, and this thread reached no verdict on that either — treating raw performance and originality of development as two separate issues.
*This article is excerpted and summarized from the 5ch (Science News+) thread “[AI] Mystery high-performance AI revealed to be China’s “GLM-5.3-Flash” — a rival to Anthropic and OpenAI.”
Leave a Reply