Mystery high-performance AI turns out to be China’s “GLM-5.3-Flash” — a rival to Anthropic

From our other sites

The story

On OpenRouter, a platform that lets users access multiple AI models, an unidentified high-performance AI called “Ox Alpha” appeared and drew attention — but the Chinese company Z.ai has now revealed itself as the developer, confirming it as “GLM-5.3-Flash.” The model’s weights exceed 300GB, and while it can be downloaded and run for free, actually running it reportedly requires substantial computing resources. On the 5ch Science News+ thread, some argued that Chinese firms are exploiting a gap left by Anthropic and other companies skimping on their cheaper model tiers, while others raised doubts about whether it’s merely a distillation of an existing model, alongside technical questions about whether it can really run on an ordinary PC.

Mystery high-performance AI revealed to be China’s “GLM-5.3-Flash” — a rival to Anthropic and OpenAI

Last week, an anonymous AI model called “Ox Alpha” appeared on OpenRouter, a platform offering access to multiple AI models.

The Chinese company Z.ai has now revealed that it is the developer behind the model.

Source: japan.cnet.com / Original article here

What people said

2AnonymousAug 29, 2026 09:25
Rakuten's "AI" is hilarious lol (w = Japanese text-equivalent of "lol")
3AnonymousAug 29, 2026 09:49
Anthropic's Claude Code is the frontrunner, and that's exactly why it's letting its cheap Haiku tier slack off.

Chinese AI companies are throwing their own frontier models right into that price bracket.
6AnonymousAug 29, 2026 13:36
Re: #1's original post:
>> Since GLM-5.3-Flash is open-weight, you can also download it and run it on your own hardware. In that case, you don't need to pay any subscription or API fees. That said, it likely requires serious computing power, since the model weights alone exceed 300GB.
11AnonymousAug 29, 2026 13:42
Re: #6-9 let you see at a glance just how capable Re: #1 actually is.


If you optimize Re: #1 to work like the AI program in Re: #10, it should become usable on an ordinary PC too.
7AnonymousAug 29, 2026 13:37
That's what it says, but people are doubting whether it'll actually start up on an average person's PC.
12AnonymousAug 29, 2026 13:43
Re: #11 If your GPU and memory usage are running high, use this as a reference for optimizing your setup:

"Colibrì," an inference engine that runs the massive 744-billion-parameter AI "GLM-5.2" on an ordinary PC with just 25GB of memory, has arrived
July 10, 2026, 4:20 PM
https://gigazine.net/news/20260710-colibri-glm/
>> Runs on a 12-core CPU with 25GB of memory
14AnonymousAug 29, 2026 13:46
Once you optimize it to the level of Re: #13, you basically end up with the same thing as a program that's already running AI locally.

Tencent releases its "Hy3" AI model as an open model — at 295B parameters it rivals GLM-5.2 and DeepSeek-V4, and even beats GPT-5.5 on science tasks
July 7, 2026, 10:51 AM
https://gigazine.net/news/20260707-tencent-ai-hy3/

US-made open model "Laguna XS 2.1" arrives — runs locally and beats Claude Haiku 4.5 on some benchmarks
July 3, 2026, 11:43 AM
https://gigazine.net/news/20260703-laguna-xs-2-1/
37AnonymousAug 30, 2026 17:00
Does it even run on Chinese-made AI chips?
When did they develop chips like that?
If they can build a whole AI stack without Nvidia, nobody's going to be able to touch China in AI anymore.
39AnonymousAug 30, 2026 19:36
Re: #1
I asked it about some obscure Japanese topics and it was completely useless.
ChatGPT and Claude seem to have trained on Japanese web content to some degree, at least.

Background and key points of the discussion

GLM-5.3-Flash is a large language model built by the Chinese AI company Z.ai, and it took an unusual path to the spotlight — quietly deployed on OpenRouter under an anonymous name before its identity was later revealed. In the thread, two separate facts tend to get conflated: that the model’s weights are a massive 300GB+, making it hard to run on a typical PC, and that it’s “open-weight,” meaning there’s no usage fee once you’ve downloaded it. Actually running it locally would require additional compression and optimization tools — something like the “Colibrì” inference engine, which lets GLM-5.2 run on a PC with just 25GB of memory. With large Chinese-made models, the recurring question is always whether they’re genuinely original or just a “distillation” trained on existing Western models, and this thread reached no verdict on that either — treating raw performance and originality of development as two separate issues.

*This article is excerpted and summarized from the 5ch (Science News+) thread “[AI] Mystery high-performance AI revealed to be China’s “GLM-5.3-Flash” — a rival to Anthropic and OpenAI.”

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *