Omen Alpha is the likely 2.0 successor; test before switching.
Omen Alpha is not literally the same model as Ox Alpha. The community position is that Omen Alpha is the newer, separately distributed stealth model — a likely 2.0 successor — with different observed pricing, behavior, context, and benchmark results. Ox Alpha was later identified as Z.ai’s GLM-5.3-Flash; Omen Alpha remains unconfirmed, though evidence points toward a GLM-5.x lineage.
Two stealth models, one likely lineage.
Ox Alpha appeared first as an anonymous coding model with a reported 1M-token context and distribution through OpenRouter and OpenCode. It was later identified as Z.ai GLM-5.3-Flash. See the Wavect account for background.
Omen Alpha appeared around September 4, 2026 — a separate model with different pricing, a reported ~500K-token context, and the OpenCode vendor path zhipu/omen-alpha. Behavioral fingerprints support a GLM-family hypothesis, but no official Z.ai statement confirms Omen Alpha’s identity.
The evidence supports this shorthand: Ox Alpha = earlier model (GLM-5.3-Flash); Omen Alpha = likely 2.0 successor (possible GLM-5.x variant, unconfirmed).
Omen Alpha vs Ox Alpha: full comparison.
The following values combine public listings and community reports. Items marked as community-reported are not vendor specifications.
| Attribute | Omen Alpha (likely 2.0) | Ox Alpha (earlier model) |
|---|---|---|
| Status | Active stealth preview | Earlier stealth model; identified as GLM-5.3-Flash |
| Likely lineage | Possible Z.ai / GLM-5.x variant; unconfirmed | Z.ai GLM-5.3-Flash |
| Input price | $0.20 / 1M tokensTokenra: $0.08 · 60% off | $0.075 / 1M tokens |
| Output price | $0.66 / 1M tokensTokenra: $0.26 · 60% off | $0.25 / 1M tokens |
| Cached input | $0.04 / 1M tokens | $0.015 / 1M tokens |
| Context window | ~500K tokens (community/platform-reported) | 1M tokens (public model listings) |
| Max output | 128K tokens (platform-reported) | 128K tokens |
| Input modalities | Text + image (community/platform-reported) | Text + image + video (public listings) |
| Reasoning | Supported (platform-reported) | Supported |
| API format | OpenAI-compatible Chat Completions | OpenAI-compatible Chat Completions |
| Model ID | omen-alpha | stealth/ox-alpha |
| Primary access | OpenCode Go · tokenra.io | OpenRouter · OpenCode · oxalpha.io |
| Subscription / offer | OpenCode Go: reported $10/mo with $100 Omen Alpha credit | Pay per use; oxalpha.io lists pricing at roughly 40% of reference rates |
| Benchmark | 23.14/40 in the dated OpenCode snapshot → | No directly comparable same-run benchmark available |
| Privacy | Not used for training · stated 0-day retention | Check the active provider’s terms |
| Vendor confirmed? | No — no company has publicly claimed it | Yes — identified as Z.ai GLM-5.3-Flash |
What each model costs right now.
Omen Alpha lists at $0.20 input / $0.66 output per million tokens, with Tokenra offering a 60% discount at $0.08 / $0.26. OpenCode Go is reported at $10 per month with $100 in monthly Omen Alpha credit. See full pricing details →
Ox Alpha lists at $0.075 input, $0.25 output, and $0.015 cached input per million tokens. The oxalpha.io offer is described as roughly 40% of reference pricing. Check the active route before budgeting.
What the available numbers say — and don’t say.
In the dated OpenCode snapshot, Omen Alpha scored 23.14 / 40, with an average cost of $0.03 per prompt and average time of 01:51.
A direct numerical comparison with Ox Alpha is not currently defensible:
- The available Ox Alpha results use different task sets, dates, or evaluators.
- Labels such as
ox-alpha-maxon other leaderboards belong to separate measurement contexts. - A fair comparison requires the same prompts, evaluator, settings, and date.
When should you choose each model?
For the active release.
Choose it when you want the newer stealth model, current OpenCode Go access, or Tokenra’s discounted route.
For known lineage.
Choose it when the confirmed GLM-5.3-Flash identity, reported 1M context, or lower listed rate matters more.
| If you prioritize | Start with | Why |
|---|---|---|
| Newest stealth release | Omen Alpha | It is the active newer release with current OpenCode Go and Tokenra access. |
| Larger reported context | Ox Alpha | Its public listings report 1M versus Omen Alpha’s reported ~500K. |
| Video input | Ox Alpha | Public listings describe text, image, and video input. |
| Lowest listed token rates | Compare live routes | Ox Alpha has lower listed rates; Tokenra’s Omen Alpha discount is close on input/output. |
| Vendor transparency | Ox Alpha | It has been identified as GLM-5.3-Flash; Omen Alpha remains unconfirmed. |
| Production coding reliability | Run both | Your repository, prompts, tools, retries, and failure policy determine the result. |
To test Omen Alpha, follow the setup guide or use the API reference.
Omen Alpha vs Ox Alpha questions.
Is Omen Alpha better than Ox Alpha?
There is no universal answer. Omen Alpha’s dated OpenCode result is 23.14/40, but a fair comparison requires the same prompts, evaluator, and date. It is a newer likely successor, not automatically better in every scenario.
Is Omen Alpha the same as Ox Alpha?
No. Ox Alpha was the earlier stealth model later identified as GLM-5.3-Flash. Omen Alpha is a newer, separately distributed model generally treated as its likely 2.0 successor.
Which model costs less?
Ox Alpha’s listed rates are $0.075 input, $0.25 output, and $0.015 cached input per million tokens. Omen Alpha’s official rates are $0.20/$0.66, while Tokenra advertises $0.08/$0.26. Confirm live pricing and promotions.
Can I still use Ox Alpha?
Availability depends on the provider route. Check OpenRouter, OpenCode, or the applicable provider for current status. Omen Alpha is the newer active model available through OpenCode Go and tokenra.io.
Why does Ox Alpha report 1M context while Omen Alpha reports ~500K?
Omen Alpha’s ~500K figure is community/platform-reported, not a vendor model card. A provider can also apply operational limits below a model’s theoretical maximum. Test the active route.