Try every model, $3 once
1M tokens across the full model catalog. No expiry, credit card required.
One API key across Claude Code, Gemini CLI, and OpenAI CLI. Top three frontier models per provider, transparent per-token pricing, credit card required.
Every plan includes the same top 3 frontier models. Pick by token volume, not by capability.
1M tokens across the full model catalog. No expiry, credit card required.
3M tokens every 5-hour window. Predictable, generous rate limits.
5M tokens every 5-hour window. Priority queue and full observability.
20M tokens every 5-hour window. Slack support, dedicated capacity.
Every token plan includes the same 9 models — 3 top-tier per provider. Pick your plan by token volume, not by capability.
Flagship reasoning with chain-of-thought, native tool use, and 256K context.
Google's flagship multimodal models with deep thinking modes and 1M context.
OpenAI's latest reasoning and chat models with native function calling and JSON mode.
Pick the plan that matches your workload. Upgrade or downgrade any time, prorated.
| Feature | Test | Pro | Max | Ultra |
|---|---|---|---|---|
Tokens included per 5-hour window | 1M (one-time) | 3M | 5M | 20M |
Models in catalog 3 top frontier per provider | 9 | 9 | 9 | 9 |
CLI tools Claude Code, Gemini CLI, OpenAI CLI | All 3 | All 3 | All 3 | All 3 |
API keys one key works across all 3 CLIs | 1 | 1 | 1 | 1 |
Priority queue | — | — | ✓ | ✓ |
Concurrent sessions | 1 | 3 | 5 | 10 |
Usage analytics | Basic | Basic | Advanced | Advanced + alerts |
Support response time | Community | Email · 24h | Email · 12h | Slack · 4h |
Annual billing 2 months free · 10 months paid | — | $250/yr $300Save $50/yr | $350/yr $420Save $70/yr | $800/yr $960Save $160/yr |
PAYG rates for the API directly. No plan required, billed per token.
7 IQx ultra fast models — GPT OSS, Llama 4 Scout, Qwen3 32B, Llama 3.3 70B, Llama 3.1 8B.
Frontier reasoning — DeepSeek V4 Pro, GLM-5.1, Qwen3.5-397B, Kimi-K2.6, Hermes-4-405B.
4 vision models — Nemotron-3-Ultra-550B, Cosmos3, MiniCPM-V-4.5, Qwen2.5-VL-72B.
Everything you need to know about Token Plan pricing and billing.
Token Factory works with Claude Code, Gemini CLI, and OpenAI CLI. Every plan includes the same top 3 frontier models from each provider — Anthropic Claude (Opus 4.8, Sonnet 4.6, Haiku 4.5), Google Gemini (2.5 Pro, Flash, Flash-Lite), and OpenAI (GPT-5.5, 5 mini, 5 nano).
One API key per account, regardless of plan. The same key works across Claude Code, Gemini CLI, and OpenAI CLI — set ANTHROPIC_API_KEY, GEMINI_API_KEY, or OPENAI_API_KEY to the same value in your environment. Even Ultra does not unlock additional keys.
Each plan refreshes its included token allowance every 5 hours. Pro gives 3M tokens per window, Max 5M, and Ultra 20M. Tokens don't roll over — unused allowance resets at the start of the next window. Window timers are visible in your dashboard.
Requests pause for the remainder of the current 5-hour window once you hit the cap. Top up via the wallet to continue, or wait for the next window reset. There is no surprise overage — you can see your usage live in the dashboard.
Annual plans charge 10 months upfront for 12 months of access. Pro is $250/yr (save $50), Max is $350/yr (save $70), Ultra is $800/yr (save $160). The Test plan is one-time only — no annual concept. Switch from monthly to annual on renewal; switch from annual to monthly at the end of the term.
Yes. Upgrades are prorated based on the days left in your current cycle — you pay the difference, and the new plan takes effect immediately. Downgrades take effect at the start of your next billing cycle. The prorated amount is shown before you confirm.
All major credit cards via Stripe. A card is required to start any plan — there is no option without a payment method. Wire transfer and ACH are available for Ultra annual contracts over $5,000.