DeepSeek V4 Pro — chain-of-thought at frontier parity
1M context, $2.10 / $4.20 per M tokens (+20% markup). Now serving in production with stable latency.
IQx Token Factory is the inference layer for modern AI — 33 open-source and frontier models through one OpenAI- and Anthropic-compatible API, with transparent per-token pricing and zero infrastructure to manage.
Frontier chain-of-thought models — DeepSeek V4 Pro, GLM-5.1, Qwen3.5-397B. Transparent per-token billing, 1M context on flagships.
Latency-critical inference on Language Processing Unit silicon. Up to 1,800 TPS for real-time agents, code completion, and interactive tools.
Production-grade open models from Qwen, Meta, Google, NVIDIA, and more. Apache 2.0 / Llama / Open-license, deployable anywhere.
Multimodal models with image + text understanding. Document Q&A, chart parsing, OCR, screenshot-to-code, and visual reasoning at production scale.
Plus 1 embeddings model for semantic search, RAG, and similarity tasks. View the embeddings model →
Your million tokens from $3, same plan.
Text, reasoning, code, vision, embeddings, and open-source general models — all through one OpenAI- and Anthropic-compatible interface. Swap providers by changing one URL.
# OpenAI & Anthropic compatible, drop-in from openai import OpenAI client = OpenAI( base_url="https://api.IQXtokenfactory.com/v1/", api_key="tf-•••••")Read the quickstart →
Peak tokens per second on GPT OSS 20B UltraFast
Time-to-first-token under load
Apache 2.0, Llama, and Open-licensed. Qwen, Meta Llama, NVIDIA Nemotron, Google Gemma, and more.
Browse open models →1M-token test plan, monthly windows, or pure pay-as-you-go.
View plans →Uptime backed by multi-region failover
On GPT OSS models — automatic prompt caching.
Works with every OpenAI and Anthropic SDK out of the box.
From your first prototype to multi-region production scale, the same API, the same billing model, the same observability stack.
Up to 1,000 tokens per second on Language Processing Unit silicon. Sub-100ms time-to-first-token, even at peak load.
Works with every OpenAI SDK, every Anthropic SDK, the Vercel AI SDK, LangChain, LlamaIndex, the OpenAI CLI, and Claude Code. Change one base URL — keep the rest.
No idle GPU cost, no minimums, no surprise overage. Stripe billing, Odoo finance sync, line-item invoices for every request.
Reasoning, code, multilingual, vision, embeddings, and open-source general models. Frontier and open-source, all under one key.
Automatic failover across three regions, real-time health monitoring, and a public status page. Your apps stay up.
Python, TypeScript, Go, and Rust SDKs. Type hints, streaming, retries, and full observability baked in. Five lines to your first request.
1M context, $2.10 / $4.20 per M tokens (+20% markup). Now serving in production with stable latency.
7 IQx ultra fast models now live on Language Processing Unit silicon. Sub-100ms TTFT under load.
Credit card required. Predictable monthly windows from $25. Pure PAYG for scale.
Sign up, get an API key, and start calling frontier models in under 60 seconds. Credit card required. $3 is the floor.
Run one command. Paste your API key. IQx Token Factory auto-configures Claude Code, Codex, and every model in your workspace based on the plan you're on — no env files, no boilerplate.
Works on macOS, Linux, and Windows. Node 18+ required.