Reduce inference spend without sacrificing speed, security, or reliability.



















[X]M+ tokens/min throughput
Production-grade reliability
Autoscaling — no hard limits on burst
Sub-[X]ms time to first token at p50
Multi-region availability
Kimi K2
SWE-bench score comparable to frontier models; strong agentic coding
GLM 5.2
Benchmark performance comparable to Claude Opus on select evals
DeepSeek V4 Flash
Compact, fast, full chain-of-thought support at lower cost
Kimi K2
$[X.XX]
$[X.XX]
MiniMax 2.5
$[X.XX]
$[X.XX]
GLM 5.2
$[X.XX]
$[X.XX]
DeepSeek V4 Flash
$[X.XX]
$[X.XX]
Dedicated capacity with guaranteed throughput. Volume pricing available.
Sign up at [link], generate an API key from your dashboard, and point your existing OpenAI-compatible client at our endpoint. Most developers are making their first call within five minutes of signing up. No credit card required to start.
Yes. Token Factory uses the same API structure as OpenAI's chat completions endpoint. In most cases, you change one line — the base_url — and your existing code runs on Corvex infrastructure. If you run into anything that doesn't work as expected, let us know in Discord.
Our model catalog includes Kimi K2, GLM 5.2, and DeepSeek V4 Flash, with more being added regularly. You can view the full catalog in the console and request models you'd like to see via [request link] or in our Discord.
[PLACEHOLDER — rate limits for the shared inference tier to be confirmed.] Default rate limits apply to shared inference. If your workload needs higher throughput, reach out to discuss Reserved Endpoints or contact us at [email].
Our per-token pricing is listed publicly on our pricing page — no quote required, no contact form. We keep pricing updated as the market moves.
Not yet — this is on our roadmap. Join our Discord to stay informed on availability. [CONFIRM — replace if fine-tuned model deployment is supported.]
Token Factory provides the inference layer — embeddings and completions — that most RAG pipelines need. Pair it with your vector store of choice (Pinecone, Weaviate, pgvector, etc.) and our embeddings models to build retrieval-augmented workflows. See our docs for a quickstart guide: [docs link]
Yes. Token Factory runs on Corvex's NVIDIA-certified GPU infrastructure with autoscaling and [PLACEHOLDER]M+ tokens/min throughput. If you have specific throughput or latency requirements, talk to us about Reserved Endpoints.
Reserved Endpoints — dedicated capacity with guaranteed throughput — is coming soon. If you have immediate dedicated capacity needs, contact us at [email].
Corvex is SOC 2 certified and HIPAA compliant. Enterprise SLAs and dedicated support are available — contact our team to discuss your requirements: [link to Talk to Us page].
Get your API key and start building — no credit card, no sales call, no approval queue. Just you and the models.