Get ready to build!
Please enter a valid work email address.
Please enter a valid work email address.
We collect your name, email, company, and usage data to provision and support your account and provide support. See our Privacy Policy for details and your privacy rights.
By submitting this form you agree to our Terms of Service and Privacy Policy.
Thank you! Your submission has been received!
‍
If you don't see an immediate response from us, please search your inbox for an email from Corvex.ai
Oops! Something went wrong while submitting the form.

Open-weight models. Zero data retention.

OpenAI- and Anthropic-compatible. GLM and DeepSeek hosted in the U.S.

SOC 2 Type II
HIPAA
U.S. GPU fleet
Docs
RequestResponseProcessed in memory on a U.S. GPUGone when the call ends. Nothing stored.
“

“Corvex Token Factory’s API access is a powerhouse! Bringing open models into my coding and research workflow has been remarkably straightforward. It’s incredibly efficient and I’m genuinely excited to build more on it.”

Brayden Lee
PhD Graduate Researcher, Oxford University

“Speed and availability make Corvex Token Factory so reliable that we’ve made it an execution layer in our engineering workflow. We send work and it runs immediately. Knowing that our code isn’t used for training was non-negotiable for the kind of engineering we do, and we trust Corvex to deliver that security.”

Abdelkader Allam
CEO, QoS Lab

“Corvex Token Factory scales effortlessly with demanding, parallel workloads. As a researcher and developer, it gave me the flexibility I need for a range of projects, from distributed systems to research tooling and model evaluation. The documentation and support have been excellent, and it’s been remarkably easy to build on.”

Ricardo Leal
PhD researcher and developer

“Open models from Corvex Token Factory are my go-to when coding in my IDE.”

Alex Haas
Principal, Haas Aero
←
→
DeepSeekZ.ai
DeepSeekZ.ai

Compatible harnesses.

Point the SDK you already use at Corvex Token Factory.
Requests land on U.S. GPUs.

Works with
Claude CodeClaude Code
CodexCodex
ClineCline
AiderAider
01

Drop-in API.

OpenAI- and Anthropic-compatible chat completions.

base_url = "api.openai.com"
base_url = "api.tokenfactory.corvex.cloud"
02

Tuned for TTFT and load.

Stack is tuned for time-to-first-token and sustained tokens per second under concurrency.

1
16
64
Tokens/s by concurrent streams
03

Nothing persisted.

Prompts and completions are processed in memory. Not logged, not stored, not used for training. Usage metadata for billing only. Details in the docs.

Memory
for the life of the call
→
Disk
nothing written

A fine-tuned catalog for optimal outcomes.

Model
Best for
Value¹
$ per million tokens
GLM 5.3
Agentic coding, multi-step reasoning, and long-context work
88% of Opus 5
28% the input cost
18% the output cost
$1.40 input
$4.40 output
DeepSeek V4 Flash 0731
Fast, high-volume workloads with full reasoning at a fraction of the cost
89% of Sonnet 5
7% the input cost
3% the output cost
$0.14 input
$0.28 output
(1) Capability is based on Artificial Analysis Intelligence Index v4.3.2 at maximum reasoning effort: GLM 5.3 (45) vs. Claude Opus 5 (51) and DeepSeek V4 Flash 0731 (34) vs. Claude Sonnet 5 (38). Cost compares Corvex pricing with the list prices of Claude Opus 5 and Sonnet 5, respectively, as of 21 September 2026. List prices exclude cached-input discounts.

Change three values.

A simple migration gets you going.

base_url
Point your client at the Token Factory endpoint.
https://api.tokenfactory.corvex.cloud/v1
API key
Generate a Corvex Token Factory key through the UI.
sk-cor-…
Model name
Enter a model slug from the catalog.
zai-org/GLM-5.3 · deepseek-ai/DeepSeek-V4-Flash-0731
from openai import OpenAI

client = OpenAI(
    api_key="sk-cor-...",  # Token Factory key
    base_url="https://api.tokenfactory.corvex.cloud/v1",
)

resp = client.chat.completions.create(
    model="zai-org/GLM-5.3",  # or deepseek-ai/DeepSeek-V4-Flash-0731
    messages=[
        {"role": "user", "content": "Summarize this diff."}
    ],
)

print(resp.choices[0].message.content)
curl https://api.tokenfactory.corvex.cloud/v1/chat/completions \
  -H "Authorization: Bearer sk-cor-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3",
    "messages": [{"role": "user", "content": "Summarize this diff."}]
  }'
import anthropic

client = anthropic.Anthropic(
    api_key="sk-cor-...",  # Token Factory key
    base_url="https://api.tokenfactory.corvex.cloud",
)

resp = client.messages.create(
    model="zai-org/GLM-5.3",  # or deepseek-ai/DeepSeek-V4-Flash-0731
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Summarize this diff."}
    ],
)

# Models reason before they answer, so the first block is "thinking".
print("".join(b.text for b in resp.content if b.type == "text"))
curl https://api.tokenfactory.corvex.cloud/v1/messages \
  -H "x-api-key: sk-cor-..." \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5.3",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Summarize this diff."}]
  }'

Zero data retention means zero data retention.

Corvex takes data security seriously. With Corvex Token Factory, that means customer data is never logged, stored, or used for training. Workloads run on GPUs in the U.S., SOC 2 Type II and HIPAA certified compliant.

Got questions? We've got answers.

Quick answers on data handling, compliance, migration and support. Can't find what you need? Contact us.

What happens to my data?

Is it really zero data retention?

Can Corvex personnel see my prompt while it's being processed?

Do the model developers receive or train on my data?

Where does inference physically run?

Are you compliant? Can I get a BAA?

How do I get an API key and make my first call?

Can I switch without rewriting my code?

Am I locked in?

What models are supported, and can I request more?

Why does it cost so much less?

How do I get support for Corvex Token Factory?

Do you offer compliance certificates and dedicated support?

Fast, reliable inference. Starts here.

Get your API key and start building with assurance from U.S.-based hardware and zero data retention.