ModelRelay

ModelRelay is a hosted model API. Call leading models through one API, and ModelRelay meters every request to the customer who made it, so you can charge your users for the AI they use.

  • Models — Choose the right model. The models page ranks models by speed, cost, and intelligence, and every model listed there is callable through the API.
  • API — One request shape for every model. Use the provider-neutral Responses API, or point an existing OpenAI or Anthropic SDK at the compatible endpoints by changing the base URL.
  • Monetize — Attribute each request to an end user with customer tokens, set limits and pricing with tiers, and bill through Stripe with customer billing.

Provider rates pass through; the models page shows each model’s price and any ModelRelay fee. New accounts get a one-time signup credit with no card required. Sign in with GitHub or Google, and pay by card or USDC.

Quickstart

Create an account at modelrelay.ai. Signing up creates a default project and its first API key; you can create more keys in the dashboard or with mrl. Then make a request:

curl https://api.modelrelay.ai/v1/chat/completions \
  -H "Authorization: Bearer $MODELRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "messages": [{"role": "user", "content": "What is the capital of France?"}]
  }'

Or use the native Responses API:

curl https://api.modelrelay.ai/api/v1/responses \
  -H "X-ModelRelay-Api-Key: $MODELRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "input": [{"type": "message", "role": "user", "content": [{"type": "text", "text": "What is the capital of France?"}]}]
  }'

Or the mrl CLI:

brew install tensor-systems/tap/mrl
mrl auth login --web
mrl keys create --name laptop
mrl "What is the capital of France?" --model claude-sonnet-5-5

Full getting started guide →

Endpoints

Endpoint Use
POST /api/v1/responses Provider-neutral Responses API
POST /v1/chat/completions OpenAI Chat Completions format
POST /v1/responses OpenAI Responses format
POST /v1/messages Anthropic Messages format
POST /v1/audio/speech Text to speech
POST /v1/audio/transcriptions, POST /v1/stt, GET /v1/stt Speech to text, including streaming transcription over WebSocket
POST /api/v1/images/generate Image generation (no image model is listed yet; see Images)
POST /v1/decisions Typed decisions about application state (beta)
GET /v1/models, GET /api/v1/models Model catalog

Documentation

  • Getting Started — Create an account, get an API key, and make a first request.
  • Authentication — Secret keys, customer tokens, and account login.
  • Core Features — Streaming, tool use, structured output, reasoning effort, and error handling.
  • SDK Compatibility — Use the OpenAI and Anthropic SDKs with ModelRelay.
  • API Reference — Every endpoint, request field, and error.
  • Billing — Customer tokens, tiers, usage, and Stripe billing.
  • CLI — mrl for requests, API keys, and project administration.