Use the Bearly API
The Bearly API lets your own code call the models in Bearly with the official OpenAI or Anthropic SDK. Change the SDK’s base URL and API key; your requests, responses, and streaming stay in the format the SDK already uses. Usage is billed to your Bearly account under your plan and credits, the same way chat is. If your code stops reading a streamed response early, the model still finishes it and the full response is billed.
Every request goes only to providers that have agreed to zero data retention: the provider does not keep your prompt or the model’s response after processing them. See zero-data-retention models for what that label means.
Create an API key
Section titled “Create an API key”- Open your account menu and select Developers. API Keys opens. If Developers or API Keys isn’t there, your organization’s policy has turned off API access; ask your team administrator.
- Select Create API Key, or New Key if you already have one. Enter a note that tells you where the key is used, such as “Production app”, then select Create.
- Select Copy on the new key and store it in a secret manager or server environment variable. The examples on this page read it from
BEARLY_API_KEY.
The key acts as your Bearly account until you delete it. Keep it out of browser code, public repositories, and email. If a key may have leaked, delete it under API Keys and create a new one.
Connect an SDK
Section titled “Connect an SDK”| SDK | Base URL | Key |
|---|---|---|
| OpenAI (Chat Completions and Responses) | https://api.bearly.ai/v1 |
Your Bearly API key as the API key |
| Anthropic (Messages) | https://api.bearly.ai |
Your Bearly API key as the API key |
You don’t need an OpenAI, Anthropic, or other provider key.
OpenAI SDK
Section titled “OpenAI SDK”Use the OpenAI SDK for every model that supports Chat Completions or Responses.
import os
from openai import OpenAI
client = OpenAI( base_url="https://api.bearly.ai/v1", api_key=os.environ["BEARLY_API_KEY"],)
response = client.responses.create( model="gpt-6.1-sol", input="In one sentence, what is an API key?",)print(response.output_text)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.bearly.ai/v1", apiKey: process.env.BEARLY_API_KEY,});
const response = await client.responses.create({ model: "gpt-6.1-sol", input: "In one sentence, what is an API key?",});console.log(response.output_text);curl https://api.bearly.ai/v1/chat/completions \ -H "Authorization: Bearer $BEARLY_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "bearly-fast", "messages": [{"role": "user", "content": "Hello"}] }'The same client also works with chat.completions.create for any model that lists Chat Completions in the model table.
Anthropic SDK
Section titled “Anthropic SDK”Use the Anthropic SDK for Claude models.
import os
from anthropic import Anthropic
client = Anthropic( base_url="https://api.bearly.ai", api_key=os.environ["BEARLY_API_KEY"],)
message = client.messages.create( model="claude-haiku-5-5", max_tokens=1024, messages=[{"role": "user", "content": "In one sentence, what is an API key?"}],)print(next(block.text for block in message.content if block.type == "text"))import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://api.bearly.ai", apiKey: process.env.BEARLY_API_KEY,});
const message = await client.messages.create({ model: "claude-haiku-5-5", max_tokens: 1024, messages: [{ role: "user", content: "In one sentence, what is an API key?" }],});console.log(message.content.find((block) => block.type === "text")?.text);curl https://api.bearly.ai/v1/messages \ -H "x-api-key: $BEARLY_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-haiku-5-5", "max_tokens": 1024, "messages": [{"role": "user", "content": "Hello"}] }'Other tools built on these SDKs usually have a base URL setting; point it at the same address and use your Bearly API key.
Streaming, tool calling, and structured output work the way each SDK documents them, with two model limits:
- GPT models,
bearly-pro, andbearly-ultra: for function calling, use the Responses API. On Chat Completions these models accept tools only withreasoning_effort: "none", whichbearly-pro,bearly-ultra, andgpt-6-lunasupport andgpt-6.1-solandgpt-6-astradon’t. - Claude Opus 5.5 and Sonnet 5.5: set
tool_choicetoauto. Forcing a tool (anyor a named tool) returns an error.
Models
Section titled “Models”Use the model name exactly as shown. Each model works only with the APIs listed for it; Claude models, for example, use the Messages API.
| Model | Chat Completions | Responses | Messages |
|---|---|---|---|
bearly-fast |
Yes | ||
bearly-pro |
Yes | Yes | |
bearly-ultra |
Yes | Yes | |
gpt-6.1-sol |
Yes | Yes | |
gpt-6-luna |
Yes | Yes | |
gpt-6-astra |
Yes | Yes | |
claude-opus-5-5 |
Yes | ||
claude-sonnet-5-5 |
Yes | ||
claude-haiku-5-5 |
Yes | ||
grok-4.7 |
Yes | Yes | |
glm-5.3 |
Yes | ||
glm-5.3-flash |
Yes | ||
deepseek-v4-pro |
Yes | ||
deepseek-v4-flash |
Yes | ||
minimax-m3 |
Yes |
This table was last checked on October 7, 2026. For the current list, call GET https://api.bearly.ai/v1/models or client.models.list() in either SDK. The OpenAI-format list shows which APIs each model supports; the Anthropic SDK’s list shows only Messages models.
Bearly models. bearly-fast, bearly-pro, and bearly-ultra are the same models as Bearly Fast, Bearly Pro, and Bearly Ultra in the app. We keep these names working and may change the model behind them as better ones arrive, so they’re the safest choice when you’d rather not track model releases. They always support Chat Completions. bearly-ultra always uses credits.
Your plan and any team policy apply. A model your plan or team doesn’t allow returns an error; Bearly never switches your request to a different model.
Turn API access on or off for your organization
Section titled “Turn API access on or off for your organization”API access is on for every member unless a policy turns it off. Team administrators control it with Developer API access in the policy’s Developer section. A policy set to Disabled by default keeps it off until you switch it on. Members whose policy turns it off don’t see API Keys in the app, and their API requests return a 403 error. See Set organization policies.
What isn’t supported
Section titled “What isn’t supported”- Stored responses. Requests to the Responses API are always sent with storage off, so
previous_response_id, background mode, and retrieving or deleting a past response aren’t available. Send the conversation history with each request, the way you would with Chat Completions. - Provider-hosted tools. Built-in web search, file search, code execution, image generation, and MCP servers hosted by a provider aren’t available, and a request that includes one returns a
400error. Function calling (tools you define and run yourself) works. - Priority processing and deferred requests. Leave
service_tierunset or set it toauto. Grok’sdeferredoption isn’t available. - Other APIs. Images, embeddings, audio, and fine-tuning aren’t available through the Bearly API.
Errors
Section titled “Errors”| Status | Meaning | What to do |
|---|---|---|
400 |
The request uses something the Bearly API doesn’t support, such as a provider-hosted tool or priority processing. | Remove the option named in the error. |
401 |
The key is missing, invalid, or deleted. | Check the key, or create a new one under API Keys. |
403 |
Your team’s policy turns off API access, or doesn’t allow this model. | Ask your team administrator, or pick an allowed model. |
402 |
The model needs a different plan, or your credits or allowance can’t cover the request. | Add credits, change plans, or pick another model. |
429 |
Your account already has 10 API requests running. | Wait for one to finish; the OpenAI and Anthropic SDKs retry automatically. |
404 model_not_found |
The model name is wrong, the model doesn’t support that API, or it’s no longer offered. | Check the model list and the API you called. |
Other errors come from the model provider in its usual format and are passed through unchanged.