API reference
Chat completions
/v1/chat/completionsExample
Works with any OpenAI SDK. Set the base URL to https://api.sweetrouter.com/v1 and the model to sweet-character-1.
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.sweetrouter.com/v1",
api_key=os.environ["SWEETROUTER_API_KEY"],
)
reply = client.chat.completions.create(
model="sweet-character-1",
messages=[
{"role": "system", "content": "You are Mia, a cheerful barista who loves bad puns."},
{"role": "user", "content": "Hi Mia! What should I order today?"},
],
max_tokens=300,
)
print(reply.choices[0].message.content)Request body
| Field | Type | Description |
|---|---|---|
modelRequired | string | Always sweet-character-1. |
messagesRequired | array | The conversation, oldest first. Each item has a role (system, user or assistant) and text content. Up to 200 messages and 60,000 characters, with at least one user message. |
max_tokensOptional | integer | Longest reply allowed, 1 to 4,096. Default 1,024. max_completion_tokens also works. |
temperatureOptional | number | 0 to 2. Higher gives more varied replies. |
top_pOptional | number | 0 to 1. An alternative to temperature; change one, not both. |
frequency_penalty / presence_penaltyOptional | number | -2 to 2. Positive values reduce repetition. |
stopOptional | string | string[] | Up to 4 sequences where the reply stops. |
streamOptional | boolean | Default false. Set true to stream the reply. See Streaming. |
stream_options.include_usageOptional | boolean | When streaming, set true to get token usage in a final chunk. |
userOptional | string | Optional id for your end user, kept with the call in your logs. Up to 256 characters. |
Response
{
"id": "chatcmpl-cmg8x2k0d0003",
"object": "chat.completion",
"created": 1791105133,
"model": "sweet-character-1",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "A latte, obviously. I'd never steer you wrong, that would be a grave misdeed... or a grave mis-bean." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 41, "completion_tokens": 27, "total_tokens": 68 },
"sweetrouter": { "job_id": "cmg8x2k0d0003", "cost_usd": "0.00012" }
}usage shows the tokens you're charged for. sweetrouter.job_id identifies the call in your console logs, and sweetrouter.cost_usd is its exact cost.
Multi-turn conversations
The API is stateless, like OpenAI's: it doesn't remember earlier calls. To continue a conversation, keep the messages in your app and send them all with each request: the system message, the earlier user and assistant turns, then the new user message.
history = [{"role": "system", "content": "You are Mia, a cheerful barista."}]
def say(text):
history.append({"role": "user", "content": text})
reply = client.chat.completions.create(model="sweet-character-1", messages=history)
answer = reply.choices[0].message.content
history.append({"role": "assistant", "content": answer})
return answer
say("Hi Mia!")
say("What did I just say to you?") # Mia remembers, because the history is sent againA request can hold up to 200 messages and 60,000 characters, and longer histories cost more input tokens. Keep the system message and the most recent turns:
MAX_TURNS = 40 # keep the last 40 user and assistant messages
def build_messages(system_prompt, history, user_text):
recent = history[-MAX_TURNS:]
return [{"role": "system", "content": system_prompt}, *recent, {"role": "user", "content": user_text}]Streaming
Set stream: true to receive the reply as it's written. The stream is OpenAI's chunk format: a first chunk with the role, one chunk per piece of text, a final chunk with finish_reason, an optional usage chunk, then data: [DONE].
stream = client.chat.completions.create(
model="sweet-character-1",
messages=[{"role": "user", "content": "Tell me a short story about bread."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Never call the API from a browser, because that would expose your key. Call it from your server and pass the stream through:
// app/api/chat/route.ts (Next.js). Your key stays on the server.
import OpenAI from "openai"
const client = new OpenAI({
baseURL: "https://api.sweetrouter.com/v1",
apiKey: process.env.SWEETROUTER_API_KEY,
})
export async function POST(request: Request) {
const { messages } = await request.json()
const stream = await client.chat.completions.create({
model: "sweet-character-1",
messages,
stream: true,
})
return new Response(stream.toReadableStream(), {
headers: { "Content-Type": "text/event-stream" },
})
}- If the call fails before the first word, you get a normal JSON error with an HTTP status.
- If it fails mid-reply, the last chunk carries an
errorobject and you aren't charged. - If your client disconnects, the reply still finishes and is charged for the tokens used.
Compatibility
response_format aren't supported, and n is always 1.See also
- Character chat guide: personas and recommended settings.
- Errors: every error code and when to retry.
- Limits: request sizes and rate limits.