Change the base URL and set the model to a Mode ticker. What carries over from the OpenAI chat completions API and what Quorum adds to the response.
Quorum serves the chat completions shape at POST /v1/chat/completions. Existing OpenAI client code runs against it after two changes: the base URL, and the model value.
from openai import OpenAI
# Base URL: was https://api.openai.com/v1. API key: a Quorum key.
# Timeout: a deliberation runs several models.
client = OpenAI(
base_url="https://www.quorum.dog/v1",
api_key="qk_live_...",
timeout=300,
)
# Model: was "gpt-4o", now a Mode ticker.
r = client.chat.completions.create(
model="QRUM:STAN",
messages=[{"role": "user", "content": "Should we shard this table now or after launch?"}],
)
print(r.choices[0].message.content)import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://www.quorum.dog/v1',
apiKey: process.env.QUORUM_API_KEY,
timeout: 300_000,
});
const r = await client.chat.completions.create({
model: 'QRUM:STAN',
messages: [{ role: 'user', content: 'Should we shard this table now or after launch?' }],
});
console.log(r.choices[0].message.content);A ticker has the form COMPANY:MODE. model names a Mode, and a Mode decides which engines sit in the seats and how the panel runs. For example, QRUM:STAN is Standard Mode. GET /v1/models lists the Modes your key can call, with pricing.
| OpenAI | Quorum |
|---|---|
POST /v1/chat/completions | Same path and method. |
Authorization: Bearer | Same header, with a Quorum key (qk_live_... or qk_test_...). |
messages with role and content | Same. Send content as a string. |
id, object, created, model | Same fields. object is chat.completion. |
choices[0].message.content | The answer. |
choices[0].finish_reason | stop, or length when the answer was cut off. |
usage.prompt_tokens, completion_tokens, total_tokens | Same fields. |
usage.completion_tokens_details.reasoning_tokens | Same field. null when the provider reported no breakdown. |
GET /v1/models | Same path. Returns the Modes your key can call. |
| Error body | { "error": { "message", "type", "code", "param" } }, the OpenAI shape. |
Idempotency-Key header | Same header. A repeat of a call that succeeded replays the answer with quorum.replayed: true. The replay matches on the key alone, so send a new key for each request. A call that ended capped is not replayed, so a retry runs and bills again. If you omit the header, Quorum derives a key from your organization, key, model and messages, so an SDK retry of the same request replays instead of billing twice. |
The request fields Quorum reads are model, messages, tier, quorum.max_cost_usd and stream. A Mode's own configuration sets how its panel runs.
The response carries a quorum block next to the standard fields. Read it to see how the seats came out.
{
"id": "…",
"object": "chat.completion",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 41, "completion_tokens": 612, "total_tokens": 653 },
"quorum": {
"request_id": "…",
"mode": "QRUM:STAN",
"rounds": 1,
"seats": 3,
"seats_answered": 3,
"converged": true,
"remaining_friction": null,
"contested_passage": null,
"capped": false,
"cap_credit_usd": 0,
"billed_usd": 0.10,
"latency_ms": 26418,
"engines": ["<seat model id>", "<seat model id>", "<seat model id>"]
}
}The values in the sample are illustrative. A replay of an earlier call carries a shorter quorum block with replayed: true: no converged, remaining_friction, contested_passage or billed_usd. Check replayed first.
| Field | Meaning |
|---|---|
request_id | The id for GET /v1/receipts/{request_id}. |
converged | false when the seats did not settle on one position. null when no judge read the round, as on the express lane. |
remaining_friction | factual_conflict when the judge flagged a factual conflict between seats, otherwise null. |
contested_passage | An object { seat, quote, why }, or null when no single passage carried the split. |
seats, seats_answered | Seats convened and seats that answered. |
billed_usd, surcharge_usd, platform_fee_usd, cap_credit_usd | The bill for this call and its parts. billed_usd is the token charge, the platform fee and the surcharge, minus cap_credit_usd. |
capped | true when the call went over its ceiling. The call is charged in full and the amount over the ceiling is credited back in cap_credit_usd. The answer is whole. |
engines | The model ids of the seats that ran. |
substitutions | Seats that ran a backup engine, and what ran instead. |
stream: true returns 400 stream_not_supported. Leave stream unset. Clients that stream by default need that turned off.413 context_length_exceeded.{"quorum": {"max_cost_usd": 1.00}} to set it lower for one call. A request can lower the ceiling and cannot raise it.Origin header returns 403 browser_origin_not_allowed. Call it from your backend.POST /v1/estimate first. It is free, takes the same body and returns the Mode's surcharge for the question and a verdict on whether a panel is expected to help. The bill adds the token charge and the platform fee."tier": "experience" (Express), "plus" (Foundation) or "pro" (Frontier) to choose which engine tier fills the seats for one call. The default is plus.An agent that called OpenAI on every step can keep that model for routine steps and call Quorum at decision points. The escalation router shows how to choose per question, and Agent frameworks wraps the call as a tool.