An engine is a model. A source is a place that serves it. Quorum picks the cheapest healthy source for each call.
An engine is a model. A source is a place that serves it: a provider, an endpoint and a price. One engine can have several sources, and adding a source gives a model another way to be served.
A Mode names the engines for its seats. Quorum decides which source serves each one.
For each seat, Quorum considers the sources that carry the Mode's engine and can take traffic.
Real cost includes the provider's own markup and per-request charge, so a source that looks cheap per token but adds a fee is ranked on the total.
A source with no price, or a price of zero, ranks below every priced source.
If the chosen source fails, Quorum tries the other sources of the same engine first. After a timeout or a provider-wide error, sources on other providers go first, since the same provider would likely fail again. If no source for that engine can run, the seat moves to a backup the Mode approved: the backup it set for that seat, or another engine from the pool a dynamic Mode approved. When a Mode is resolved, the seat stays inside its approved engines.
If neither the engine nor its backup can run, the seat is refused and the receipt says why: unavailable: reason. See Reading a receipt.
A new model enters the catalog with an exact published price and a source that passes a probe. Once live, the Mode builder can seat it, and a Mode that draws from a pool can pick it up. A Mode that names its engines keeps them until its maker changes them.
A probe job is scheduled every 15 minutes. It re-tests the lowest-scoring and longest-unchecked sources, so a recovered source returns to ranking and a failing one drops out. A sync scheduled once a week checks each provider for new and retired models.