AI API
Endpoints for interacting with AI models through a unified provider layer. These endpoints are served by the backend API service, not the web application. They are available at the backend base URL, which may differ from the web API base URL depending on your deployment.These endpoints are internal backend endpoints and are not exposed through the web application’s
/api routes. All /api/ai/* endpoints require bearer token (API key) authentication — the authenticate middleware is applied to the entire /api/ai router mount. The chat and cost estimation endpoints additionally require a valid subscription plan via header-based plan enforcement. All POST requests must include the Content-Type: application/json header.Health check
Response
status field is healthy when the OpenRouter provider is reachable and degraded when it is not. Only the OpenRouter provider is checked.
Error response
When the provider check fails, the response usesstatus: "error" and includes the error message:
List models
Response
Errors
List models by provider
Path parameters
Response
Select model
Request body
Response
Errors
Chat completion
The chat endpoint uses header-based authentication. Access control is enforced through the
x-user-plan and x-stripe-subscription-id headers. When a database is available, the plan middleware cross-references the x-stripe-subscription-id header against the user’s record in the database to prevent subscription forgery. The plan stored in the database is used instead of the header value. If the database is unavailable, the middleware falls back to header-based validation with a format check on the subscription ID. Admin emails (configured via ADMIN_EMAILS) bypass both plan and subscription requirements.Request headers
The following headers are required for plan enforcement:Request body
Example request
Example request with Algorithm mode
WhenalgorithmMode is enabled, the agent responds using a structured 7-phase format for non-trivial tasks:
Response
Returns a structured response with the following shape:Errors
402 error examples
x-stripe-subscription-id header does not match the subscription stored in the database for the authenticated user:
403 error example
Token quotas
Token quotas are enforced per user on a calendar-month basis. Each chat completion request checks the user’s cumulative token usage for the current month against their plan limit before calling the AI provider. Requests that would exceed the quota are rejected with a429 status and a QUOTA_EXCEEDED error code. The quota resets automatically at the start of each calendar month.
Usage is tracked in the model_metrics table. Each successful and failed chat request logs the model, token counts, latency, and outcome for auditing and quota enforcement.
429 error example
If the database is temporarily unreachable, quota enforcement fails open — the request proceeds without a usage check. A warning is logged server-side.
Plan-based model access
Each subscription plan grants access to a specific set of AI models. The chat endpoint enforces these limits automatically via the plan middleware.Admin users are automatically granted
network-level access regardless of their subscription plan.The plan middleware (
x-user-plan header) enforces model access, skill limits, and A2A message quotas. The provisioning endpoint enforces separate agent creation limits: solo 1, collective 3, label 10, network unlimited. The agent count is per-user across all plans — all active agents count toward the current plan’s limit. The provisioning limits determine how many agents you can create, while the middleware limits in the table above apply to per-request AI model access and skill usage.Model fallbacks
The backend AI service uses a tier-based fallback system. Each tier has a primary model and one or more fallback models. When the primary model is unavailable, times out, or returns an error, the system automatically retries the request using the next fallback model in order. Each model attempt is bounded by a configurable timeout (default 30 seconds) to prevent hangs.
Fallback routing is handled transparently. The response always indicates which model ultimately served the request via the
model field. All tier-based requests are routed through OpenRouter.
Task-based model selection
In addition to provider-level fallbacks, the backend AI service uses tag-based model selection that picks the best available model based on the type of work being performed. When you specify ataskType, the system searches the available OpenRouter models for matching capability tags and selects the first match.
When no model matches the requested task tags, the first available model from the OpenRouter catalog is used as a fallback. All task-based requests are routed through OpenRouter.
Algorithm mode
The chat endpoint supports an optional structured problem-solving mode called PAI Algorithm mode. When enabled via thealgorithmMode parameter, a system prompt is injected that instructs the model to process non-trivial tasks using a 7-phase format:
The Algorithm system prompt is prepended to the
messages array as a system message. If the messages already contain a system message with Algorithm phase markers, the prompt is not duplicated. For simple greetings or acknowledgments, the model skips the 7-phase format and responds naturally.
Algorithm mode is opt-in and does not affect billing or model selection. It only modifies the system prompt sent to the model. Based on Daniel Miessler’s TheAlgorithm v0.2.24.