Inference settings
The inference gateway serves chat, embeddings and transcription from one endpoint. This page explains the settings that control it, how your own provider keys work, and which numbers are estimates. The request and response formats of the gateway are in the inference API reference.
The gateway and its keys
The gateway lives under /api/inference/v1. It accepts an API key that has the scope inference:run. That scope runs inference and nothing else. A key with only inference:run cannot read or change a setting, a key, a guardrail or a harness.
The settings on this page live under /api/v1/ai. They need a key with the scope read, write or admin. The organization always comes from the key. No request field names an organization.
| Parameter | Type | Description |
|---|---|---|
| read | scope | Read the config, the model list, usage, quota, BYOK rows, guardrails and prompts. |
| write | scope | Change the config, guardrails and prompts. Run a prompt. |
| admin | scope | Save, change or delete a provider key. Delete a harness. |
| inference:run | scope | Call the gateway. No access to the settings. |
Settings
A setting is either enforced or not enforced. The gateway reads an enforced setting on each request. A setting that is not enforced has no reader. The platform refuses to save a value for it, so a stored value never looks like a rule that does not exist.
Read the current state from the server. The call returns each setting and an enforcement map with one entry per setting.
| Parameter | Type | Description |
|---|---|---|
| allow_list | enforced | The models your organization may call. A request for any other model is refused. |
| default_model | enforced | The model used when a request names none. The same allow-list and quota rules apply to it. |
| fallback_model | enforced | The model tried once when the platform providers fail. The response header x-inference-fallback names the model that was asked for. |
| prevent_overrides | enforced | A chat request that names a model other than the default is refused with 403 model_override_not_allowed. It needs a default model. |
| inference_enabled | enforced | The switch for the whole gateway for your organization. |
| output_index_enabled | enforced | Whether the platform indexes your served output so that a stripped assertion can still be identified. |
| retention_days | enforced | How long the output index keeps an entry. The platform window of 90 days is the upper limit. Usage records are billing records and do not follow this setting. |
| monthly_budget_cents | enforced | The gateway checks the organization budget before each chat request. The figure is an estimate at the plan overage rates. Null means no budget. |
| include_byok_spend | enforced | Whether tokens that go through your own provider keys count against the organization budget. |
| cost_tier | not enforced | Not enforced. The platform has no model price tiers. Only the value balanced is accepted. |
| provider_sort | not enforced | Not enforced. Each model has one platform provider, so there is nothing to order. Only the value balanced is accepted. |
| capture_content_enabled | not enforced | Not enforced. The platform does not capture prompts or responses. Only false is accepted. |
Your own provider keys
You can save a key from your own provider account. The vendor bills you for those requests. NuPaaS does not bill the tokens and adds no fee. The gateway shows only the last four characters of a saved key. It never returns the key.
Each key has a priority. The priority decides when the gateway uses it.
| Parameter | Type | Description |
|---|---|---|
| always | priority | The only route for models the key serves. A failure is returned as it is. The platform providers and the fallback model are never called. |
| prefer | priority | Tried first. A rate limit, an outage or a refused key moves the request to the platform providers. |
| fallback | priority | Tried after the platform providers are used up. |
The gateway can route to these providers: OpenAI, Mistral, DeepSeek, Groq, Together, Fireworks, Cerebras, DeepInfra and xAI. It can also route to any OpenAI-compatible HTTPS endpoint that you give as a custom provider. A key for any other provider is refused when you save it. A key counts as routable only when the provider lists the model for your key.
Guardrails and prompts
A guardrail is a policy for API keys: allowed models, a monthly budget, personal-data redaction, retention and moderation. One guardrail can be the default for keys that have none. You can assign a guardrail to one key. Assigning none returns the key to the default policy.
A saved prompt is a template that you can run through the gateway. A run passes the same checks as a chat request and uses tokens. The result of a run is not stored.
Budgets, quotas and rate limits
Three limits can stop a request. Each one has its own refusal.
| Parameter | Type | Description |
|---|---|---|
| Plan quota | hard limit | The token limit of your plan for the billing period. A request over it is refused. |
| Key budget | estimate | A budget on one API key, from its guardrail. The gateway prices the tokens at the plan overage rates and compares them with the budget. |
| Organization budget | estimate | The monthly_budget_cents setting. It is priced the same way and answers 429 inference_org_budget_exceeded. |
| Rate limit | per organization | A limit on requests per period. Chat requests and prompt runs count against it. |
Provenance marks
The gateway can mark generated output with a signed assertion in the field platform_provenance. A prompt run returns the mark unchanged. If a call returns no generated output, the field is null. Verify a mark with the public key and the verify endpoint. The reference explains the claims and the limits of a mark.
Usage
Usage is the input and output tokens of the billing period. Pass a harness run id to see only the rows that the gateway tagged with that run. A harness key sends the run id in a header that the gateway sets. The tag counts only calls that carried it.
From the SDK and the CLI
These calls use the public API, so they follow the scopes above.
import { PlatformClient } from "@type-driven/platform-sdk";
const client = PlatformClient.fromEnv();
const config = await client.ai.getConfig();
console.log(config.enforcement.cost_tier); // "not_enforced"
await client.ai.updateConfig({ default_model: "mistral-small-latest" });
await client.ai.putByokKey("openai", {
label: "main",
api_key: process.env.OPENAI_API_KEY!,
priority: "prefer",
});
const usage = await client.ai.getUsage({ harness_run_id: "RUN_ID" });
console.log(usage.total_tokens, usage.tagged_rows);platform inference config get
platform inference config set --default-model mistral-small-latest
platform inference byok list
platform inference byok put openai --label main --key-file ./openai.key --priority prefer
platform inference byok delete openai --yes
platform inference guardrails list
platform inference prompts run --user "Write one full sentence."
platform inference usage --harness-run RUN_IDThe key file holds the provider key. The CLI never takes a key as a command argument, so the key stays out of your shell history.