Tenzro LabsPlatform

Inference API

Call Tenzro Labs models with one API key. Name the model; the Platform routes the call to a machine on the Tenzro Network that serves it, and handles the network's keys, endpoints and methods for you.

Base URLhttps://api.tenzro.xyz/v1
AuthYour Platform API key with the inference permission: Authorization: Bearer tzl_…
Modelmodel: a model id from GET /models
Machinemachine (optional): where to run it. Defaults to tenzro-partner-cluster, the Tenzro Partner Cluster.
BillingFrom your credits, or a rental or subscription the key belongs to. Free on testnet.
OpenAPI/v1/openapi.json

Try it

Send a real request as your account and see what comes back. Each panel shows the same call made with an API key.

No model is serving right now, so there is nothing to try yet.

How a call runs

The models run on operators' machines on the Tenzro Network that Tenzro Labs serves from. Each model kind has its own endpoint, because each is called differently: a conversation, texts to embed, a series to forecast, a prompt to render. The Platform checks your key, finds a machine that is serving the model right now, and makes the call there with the network's native method for that kind. You never handle a node address, a network key or a provider id. Calling a model at the wrong endpoint returns 400 wrong_endpoint naming the right one.

Model kindEndpoint
LanguagePOST /v1/chat/completions
EmbeddingsPOST /v1/embeddings
ForecastingPOST /v1/forecasts
Image generationPOST /v1/images/generations
List modelsGET /v1/models

Models

Model idKindEndpointContextPrice
qwen3.8-27bLanguage, image input/v1/chat/completions256KFree in · Free out / 1M tokens
qwen3.5-4bLanguage/v1/chat/completions128KFree in · Free out / 1M tokens
timesfm-2.5-200mForecasting/v1/forecasts2,048 pointsFree per request
qwen3-embedding-0.6bEmbeddings/v1/embeddings—Free per request
flux2-klein-4bImage generation/v1/images/generations—Free per request

A model is offered only while a machine is serving it. GET https://api.tenzro.xyz/v1/models lists what can be called right now, with each model's endpoint and machines. It needs no key.

json
{
  "object": "list",
  "default_machine": "tenzro-partner-cluster",
  "data": [
    {
      "id": "qwen3.5-4b",
      "object": "model",
      "owned_by": "tenzro-labs",
      "name": "Qwen 3.5 4B",
      "modality": "text",
      "endpoint": "/v1/chat/completions",
      "context_length": 131072,
      "max_output_tokens": 2048,
      "features": ["tools"],
      "machines": [{ "id": "tenzro-partner-cluster", "name": "Tenzro Partner Cluster", "region": "Europe" }]
    }
  ]
}

Language models: chat completions

OpenAI-compatible. Point an OpenAI SDK at the base URL and use your Platform key. Streaming follows the OpenAI server-sent events format; the last chunk carries token usage. max_tokens is capped at the model's max_output_tokens.

FieldType
modelstringRequired. A language model id.
messagesarrayRequired. OpenAI chat messages: role and content.
streambooleanStream the reply as it is generated.
max_tokensintegerLongest reply, up to the model's limit.
temperature, top_p, stop, seedAs in the OpenAI API.
tools, tool_choiceFunction calling, for models with the tools feature.
enable_thinkingbooleanReasoning models only: think before answering (off by default).
machinestringOptional. Defaults to tenzro-partner-cluster.

curl

shell
curl https://api.tenzro.xyz/v1/chat/completions \
  -H "Authorization: Bearer $TENZRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5-4b",
    "messages": [{"role": "user", "content": "Write a haiku about GPUs"}],
    "stream": true
  }'

Python

python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.tenzro.xyz/v1", api_key=os.environ["TENZRO_API_KEY"])

reply = client.chat.completions.create(
    model="qwen3.5-4b",
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={"machine": "tenzro-partner-cluster"},  # optional
)
print(reply.choices[0].message.content)

TypeScript

ts
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.tenzro.xyz/v1", apiKey: process.env.TENZRO_API_KEY });

const stream = await client.chat.completions.create({
  model: "qwen3.8-27b",
  messages: [{ role: "user", content: "Explain proof of stake in two sentences" }],
  stream: true,
});
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");

Embedding models: embeddings

OpenAI-compatible. Vectors are unit length, so cosine similarity is a dot product.

FieldType
modelstringRequired. An embedding model id.
inputstring | string[]Required. Up to 256 texts per request.
dimensionsintegerOptional. A shorter vector, where the model supports it (1024, 768, 512, 256, 128, 64, 32 for Qwen3-Embedding 0.6B).
input_type"document" | "query"Optional. Embed search queries as query; everything else as document (the default).
machinestringOptional. Defaults to tenzro-partner-cluster.

curl

shell
curl https://api.tenzro.xyz/v1/embeddings \
  -H "Authorization: Bearer $TENZRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-embedding-0.6b", "input": ["Tenzro settles in TNZO", "What does Tenzro settle in?"]}'

Python

python
vectors = client.embeddings.create(
    model="qwen3-embedding-0.6b",
    input=["first passage", "second passage"],
    dimensions=512,
)
print(len(vectors.data[0].embedding))  # 512

Response

json
{
  "object": "list",
  "model": "qwen3-embedding-0.6b",
  "data": [{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, ...] }],
  "usage": { "prompt_tokens": 12, "total_tokens": 12 }
}

Forecasting models: forecasts

Send a series, oldest value first, and how many steps to forecast. Quantiles give an uncertainty band.

FieldType
modelstringRequired. A forecasting model id.
historynumber[]Required. At least two values, oldest first; the model uses up to its history limit.
horizonintegerRequired. Steps ahead, up to the model's horizon limit.
quantilesnumber[]Optional. Levels between 0 and 1, e.g. [0.1, 0.5, 0.9].
frequency_secondsintegerOptional. The spacing of the series, when it is regular.
machinestringOptional. Defaults to tenzro-partner-cluster.

curl

shell
curl https://api.tenzro.xyz/v1/forecasts \
  -H "Authorization: Bearer $TENZRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "timesfm-2.5-200m", "history": [112, 118, 132, 129, 121, 135, 148, 148, 136, 119], "horizon": 6, "quantiles": [0.1, 0.5, 0.9]}'

Python

python
import os, requests

r = requests.post(
    "https://api.tenzro.xyz/v1/forecasts",
    headers={"Authorization": f"Bearer {os.environ['TENZRO_API_KEY']}"},
    json={"model": "timesfm-2.5-200m", "history": series, "horizon": 24, "quantiles": [0.1, 0.9]},
)
forecast = r.json()
print(forecast["point"])

Response

json
{
  "object": "forecast",
  "model": "timesfm-2.5-200m",
  "history_used": 10,
  "point": [141.2, 137.9, 139.4, 144.0, 146.1, 143.3],
  "quantile_levels": [0.1, 0.5, 0.9],
  "quantiles": [[128.4, ...], [141.2, ...], [155.0, ...]],
  "generation_time_ms": 84
}

Image models: image generations

The endpoint is live and routes like the others. While no image model is serving, it answers 503 no_capacity; the models list shows when one is.

shell
curl https://api.tenzro.xyz/v1/images/generations \
  -H "Authorization: Bearer $TENZRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "flux2-klein-4b", "prompt": "a lighthouse at dawn, watercolour"}'

Billing

A call through the API or MCP is paid from your credits, or covered by the rental or subscription its key belongs to. (Cortex, for people, runs on credits alone.) Credit charges are made after the call completes: language models by input and output tokens, embeddings by input tokens, forecasting and image models per request. Every price is zero on testnet, and a free model needs no balance. When a priced model is called with an empty balance, the answer is 402 insufficient_credit until you top up.

Errors

Errors use the OpenAI shape, {"error": {"message", "type", "code"}}, with a matching status.

StatusCodeWhat to do
400missing_model · invalid_json · invalid_input · invalid_history · invalid_horizonFix the request as the message says.
400wrong_endpointCall the model at the endpoint the message names.
401missing_api_key · invalid_api_keySend a valid Platform key; it may be revoked.
402insufficient_creditTop up credit.
403key_lacks_inference_permissionCreate a key with the inference permission.
403key_not_scoped_to_machineThe key is pinned to another machine; use that one or an unpinned key.
403subscription_inactive · model_not_in_plan · rental_endedThe plan or rental behind the key does not allow this call.
404model_not_found · machine_not_foundUse an id from GET /models.
502machine_unreachable · machine_errorRetry with backoff.
503no_capacityThe model is not serving on that machine right now; retry later or pick another model.
Retry 502 and 503 with exponential backoff. The other errors will not succeed on retry.