Kindling / DocumentationOpen chat ↗
KINDLING GUIDE

Start with clarity.

One simple way to build with AI. Ask in chat, or connect the API to an inference service you already run.

What is Kindling?

Kindling gives your product one clear interface for AI answers. Choose a response mode for each request. Kindling returns a unified response, so your app can stay focused on the result.

✳

Keep it simple. Use the default mode for everyday questions. Switch to Spark when a request benefits from more deliberate reasoning.

Choose a response mode

Modes are selected per request in chat or with the API’s mode field.

✳

Kindling DEFAULT

Quick, efficient answers for everyday tasks, ideas, and questions. It is the right place to start when you want a helpful response in a breeze.

✳

Spark MORE CONSIDERED

Give complex, nuanced, or important questions a more deliberate pass. Spark is designed to help with careful answers; no mode can guarantee correctness, so check important information.

The modes describe the experience. The service handles the underlying routing and returns one answer.

Use the chat

  1. Open Kindling chat.
  2. Choose Kindling for a quick everyday answer, or Spark for a more considered response.
  3. Write your question and press Enter. Use Shift + Enter for a new line.

You can start a fresh conversation from the sidebar. Suggestions on the welcome screen fill in a starter prompt.

Call the API

The local app serves a non-streaming, OpenAI-compatible chat completions endpoint:

POST /v1/chat/completions
curl http://127.0.0.1:5173/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"mode":"kindling","messages":[{"role":"user","content":"Hello!"}]}'

For a more considered reply, set "mode":"spark". The response uses the standard chat-completion shape with the answer in choices[0].message.content.

JavaScript
const response = await fetch("http://127.0.0.1:5173/v1/chat/completions", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    mode: "spark",
    messages: [{ role: "user", content: "Compare these options…" }]
  })
});
const result = await response.json();
console.log(result.choices[0].message.content);

Connect an inference service

Kindling needs an already-running OpenAI-compatible inference server. Set these environment variables before starting the local server:

PowerShell
$env:KINDLING_UPSTREAM_URL = "http://127.0.0.1:8000/v1/chat/completions"
$env:KINDLING_UPSTREAM_MODEL = "your-existing-default-route"
$env:KINDLING_SPARK_UPSTREAM_MODEL = "your-existing-spark-route"
$env:KINDLING_UPSTREAM_API_KEY = "your-upstream-key" # if required
./.venv311/Scripts/python.exe dashboard/server.py --host 127.0.0.1 --port 5173

Use KINDLING_API_KEY if you want Kindling itself to require a bearer key from API clients. The public mode names are kindling and spark; the two model variables map those names to routes your service already provides.

i

Kindling does not download model files. Requests are forwarded only to the upstream URL you configure.

API reference

GET /healthzReports whether an upstream URL is configured. It does not probe that service.
GET /v1/modelsLists Kindling’s public mode aliases. If API keys are enabled, send a bearer key.
POST /v1/chat/completionsCreates one non-streaming chat completion.

Request fields

  • mode or model: kindling or spark; defaults to Kindling.
  • messages: non-empty array of role and text content objects.
  • stream: must be false; streaming is not available yet.
  • max_tokens: optional integer from 1 to 8192.
  • temperature: optional number from 0 to 2.

Current capabilities

  • Chat text requests and single unified answers.
  • Optional Spark response mode.
  • OpenAI-compatible, non-streaming completion endpoint.
  • Media, document, and file choices in the chat composer currently add local attachment chips only; the API does not yet upload or read those files. Links are shown in the composer and are not fetched by the API.

When inference has not been configured, the API returns 503 inference_not_configured. Configure an existing server using the steps above to enable live answers.