Start with clarity.
One simple way to build with AI. Ask in chat, or connect the API to an inference service you already run.
What is Kindling?
Kindling gives your product one clear interface for AI answers. Choose a response mode for each request. Kindling returns a unified response, so your app can stay focused on the result.
Keep it simple. Use the default mode for everyday questions. Switch to Spark when a request benefits from more deliberate reasoning.
Choose a response mode
Modes are selected per request in chat or with the API’s mode field.
Kindling DEFAULT
Quick, efficient answers for everyday tasks, ideas, and questions. It is the right place to start when you want a helpful response in a breeze.
Spark MORE CONSIDERED
Give complex, nuanced, or important questions a more deliberate pass. Spark is designed to help with careful answers; no mode can guarantee correctness, so check important information.
The modes describe the experience. The service handles the underlying routing and returns one answer.
Use the chat
- Open Kindling chat.
- Choose Kindling for a quick everyday answer, or Spark for a more considered response.
- Write your question and press Enter. Use Shift + Enter for a new line.
You can start a fresh conversation from the sidebar. Suggestions on the welcome screen fill in a starter prompt.
Call the API
The local app serves a non-streaming, OpenAI-compatible chat completions endpoint:
curl http://127.0.0.1:5173/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"mode":"kindling","messages":[{"role":"user","content":"Hello!"}]}'For a more considered reply, set "mode":"spark". The response uses the standard chat-completion shape with the answer in choices[0].message.content.
const response = await fetch("http://127.0.0.1:5173/v1/chat/completions", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
mode: "spark",
messages: [{ role: "user", content: "Compare these options…" }]
})
});
const result = await response.json();
console.log(result.choices[0].message.content);Connect an inference service
Kindling needs an already-running OpenAI-compatible inference server. Set these environment variables before starting the local server:
$env:KINDLING_UPSTREAM_URL = "http://127.0.0.1:8000/v1/chat/completions"
$env:KINDLING_UPSTREAM_MODEL = "your-existing-default-route"
$env:KINDLING_SPARK_UPSTREAM_MODEL = "your-existing-spark-route"
$env:KINDLING_UPSTREAM_API_KEY = "your-upstream-key" # if required
./.venv311/Scripts/python.exe dashboard/server.py --host 127.0.0.1 --port 5173Use KINDLING_API_KEY if you want Kindling itself to require a bearer key from API clients. The public mode names are kindling and spark; the two model variables map those names to routes your service already provides.
Kindling does not download model files. Requests are forwarded only to the upstream URL you configure.
API reference
GET /healthzReports whether an upstream URL is configured. It does not probe that service.GET /v1/modelsLists Kindling’s public mode aliases. If API keys are enabled, send a bearer key.POST /v1/chat/completionsCreates one non-streaming chat completion.Request fields
modeormodel:kindlingorspark; defaults to Kindling.messages: non-empty array of role and text content objects.stream: must be false; streaming is not available yet.max_tokens: optional integer from 1 to 8192.temperature: optional number from 0 to 2.
Current capabilities
- Chat text requests and single unified answers.
- Optional Spark response mode.
- OpenAI-compatible, non-streaming completion endpoint.
- Media, document, and file choices in the chat composer currently add local attachment chips only; the API does not yet upload or read those files. Links are shown in the composer and are not fetched by the API.
When inference has not been configured, the API returns 503 inference_not_configured. Configure an existing server using the steps above to enable live answers.