Skip to content
Quancis
Kael, in beta

Models That Check Each Other's Work.

Kael is a system of specialist models that drafts, checks and refines every answer before it reaches you. It is built to catch bugs, security holes and wrong answers early, behind an API you already know how to call.

Overview

The Facts, Up Front.

Everything you need to decide whether Kael fits, before the explanation. Limits, formats, what it accepts and what it does not. Each item below is covered in more detail further down the page.

Model ID
kael-betaThe string you pass as "model".
Status
BetaThings will change as we learn from real use.
Context window
1M tokensInput per request.
Max output
128k tokensPer response.
Input
Text and imagesPDFs, video and other file types are on the way, not in the beta yet.
Output
Text
Languages
EnglishOther languages have not been tested, so they are not listed as supported.
Knowledge cutoff
July 2026Kael is several models working together, so the cutoff can differ between components.
API formats
Chat Completions, Responses API, Anthropic MessagesUse the SDK you already have.
Weights
ClosedAvailable through the API and our apps.
Thinking levels
FiveLow, High and Max today. z-low and z-high are not released yet.
Speed
Accuracy over speedDetails in the Thinking section below.
How It Works

Not One Model. A System Of Them.

Most AI companies ship one model and call it a day. Kael is a system, and we call it the Composite Intelligence System, or CIS. Behind a single API call, fine-tuned language models, small language models, retrieval and other specialist models work together as one: a draft is produced, then checked, then refined, the way a good engineering team works.

You never see any of it. The request and the response look exactly like any other chat model’s: the same endpoints, the same message format, the same SDKs. The system is the part you don’t have to learn.

  1. 01

    Draft

    A first pass at the answer is produced, the same way any single model would attempt it.

  2. 02

    Check

    The draft is reviewed against your request before it ever reaches you, catching what a single pass tends to miss. For code, this is where bugs and security vulnerabilities get caught.

  3. 03

    Refine

    The system corrects what the check found, and only the refined result comes back as the response.

Why a system, not a bigger model

Models keep getting smarter, but not more accurate. The gap between impressive and correct is where bugs, vulnerabilities and confident wrong answers live. We believe the next gains come from systems of models checking one another, not from one ever-larger model, and that this holds well beyond coding.

Checked on the way in and the way out

CIS sees your request and the response before the response is sent. If something harmful is in either one, the system can stop it before it goes anywhere, instead of cleaning up after the fact.

Built on open-weight models, tuned for the system

Kael is not trained from scratch. Its models start from open-weight models and are fine-tuned so each one understands the environment it runs in and works properly inside CIS. The tuning data came from an earlier internal version of Kael and was not meant to add knowledge. It is the same idea as tuning a model to call tools more reliably. The weights are not released.

Your system prompt stays yours

A system with its own internal instructions and skills has a problem a single model doesn’t: your system prompt and ours can compete for authority. It was the hardest problem we solved, and the reason we tuned the models themselves instead of only wrapping them. You write your system prompt the way you would for any other model.

Efficient inside

Running a whole system costs more tokens inside than running one model, and a lot of them are spent before an answer comes out. We put real effort into optimising CIS to keep output quality as high as possible while keeping that internal cost down.

Strengths

Where Kael Is Strongest.

Kael is built for people who want accuracy more than speed, and it shows most in technical work. These are the jobs it is tuned for.

Coding
Kael is built to write code with fewer bugs, and to catch the bugs that are already there. The check step reads a draft against your request before you see it, so a mistake a single pass would have shipped is more likely to be caught first.
Security
It catches security vulnerabilities in existing code, and it aims not to introduce them in the first place. Both matter: a model that only finds problems after writing them has already cost you a review cycle.
Software engineering
Real engineering tasks, like the distributed-cache refactor in the example below, not just single functions. These are the jobs where a small wrong assumption early on spreads through everything built on top of it.
Agentic work
Long, multi-step tasks where the model acts on its own earlier output. Accuracy matters more here than anywhere, because one wrong step is carried into every step after it.
Math
Multi-step reasoning where the answer has to hold up, not just look plausible. Higher thinking levels give Kael more room to work a problem through before it answers.

Is Kael the right fit?

It is not built for everything, and we would rather say so plainly. Kael can write a story, but that is not what it is tuned for.

Reach for Kael when

  • You care more about a correct answer than a fast one.
  • The work is code review, debugging, refactoring or security review.
  • An agent will act on the output, and a wrong step is expensive.
  • You already call a chat API and want to change as little as possible.
  • Your work is in English.

Look elsewhere when

  • You need the fastest possible response. Kael thinks longer than most models, on purpose.
  • The work is writing-first: fiction, marketing copy, anything where style is the point.
  • You need to send PDFs, video or other file types today. The beta takes text and images.
  • You need a language other than English. Other languages have not been tested.
  • You need weights you can download or run yourself. Kael is closed.
Thinking

Accuracy Takes Time. Here Is How Much.

Kael is slower than most models on purpose. It has five thinking levels, and you choose how much thinking a request gets. Pick a level to see what it costs and where it fits.

With thinking on, at any level, Kael takes about 2.2 times as long to think as an average AI model. Once it starts writing, the answer streams out between 220 and 340 tokens per second. With thinking switched off completely, the first word usually arrives within 0.7 to 3 seconds, at roughly 90 to 140 tokens per second.

Thinking level

Low

Default

The default. Thinking is on, at the lowest of the five levels.

Input
$2per 1M tokens, at up to 80% cache hit
Output
$6per 1M tokens

Good for: Everyday coding and math, when you want Kael’s checking without the longest wait.

Thinking on, at any level

Thinking takes about 2.2× as long as an average AI model’s. Once it starts writing, the answer streams out fast.

Thinking off

With auto-thinking off too, the first word usually arrives in 0.7 to 3 seconds.

Thinking on220–340 tokens/sec
Thinking off90–140 tokens/sec
Pricing

Pay For What You Use.

Pricing is usage-based, per million tokens. The three levels you can use today cost the same. The Z levels, when they launch, cost more because they send requests through a different internal route that produces better results.

Kael pricing by thinking level, per 1M tokens
Thinking levelStatusInputOutput
LowDefault$2$6
HighAvailable$2$6
MaxAvailable$2$6
z-lowNot released yet$6$15
z-highNot released yet$6$15

Prices are per 1M tokens. Input prices are quoted at up to an 80% cache hit rate. See the console for the exact numbers.

Benchmarks

Independent Results First.

We haven’t published benchmark numbers for Kael yet, and we aren’t publishing numbers we ran ourselves. Independent evaluations, including the Artificial Analysis Intelligence Index, are in progress. When results come in, we plan to share them here.

Until then, the best test is your own. Because Kael is a base-URL change, you can point the SDK you already use at it and run your hardest prompts in a few minutes.

Your data

What We Keep, And For How Long.

We don’t need your data. Kael isn’t trained on it, so there is nothing for us to gain from holding on to it. Here is exactly what we do keep.

Normal requests
Held for up to 15 minutes so cache hits can be calculated. After that, they are not stored in a database.
Flagged requests
If the safety system flags a request or a response as harmful, we keep it until our support team has reviewed it. We do this to improve the safety system, and to keep a record in case of a dispute, for example when an account has been blocked or suspended and the account holder contacts us.
Training
Your requests are not used to train Kael. Its models start from open-weight models and are tuned on data generated by an earlier internal version of Kael.

This is a plain-language summary. The Privacy Policy is the binding document.

Integration

Swap The Base URL.

No custom SDK and no migration. Kael accepts Chat Completions, the Responses API and the Anthropic Messages format, so the SDK you already have keeps working. For OpenAI-style requests, point it at our base URL, use the model name kael-beta, keep your existing auth pattern, and everything downstream stays the same.

https://api.quancis.space/v1
import openai

# Quancis is 100% drop-in compatible with the OpenAI SDK.
client = openai.OpenAI(
    api_key="sk-quan-...",
    base_url="https://api.quancis.space/v1",
)

response = client.chat.completions.create(
    model="kael-beta",
    messages=[
        {"role": "system", "content": "You are a helpful coding assistant."},
        {"role": "user", "content": "Refactor our distributed cache..."},
    ],
)

print(response.choices[0].message.content)

This covers the shape of a request. Request and response schemas, streaming, error codes and every parameter live in the full docs, which this page isn't trying to replace.

Full docs and code samples →
Where to use it

One System, Three Ways In.

The same Kael is behind all three. Pick the one that matches how you work.

The API
For developers and teams who already call a model from their own software. This is the page you are on.Developer platform →
Quan Harness
Our coding agent, built on Kael. It works inside a real project environment with a file system, live preview and git.About Harness →
Quan Chat
Kael in a conversation, with deep reasoning and web access you can switch on from the composer.About Chat →
FAQ

Questions People Ask.

The things developers ask first, answered straight. If something here is missing, the docs go deeper.

Is Kael one model?

No. Kael is a system, which we call the Composite Intelligence System (CIS). Fine-tuned language models, small language models, retrieval and other specialist models work together behind a single API call.

From the outside it looks like one model: you send a request and get a response.

Will it work with the SDK I already use?

Yes. Kael accepts Chat Completions, the Responses API and the Anthropic Messages format, so existing SDKs and tools work without a custom client.

For OpenAI-style requests you change the base URL and the model name (kael-beta) and keep the rest of your code.

Why is Kael slower than other models?

It does more work per answer. Every response goes through a draft, a check and a refinement before it reaches you. With thinking on, at any level, Kael takes about 2.2 times as long to think as an average AI model.

That is a deliberate trade: we chose accuracy over speed. If you need the fastest response, turn thinking off. With thinking and auto-thinking both off, the first word usually arrives in 0.7 to 3 seconds.

Which thinking level should I use?

Start with Low, the default. Move up to High or Max when a task is hard enough that being right matters more than being quick, such as a security review or a large refactor.

z-low and z-high are the highest levels and are not released yet.

Can Kael read images, PDFs and video?

Images, yes. Kael takes text and images as input and returns text.

PDFs, video and other file types are on the way, but they are not available in the beta.

Which languages does Kael support?

English. Other languages have not been tested yet, so we don’t list them as supported.

What is the context window?

1 million tokens of input and up to 128,000 tokens of output per request.

What is Kael’s knowledge cutoff?

July 2026. Because Kael is several models working together, the edge of knowledge can differ between components.

Does Kael train on my data?

No. We don’t use your requests to train Kael. Its models start from open-weight models and are tuned on data generated by an earlier internal version of Kael, so nothing in that process needs customer data.

How long do you keep my requests?

Up to 15 minutes, so cache hits can be calculated. After that, requests are not stored in a database.

The exception is anything our safety system flags as harmful. Those are kept until our support team has reviewed them. The data section above and the Privacy Policy cover the details.

Can I download or self-host Kael?

No. Kael is closed: the weights are not released. It is available through the API and our apps.

Is Kael good at creative writing?

It can write stories and other creative text, but that is not what it is tuned for. Kael is strongest at coding, security, software engineering, agentic work and math.

If writing is your main use case, Kael is probably not the best fit today.

How is Kael priced?

Per million tokens. At the Low, High and Max thinking levels it is $2 for input, at up to an 80% cache hit rate, and $6 for output.

The Z levels will be $6 for input and $15 for output when they launch. The pricing section and the console have the exact numbers.

Where are the benchmarks?

We haven’t published any. We don’t want to grade ourselves, so we are waiting on independent evaluations, including the Artificial Analysis Intelligence Index, and we plan to share the results here once they are in.

When will z-low and z-high be available?

We don’t have a date yet. They are the highest thinking levels. They send requests through a different internal route that costs more and produces better results, priced at $6 for input and $15 for output per million tokens.

Start at the default thinking level and move up when a task needs it. Usage-based pricing, with the exact numbers on the console.