Why Kael Is Slower, On Purpose
Kael takes longer to answer than most models. That is not a limitation we are working around. It is the trade we chose: accuracy over speed.
With thinking on, at any level, Kael takes about 2.2 times as long to think as an average AI model. Once it starts writing, the answer streams out at between 220 and 340 tokens per second. If you switch thinking off entirely, including auto-thinking, the first word usually arrives within 0.7 to 3 seconds, at roughly 90 to 140 tokens per second.
You choose how much thinking a request gets. Kael has five thinking levels. Low is the default, and High and Max add more thinking. At those three levels, input costs $2 per million tokens, at up to an 80% cache hit rate, and output costs $6 per million. Two more levels, z-low and z-high, are not released yet. They are the highest levels of thinking, and they send requests through a different internal route that costs more and produces better results. They will be priced at $6 per million input tokens, at up to an 80% cache hit rate, and $15 per million output tokens.
In practice: start at Low. Move up to High or Max when a task is hard enough that being right matters more than being quick, such as a security review, a large refactor, or a bug that touches several files. Turn thinking off when you need the quickest possible reply and the question is simple.
None of this makes Kael the right tool for every job. If you need the fastest response available, or your work is mostly creative writing, there are better fits. If you need the answer to be correct, that is what Kael is for.