Reference

Models inside Kotia services

11 models from 4 companies. You pick a model for every message and can switch mid-conversation. If a model doesn't suit you, you can always choose another and find the one that performs best for your tasks.

Model comparison table

ModelContext windowCostFilesWeb searchAccess
GPT-5.6 Luna default
OpenAI · Fast, cheapest
≈ 2042 pages
1050000 tokens
low
the cheapest one
Images, PDF, text
Built-in
Everyone
Gemini 3.5 Flash-Lite
Google · Editing, fast
≈ 1944 pages
1000000 tokens
low
×1.5 vs the cheapest
Images, PDF, text
Built-in
Everyone
DeepSeek V4 Flash
DeepSeek · Cheap, text only
≈ 1611 pages
1000000 tokens
low
×2.2 vs the cheapest
Text only
Via Kotia
Everyone
GPT-5.4 mini
OpenAI · Fast, research
≈ 778 pages
400000 tokens
medium
×3.8 vs the cheapest
Images, PDF, text
Built-in
Subscription
Claude Haiku 4.5
Anthropic · Agent tasks
≈ 222 pages
200000 tokens
medium
×5 vs the cheapest
Images, PDF, text
Built-in
Everyone
DeepSeek V4 Pro
DeepSeek · Strong reasoning, text only
≈ 1611 pages
1000000 tokens
medium
×6.6 vs the cheapest
Text only
Via Kotia
Subscription
Gemini 3.5 Flash
Google · Best for writing
≈ 1944 pages
1000000 tokens
medium
×7.5 vs the cheapest
Images, PDF, text
Built-in
Subscription
Claude Sonnet 5
Anthropic · Balanced
≈ 1111 pages
1000000 tokens
medium
×10 vs the cheapest
Images, PDF, text
Built-in
Subscription
GPT-5.4
OpenAI · Deep reasoning
≈ 2042 pages
1050000 tokens
high
×12.5 vs the cheapest
Images, PDF, text
Built-in
Subscription
GPT-5.6 Sol
OpenAI · Strongest for documents
≈ 2042 pages
1050000 tokens
high
×20 vs the cheapest
Images, PDF, text
Built-in
Subscription
Claude Opus 5
Anthropic · Most capable
≈ 1111 pages
1000000 tokens
high
×25 vs the cheapest
Images, PDF, text
Built-in
Subscription

Sorted from the cheapest to the most expensive. A page is 1,800 characters.

What the terms above mean

Context window
How much text the model keeps in mind while replying: your question, the whole conversation, attached files and its own answer. Once the window fills up, older messages stop being taken into account. We show it in pages of plain text so it's easier to tell whether a document will fit.
Page
A rough unit of volume: 1,800 characters including spaces, about one A4 page in a regular font. The page count in the table is an estimate — the exact figure depends on the language of the text.
Token
A chunk of text that models count with: a word or part of a word. Providers measure the window and charge in tokens. As a rule of thumb, one word is one to three tokens.
Cost
How much is deducted from your balance for one reply. It depends on the model and on length: the longer the conversation and the files, the more each following reply costs. The table shows the tier and how many times more expensive the model is than the cheapest one.
Files
What you can attach to a message. «Images, PDF, text» — the model sees the image and the PDF pages. «Text only» — the model accepts TXT, MD, CSV and JSON and rejects an image or PDF right when you send it.
Web search
«Built-in» — the model searches the web itself, inside its reply. «Via Kotia» — the platform runs the search and opens the pages, then hands the findings to the model. The results are comparable; what differs is who searched and how it shows in the feed.
Access
«Everyone» — the model is available on any account. «Subscription» — only with an active subscription. There are no free models: tokens are deducted for every reply from any model; the subscription only decides which models you can choose.
Default
The model the chat runs on until you pick another one. You can switch at any point in the conversation — the next reply comes from the new model, previous ones stay as they are.

What makes each model different and which one to choose

Everything below comes from working with the models in Kotia, not from provider marketing.

ModelWhat makes it differentWhen to choose itWhen to avoid it
GPT-5.6 Luna default
OpenAI · Fast, cheapest
  • The cheapest model in the chat.
  • Runs by default until you pick another one.
  • Medium: thinks longer on complex tasks.
  • Context window usage is shown with a margin: the provider has no exact counter.
  • Conversations are not retained by the provider.
Long professional documents, analysis, structured texts. The junior model is the cheapest in the chat.Precise tracking of conversation length — the provider offers no counter, so the remaining window is shown with a safety margin.
Gemini 3.5 Flash-Lite
Google · Editing, fast
  • Cost low: ×1.5 vs the cheapest.
  • The fastest group in the chat.
  • Search happens inside the reply — you won't see a separate «searched the web» line.
  • Conversations are not retained by the provider.
Text: writing, editing, conversation, working with large volumes of material. The fastest replies in the chat.Long agent chains with tools — on multi-step tasks it falls behind Anthropic.
DeepSeek V4 Flash
DeepSeek · Cheap, text only
  • Cost low: ×2.2 vs the cheapest.
  • Slow: a reply can take up to a minute.
  • On the junior model, internal reasoning occasionally leaks into the visible reply — it looks like someone else's monologue before the answer.
  • Conversations are retained by the provider — this cannot be turned off.
Reasoning and problem analysis on a budget, when the work is text only.Images and PDF (not accepted), confidential materials, urgent short questions.
GPT-5.4 mini
OpenAI · Fast, research
  • Cost medium: ×3.8 vs the cheapest.
  • Available with a subscription.
  • Medium: thinks longer on complex tasks.
  • Context window usage is shown with a margin: the provider has no exact counter.
  • Conversations are not retained by the provider.
Long professional documents, analysis, structured texts. The junior model is the cheapest in the chat.Precise tracking of conversation length — the provider offers no counter, so the remaining window is shown with a safety margin.
Claude Haiku 4.5
Anthropic · Agent tasks
  • Cost medium: ×5 vs the cheapest.
  • The shortest context window on the list — about 222 pages.
  • Replies evenly, without long pauses.
  • Conversations are not retained by the provider.
Agent-style work: canvas, tool chains, pipelines, digging through materials with search and source reading. Holds a long sequence of steps without losing the task.Pure text work — style and editing are drier than Google's. There are no cheap models in this family.
DeepSeek V4 Pro
DeepSeek · Strong reasoning, text only
  • Cost medium: ×6.6 vs the cheapest.
  • Available with a subscription.
  • Slow: a reply can take up to a minute.
  • On the junior model, internal reasoning occasionally leaks into the visible reply — it looks like someone else's monologue before the answer.
  • Conversations are retained by the provider — this cannot be turned off.
Reasoning and problem analysis on a budget, when the work is text only.Images and PDF (not accepted), confidential materials, urgent short questions.
Gemini 3.5 Flash
Google · Best for writing
  • Cost medium: ×7.5 vs the cheapest.
  • Available with a subscription.
  • The fastest group in the chat.
  • Search happens inside the reply — you won't see a separate «searched the web» line.
  • Conversations are not retained by the provider.
Text: writing, editing, conversation, working with large volumes of material. The fastest replies in the chat.Long agent chains with tools — on multi-step tasks it falls behind Anthropic.
Claude Sonnet 5
Anthropic · Balanced
  • Cost medium: ×10 vs the cheapest.
  • Available with a subscription.
  • Replies evenly, without long pauses.
  • Conversations are not retained by the provider.
Agent-style work: canvas, tool chains, pipelines, digging through materials with search and source reading. Holds a long sequence of steps without losing the task.Pure text work — style and editing are drier than Google's. There are no cheap models in this family.
GPT-5.4
OpenAI · Deep reasoning
  • Cost high: ×12.5 vs the cheapest.
  • Available with a subscription.
  • Medium: thinks longer on complex tasks.
  • Context window usage is shown with a margin: the provider has no exact counter.
  • Conversations are not retained by the provider.
Long professional documents, analysis, structured texts. The junior model is the cheapest in the chat.Precise tracking of conversation length — the provider offers no counter, so the remaining window is shown with a safety margin.
GPT-5.6 Sol
OpenAI · Strongest for documents
  • Cost high: ×20 vs the cheapest.
  • Available with a subscription.
  • Medium: thinks longer on complex tasks.
  • Context window usage is shown with a margin: the provider has no exact counter.
  • Conversations are not retained by the provider.
Long professional documents, analysis, structured texts. The junior model is the cheapest in the chat.Precise tracking of conversation length — the provider offers no counter, so the remaining window is shown with a safety margin.
Claude Opus 5
Anthropic · Most capable
  • Cost high: ×25 vs the cheapest.
  • Available with a subscription.
  • Replies evenly, without long pauses.
  • Conversations are not retained by the provider.
Agent-style work: canvas, tool chains, pipelines, digging through materials with search and source reading. Holds a long sequence of steps without losing the task.Pure text work — style and editing are drier than Google's. There are no cheap models in this family.

Questions and answers

In the chat and on the canvas — where you talk to the agent and build chains of steps. Article generation, research, news and AI detection run on their own models, and the choice in the chat does not affect them.
No. Tokens are deducted for every reply from any model — free daily ones or purchased. A subscription doesn't make models free; it unlocks the strong models that are not on the list without it.
The model selector sits next to the input field and can be changed at any time. The next reply comes from the new model, and the whole conversation stays in place — the new model sees all of it.
Kotia makes up to three attempts on the same model. If that fails, you'll see an error and decide what to do: retry or switch to another model. Kotia never swaps the model silently — you always know who replied.
The price of a reply depends on length: the model reads not only the question but the whole conversation, attached files and search results. The longer the conversation, the more each following reply costs. A new chat for a new topic is the simplest way to pay less.
The reply won't be cut off or lost: you'll receive it, and the balance may go slightly negative. The next daily top-up covers it. To start a conversation you need a minimum balance.
Yes, with models marked «Images, PDF, text». A model marked «Text only» accepts TXT, MD, CSV and JSON and rejects an image or PDF right when you send it, rather than replying about an empty file.
The model can no longer fit the whole conversation, and older messages stop being taken into account. The usage meter next to the input field shows the remaining space with a margin. For a large document, pick a model with a wider window or start a new chat.
Conversations are stored in Kotia. With most models the provider doesn't retain them. The exception is DeepSeek: there the conversation stays with the provider and this cannot be disabled, so it's better not to send confidential materials there.
That's the model's internal reasoning, which occasionally leaks into the visible reply — it happens with the junior DeepSeek model. The answer itself is correct. If it bothers you, send the message again or choose another model.