# searchit — model price list

> Your token price for the language models searchit offers for AI extraction and the AI Agent, for embedding and for image analysis — outside and inside the EU.

## See every model price.

This list shows your token price for the language models we offer with searchit — for AI extraction and the AI Agent —, for the embedding behind semantic search and for image analysis.

As of 25 September 2026

## Language models

A selection: language models that are particularly well suited to AI extraction — that is, to deriving summaries, entities and relations from your documents. For each model the token price with processing outside and inside the EU, ordered by the result of our own extraction measurement. A dash means: we do not offer the model at that location.

**Token prices per 1 million tokens, in euros excluding VAT**

| Model | Extraction | Self-hostable | In the EU: Input | In the EU: From cache | In the EU: Output | Outside the EU: Input | Outside the EU: From cache | Outside the EU: Output |
|---|---:|---|---:|---:|---:|---:|---:|---:|
| GPT-6 astra (OpenAI, from the USA) | 7.7 | no | — | — | — | €10.51 | €1.05 | €52.54 |
| DeepSeek V4.1 Flash (DeepSeek, from China) | 6.8 | no | — | — | — | €0.3152 | €0.0063 | €1.26 |
| DeepSeek V4 Pro (DeepSeek, from China) | 6.4 | yes | €1.84 | — | €3.68 | €1.00 | — | €2.01 |
| Mistral Medium 3.5 (Mistral AI, from France) | 6.1 | yes | €1.80 | — | €7.20 | €1.50 | €0.1500 | €7.68 |
| Kimi K2.6 (Moonshot AI, from China) | 6.0 | yes | €0.7920 | €0.2640 | €4.50 | €0.7200 | — | €3.60 |
| Qwen3.5-27B (Alibaba Cloud, from China) | 5.5 | yes | — | — | — | €0.2049 | — | €1.64 |
| GLM-5 (Z.ai, from China) | 4.9 | yes | €1.14 | — | €3.48 | €0.6305 | — | €2.02 |
| Mistral Large 2.1 (Mistral AI, from France) | 4.6 | yes | — | — | — | €2.10 | — | €6.30 |
| Qwen3.5-397B-A17B (Alibaba Cloud, from China) | 4.5 | yes | €0.6050 | — | €3.63 | €0.5779 | — | €3.68 |
| Gemma 4 31B (Google, from the USA) | 4.3 | yes | €0.2400 | — | €0.4200 | €0.0946 | — | €0.3573 |
| Qwen3.7-Plus (Alibaba Cloud, from China) | 3.3 | no | — | — | — | €0.4203 | €0.0420 | €1.68 |
| Qwen3.6-35B-A3B (Alibaba Cloud, from China) | 3.3 | yes | €0.1800 | — | €0.6000 | €0.1576 | — | €1.05 |
| Qwen3.5-9B (Alibaba Cloud, from China) | 2.9 | yes | €0.1008 | — | €0.1513 | €0.1051 | — | €0.1576 |

- Extraction: quality score from 0 to 10 from our own measurement — the mean hit rate for recognised entities and relations against an independently prepared reference, divided by ten. It serves to compare the models with each other. Measured in 2026 on eight sample documents — it holds for these documents, not for every collection.
- Self-hostable: the model weights are openly available; the model can also run on your own hardware.
- In the EU: we offer a model there only through providers that process your data in the EU and neither store your inputs and outputs nor use them for training — documented by their terms. Where this is not documented, there is a dash.

**Performance profile per model**

| Model | Entities found | Relations found | Context window | Model size |
|---|---:|---:|---|---|
| GPT-6 astra | 8.1 | 7.4 | not published | not published |
| DeepSeek V4.1 Flash | 7.7 | 5.9 | 1,000,000 tokens | 763.2bn, 16bn of them active |
| DeepSeek V4 Pro | 7.5 | 5.2 | 1,048,576 tokens | 1.6tn, 49bn of them active |
| Mistral Medium 3.5 | 7.0 | 5.3 | 262,144 tokens | 127.7bn |
| Kimi K2.6 | 7.5 | 4.5 | 262,144 tokens | 1tn, 32bn of them active |
| Qwen3.5-27B | 6.7 | 4.3 | 262,144 tokens | 27.8bn |
| GLM-5 | 6.0 | 3.8 | 204,800 tokens | 753.9bn, 40bn of them active |
| Mistral Large 2.1 | 6.1 | 3.1 | 131,072 tokens | 123bn |
| Qwen3.5-397B-A17B | 6.0 | 3.0 | 262,144 tokens | 403.4bn, 17bn of them active |
| Gemma 4 31B | 5.5 | 3.1 | 262,144 tokens | 31bn |
| Qwen3.7-Plus | 4.3 | 2.3 | 1,000,000 tokens | 397bn, 17bn of them active |
| Qwen3.6-35B-A3B | 4.7 | 1.8 | 262,144 tokens | 36bn, 3bn of them active |
| Qwen3.5-9B | 4.7 | 1.1 | 262,144 tokens | 9.7bn |

- Quality score 0 to 10 per subtask, from our own extraction measurement in 2026 on eight sample documents.
- Performance profile: the context window is how many tokens the model processes at once; model size is its number of parameters — for models that compute only part of them per token, the active part follows.

## Embedding

For semantic search every text section is translated once into a vector; only input is charged. Which model your index uses is fixed at set-up — changing it means rebuilding the index. The two standard models come first, all others below by vendor.

**Token prices of embedding per 1 million input tokens, in euros excluding VAT**

| Model | Search German | Self-hostable | In the EU | Outside the EU |
|---|---:|---|---:|---:|
| Qwen3 Embedding 8B (Alibaba Cloud, from China) — standard outside the EU | 0.80 | yes | €0.0105 | €0.0105 |
| BGE-M3 (BAAI, from China) — standard in the EU | 0.69 | yes | €0.0120 | €0.0105 |
| Qwen Text Embedding v4 (Alibaba Cloud, from China) | — | no | — | €0.0736 |
| Qwen3-VL Embedding 8B (Alibaba Cloud, from China) | — | yes | €0.0960 | — |
| Amazon Titan Text Embeddings V2 (Amazon, from the USA) | — | no | €0.2102 | €0.0210 |
| BGE Multilingual Gemma2 (BAAI, from China) | — | yes | €0.0120 | — |
| Cohere Embed 4 (Cohere, from Canada) | 0.81 | no | — | €0.1261 |
| Gemini Embedding 2 (Google, from the USA) | — | no | — | €0.2102 |
| E5-Mistral-7B (Microsoft, from the USA) | 0.72 | yes | €0.0240 | — |
| Codestral Embed (Mistral AI, from France) | — | no | — | €0.1800 |
| Mistral Embed (Mistral AI, from France) | — | no | — | €0.1200 |
| text-embedding-3-large (OpenAI, from the USA) | 0.75 | no | — | €0.1366 |
| text-embedding-3-small (OpenAI, from the USA) | 0.71 | no | — | €0.0210 |
| Voyage 4 (Voyage AI, from the USA) | 0.83 | no | — | €0.0630 |
| Voyage 4 large (Voyage AI, from the USA) | 0.84 | no | — | €0.1261 |
| Voyage 4 lite (Voyage AI, from the USA) | 0.80 | no | — | €0.0210 |

- Search German: mean of the four German retrieval tasks of the public benchmark MTEB (nDCG@10, from 0 to 1) — the higher, the better the model finds the text sections that match a question. It serves to compare the models with each other, not as a promise for your data. A dash: no value has been published for the model.
- Standard outside the EU / standard in the EU: we use this model unless you choose otherwise — without or with fully European operation.
- In the EU: we offer a model there only through providers that process your data in the EU and neither store your inputs and outputs nor use them for training — documented by their terms. Where this is not documented, there is a dash.

**Performance profile per model**

| Model | General texts | Healthcare | Law | Legal questions | Vector length | Context window | Model size |
|---|---:|---:|---:|---:|---:|---|---|
| Qwen3 Embedding 8B | 0.97 | 0.88 | 0.70 | 0.67 | 4,096 | 32,768 tokens | 7.6bn |
| BGE-M3 | 0.92 | 0.72 | 0.61 | 0.50 | 1,024 | 8,192 tokens | 568m |
| Qwen Text Embedding v4 | — | — | — | — | 1,024 | 8,192 tokens | not published |
| Qwen3-VL Embedding 8B | — | — | — | — | 4,096 | 32,768 tokens | not published |
| Amazon Titan Text Embeddings V2 | — | — | — | — | 1,024 | 8,192 tokens | not published |
| BGE Multilingual Gemma2 | — | — | — | — | 3,584 | 8,192 tokens | 9.2bn |
| Cohere Embed 4 | 0.95 | 0.85 | 0.73 | 0.72 | 1,536 | 128,000 tokens | not published |
| Gemini Embedding 2 | — | — | — | — | 3,072 | 8,192 tokens | not published |
| E5-Mistral-7B | 0.97 | 0.79 | 0.69 | 0.44 | 4,096 | 4,096 tokens | 7.1bn |
| Codestral Embed | — | — | — | — | 1,536 | 8,192 tokens | not published |
| Mistral Embed | — | — | — | — | 1,024 | 8,192 tokens | not published |
| text-embedding-3-large | 0.98 | 0.81 | 0.67 | 0.56 | 3,072 | 8,192 tokens | not published |
| text-embedding-3-small | 0.96 | 0.70 | 0.62 | 0.55 | 1,536 | 8,192 tokens | not published |
| Voyage 4 | 0.98 | 0.88 | 0.75 | 0.71 | 1,024 | 32,000 tokens | not published |
| Voyage 4 large | 0.98 | 0.91 | 0.76 | 0.72 | 1,024 | 32,000 tokens | not published |
| Voyage 4 lite | 0.97 | 0.83 | 0.73 | 0.68 | 1,024 | 32,000 tokens | not published |

- Search quality per German task of the public MTEB benchmark (nDCG@10, 0 to 1); their mean is the “Search German” value.
- Performance profile: the context window is how many tokens the model processes at once; model size is its number of parameters — for models that compute only part of them per token, the active part follows.

## Image models

For image analysis an image model reads every image and describes its content. For each model the token price with processing outside and inside the EU, ordered by the public benchmark MMMU-Pro. A dash means: we do not offer the model at that location.

**Token prices of the image models per 1 million tokens, in euros excluding VAT**

| Model | MMMU-Pro | Self-hostable | In the EU: Input | In the EU: Output | Outside the EU: Input | Outside the EU: Output |
|---|---:|---|---:|---:|---:|---:|
| Kimi K2.6 (Moonshot AI, from China) | 79.4% | yes | €0.8400 | €4.20 | €0.7200 | €3.60 |
| Qwen3.5-397B-A17B (Alibaba Cloud, from China) | 79.0% | yes | €0.6050 | €3.63 | €0.4098 | €2.46 |
| MiniMax M3 (MiniMax, from China) | 78.1% | yes | €0.4203 | €2.10 | €0.3152 | €1.26 |
| Gemma 4 31B (Google, from the USA) — standard outside the EU, standard in the EU | 76.9% | yes | €0.1200 | €0.3600 | €0.0946 | €0.3573 |
| Qwen3.6-27B (Alibaba Cloud, from China) | 75.8% | yes | €0.5400 | €0.7800 | €0.3152 | €2.10 |
| Qwen3.6-35B-A3B (Alibaba Cloud, from China) | 75.3% | yes | €0.3600 | €1.44 | €0.1051 | €0.9457 |
| Gemma 4 26B A4B (Google, from the USA) | 73.8% | yes | €0.1200 | €0.6000 | €0.0736 | €0.3573 |
| Qwen3.5-9B (Alibaba Cloud, from China) | 70.1% | yes | €0.1008 | €0.1513 | €0.1051 | €0.1576 |

- MMMU-Pro: public benchmark with questions about images, charts and document pages; shown is the share of correctly solved tasks, as published by the respective vendor. It serves to compare the models with each other, not as a promise for your images.
- Standard outside the EU / standard in the EU: we use this model unless you choose otherwise — without or with fully European operation.
- In the EU: we offer a model there only through providers that process your data in the EU and neither store your inputs and outputs nor use them for training — documented by their terms. Where this is not documented, there is a dash.

**Performance profile per model**

| Model | MMMU-Pro | Context window | Model size |
|---|---:|---|---|
| Kimi K2.6 | 79.4% | 262,144 tokens | 1tn, 32bn of them active |
| Qwen3.5-397B-A17B | 79.0% | 262,144 tokens | 403.4bn, 17bn of them active |
| MiniMax M3 | 78.1% | 1,048,576 tokens | 427bn, 23bn of them active |
| Gemma 4 31B | 76.9% | 262,144 tokens | 31.3bn |
| Qwen3.6-27B | 75.8% | 262,144 tokens | 27.8bn |
| Qwen3.6-35B-A3B | 75.3% | 262,144 tokens | 36bn, 3bn of them active |
| Gemma 4 26B A4B | 73.8% | 262,144 tokens | 25.8bn, 4bn of them active |
| Qwen3.5-9B | 70.1% | 262,144 tokens | 9.7bn |

- Share of correctly solved tasks in the public MMMU-Pro benchmark, as published by the vendor.
- Performance profile: the context window is how many tokens the model processes at once; model size is its number of parameters — for models that compute only part of them per token, the active part follows.

All prices in euros plus VAT.
