Security, Compliance & Certifications
GDPR Compliant
EU AI Act Compliant
AI 100% Sovereign
No Training of Models
World Class Models
Product Presentation
Watch our keynote presentation to discover how IG1 AI API delivers World-class AI performance with full data sovereignty — at up to 3x lower cost.
Official GLM 5.3 Flash benchmarks highlight leading agentic coding, long-horizon terminal execution, a one-million-token context, and native multimodal input.
Terminal-Bench 2.1
GLM 5.3 Flash scores 84.3 on Terminal-Bench 2.1 in Z.ai’s official benchmark table, level with the strongest proprietary models on agentic terminal execution.
Context Window
GLM 5.3 Flash supports an official one-million-token context window and up to 128K output tokens, making it suitable for long documents, large codebases, and multi-step agent workflows.
Native Multimodal
GLM 5.3 Flash is the first natively multimodal model in the GLM-5 series: video, image, text, and file input in a sparse MoE architecture with 18 billion active parameters.
A filmed testimonial: EasyBourse shares how sovereign AI helps its teams scale software development.
Watch on YouTubeLeveraging the most powerful open-weight models available today. No vendor lock-in. Full transparency.
GLM 5.3 Flash Frontier
by Z.ai — Beijing, China
Efficient Frontier LLM
The flagship of our catalogue for reasoning, code and agentic workflows — and its lowest token price.
Qwen 3.5
by Alibaba Cloud — Hangzhou, China
122B-A10B LLM
Efficient high-performance model for production chat, code, reasoning, and multi-turn workflows.
Qwen Image
by Alibaba Cloud — Hangzhou, China
Vision & Multimodal
Image understanding, visual analysis, and multimodal reasoning for enterprise workflows.
Qwen3-VL-Embedding-8B SOTA
by Alibaba Cloud — Hangzhou, China
Multimodal Embedding
State-of-the-art retrieval with multimodal embeddings and dynamic dimensions up to 4096 — far beyond fixed-size text encoders.
BGE
by BAAI — Beijing, China
Fast Embedding & Reranking
Fast, fixed-dimension text embeddings and reranking — a quality reference for high-throughput semantic search and RAG pipelines.
Two models are joining our sovereign infrastructure, covering the compact tier and the mid-size tier.
Qwen 3.8 Flash Next September 2026
by Alibaba Cloud — Hangzhou, China
Mid-Size LLM
The next mid-size tier, set to replace Qwen 3.5 122B-A10B for production chat, code, and multi-turn workflows.
Qwen 3.8 27B September 2026
by Alibaba Cloud — Hangzhou, China
27B Dense LLM
Latest-generation compact dense model for production chat, code, extraction, and multi-turn workflows.
Transparent per-token pricing. All models are sovereign, hosted in France, with no hidden fees and no minimum commitment.
| Model | Description | 1M Input Tokens | 1M Output Tokens |
|---|---|---|---|
| Natively multimodal efficient frontier model (video, image, text, file) for reasoning, code and agentic workflows. Max context: 1M tokens, 128K output. | 0,40 €Cached input: 0,09 € | 1,30 € | |
| Efficient high-performance model for production chat, code, and reasoning. Default instruct profile. | 0,60 € | 2,50 € | |
| Same as Qwen3.5-122B-A10B with hyperparameters tuned for creative ideation, synthesis, and high-quality content generation. | 0,60 € | 2,50 € | |
| Same as Qwen3.5-122B-A10B with thinking mode enforced for structured reasoning and analysis. Max context: 256K tokens. | 0,60 € | 2,50 € | |
| Qwen3.5-122B-A10B configured for advanced coding, architecture, refactoring, debugging, and thinking-enforced development workflows. | 0,60 € | 2,50 € | |
| Next mid-size tier, set to replace Qwen3.5-122B-A10B for production chat, code and multi-turn workflows. | 0,60 € | 2,50 € | |
| Newer-generation compact model for production chat, code, extraction, and multi-turn workflows. Default instruct profile. | 0,60 € | 2,50 € | |
| Same as Qwen3.8-27B with hyperparameters tuned for creative ideation, synthesis, and high-quality content generation. | 0,60 € | 2,50 € | |
| Same as Qwen3.8-27B with thinking mode enforced for structured reasoning and analysis. | 0,60 € | 2,50 € | |
| Qwen3.8-27B configured for advanced coding, refactoring, debugging, and thinking-enforced development workflows. | 0,60 € | 2,50 € | |
| BGE-m3 | Text embeddings for RAG and semantic search. | 0,05 € | N/A |
| BGE-reranker-v2-m3 | Reranking for search result optimization. | 0,30 € | N/A |
| Document | Convert any document in Markdown to be used by LLM. Useful for CAG – Cached RAG including vision RAG use cases. Document images uses IG1 Standard to be described as text. | N/A | 2,00 € |
| Model | Description | 1M Input Tokens | 1M Output Tokens |
|---|---|---|---|
| Image generation. Token count based on image resolution and quality level. | 0,60 € | 2,50 € | |
| Image generation with Prompt Enhancing by LLM. Prompt enhancement uses Qwen3.5-122B-A10B (charged separately). Token count based on resolution and quality level. | 0,60 € | 2,50 € | |
| Image editing. Token count based on source and generated image resolution and quality level. | 0,60 € | 2,50 € | |
| Image editing with Prompt Enhancing by LLM. Prompt enhancement uses Qwen3.5-122B-A10B (charged separately). Token count based on source and generated image resolution and quality. | 0,60 € | 2,50 € |
Get your access keys. Integrate in minutes. Test fast. Scale with confidence.
1. Get Keys
Request your API access keys from our team.
2. Integrate
Connect in minutes with our OpenAI-compatible API.
3. Scale
Test fast, iterate, and scale with confidence.
Everything you need to know about IG1 AI API — sovereignty, models, pricing, and getting started.
Sovereignty & Security
Models & Performance
Pricing & Billing
Getting Started