IG1 AI API
SOVEREIGN AI API — HOSTED IN FRANCE, NO DATA RETENTION

Sovereign AI API.
Frontier Models.
Zero Data Leakage.

Frontier reasoning, code generation,
long-context analysis and agentic workflows through a sovereign API hosted in France.
No data retention. No compromise.

Get Your API Keys Contact Sales
SCROLL

Security, Compliance & Certifications

GDPR

GDPR Compliant

EU AI Act

EU AI Act Compliant

AI Sovereign

AI 100% Sovereign

No Training

No Training of Models

World Class Models

World Class Models


Product Presentation

See IG1 AI API
in Action.

Watch our keynote presentation to discover how IG1 AI API delivers World-class AI performance with full data sovereignty — at up to 3x lower cost.


Z.ai Performance

GLM 5.3 Flash
Performance Signals.

Official GLM 5.3 Flash benchmarks highlight leading agentic coding, long-horizon terminal execution, a one-million-token context, and native multimodal input.

Terminal-Bench 2.1

84.3 agentic benchmark

GLM 5.3 Flash scores 84.3 on Terminal-Bench 2.1 in Z.ai’s official benchmark table, level with the strongest proprietary models on agentic terminal execution.

  • —Agentic coding and terminal execution
  • —320B MoE architecture, 18B active parameters
  • —No data retention, fully sovereign
  • —Frontier reasoning through IG1 AI API

Context Window

1M tokens

GLM 5.3 Flash supports an official one-million-token context window and up to 128K output tokens, making it suitable for long documents, large codebases, and multi-step agent workflows.

  • —1M-token official context length
  • —Up to 128K output tokens
  • —Hybrid sparse and linear attention

Native Multimodal

320B parameters — 18B active

GLM 5.3 Flash is the first natively multimodal model in the GLM-5 series: video, image, text, and file input in a sparse MoE architecture with 18 billion active parameters.

  • —Video, image, text, and file input
  • —3.01× less attention compute vs GLM-5.3
  • —4.44× smaller KV cache vs GLM-5.3
  • —Open weights under MIT licence

Z.ai Efficient Frontier

GLM 5.3 Flash.
The Flagship.

GLM 5.3 Flash is the efficient frontier model of IG1 AI API: a natively multimodal reasoning, coding, and agentic execution engine with a one-million-token context, delivered with sovereign inference and transparent per-token pricing.

Model Positioning Best for Offered by IG1 AI
Small model Compact Simple automation ✓ Qwen 3.8 27B
Mid-size model Balanced Fast enterprise workloads ✓ Qwen 3.8 Flash Next
Efficient frontier Efficient World-class reasoning, code, and agents ✓ GLM 5.3 Flash

What GLM 5.3 Flash changes in practice

GLM 5.3 Flash is built for tasks where quality, precision, and persistence matter more than a quick generic answer. Here's where the difference is immediately visible:

Long contract analysis

GLM 5.3 Flash follows cross-references, nested clauses, exceptions, and definitions across very long documents — up to a one-million-token context — and accepts scanned pages and images as readily as text.

Multi-file code generation

From architecture planning to implementation details, GLM 5.3 Flash is tuned for coherent multi-file development, refactoring, debugging, and terminal execution, scoring 84.3 on Terminal-Bench 2.1.

Multi-constraint reasoning

Budget, timeline, stack, compliance, security, UX, scalability: GLM 5.3 Flash keeps simultaneous constraints active and turns them into clear decisions.

Complex instructions over long context

For long prompts with nested requirements, GLM 5.3 Flash is the model to use when your team needs the answer to stay faithful to every instruction through the end of the generation.

Why choose GLM 5.3 Flash?

Use GLM 5.3 Flash when the work is strategic: hard reasoning, product ideation, advanced coding, agentic planning, deep analysis, and documents and media where small mistakes are expensive. It is also the most affordable model in the IG1 AI catalog — the default choice for everyday work too.

Frontier performance at the lowest token price

GLM 5.3 Flash is the flagship model of the catalogue — and also its most affordable: 0,40 € / 1M input tokens (0,09 € cached) and 1,30 € / 1M output tokens. Qwen 3.8 Flash Next, Qwen 3.8 27B and Qwen 3.5 122B-A10B complete the catalogue for production workloads, all three at 0,60 € / 2,50 €.


Qwen Mid-Size Model

Qwen 3.8 Flash Next.
Mid-Size Power, Production Ready.

GLM 5.3 Flash is the flagship model. Qwen 3.8 Flash Next is the mid-size tier: a new-architecture model for production traffic — with Qwen 3.8 27B in the compact tier and Qwen 3.5 122B-A10B for a larger knowledge base, all three at exactly the same rate.

New-generation mid-size model

Qwen 3.8 Flash Next is the newest generation of the Qwen family: a sparse mid-size model that engages under 5% of its parameters per token — mid-size capability with a compact inference footprint.

Built for production traffic

At 0,60 € input and 2,50 € output per 1M tokens, Qwen 3.8 Flash Next is built for production traffic, internal copilots, and high-volume agent pipelines.

Production-grade versatility

Use it for chat, coding assistance, OCR/vision workflows, structured extraction, summarization, and tool-using agents when frontier-level reasoning is not required.

Three models, one price

Qwen 3.8 27B and Qwen 3.5 122B-A10B stay available at exactly the same 0,60 € / 2,50 € rate: drop down to the compact dense model, or switch to the larger Mixture-of-Experts knowledge base, whenever a workload needs it — without touching your budget.

New Qwen

Qwen 3.8 Flash Next: a new architecture

September 2026

Qwen 3.8 Flash Next is the public preview of the Qwen4 architecture. Where almost every LLM still relies on dense attention and a cache that grows with the context, it rebuilds four layers at once: attention, residual, embeddings and optimiser.

3:1 hybrid attention

Three Gated DeltaNet layers to one Qwen Sparse Attention layer, repeated twelve times across 48 layers. Gated DeltaNet compresses the history instead of keeping all of it; Qwen Sparse Attention uses a lightweight indexer to pick only the context that matters, at micro-block granularity. That is what produces the very-long-context gains.

Sparsity pushed to the limit

512 experts, 10 routed plus 1 shared per token: 6 billion active parameters out of 125, under 5% of the model engaged. The sparsest model in the IG1 catalogue, ahead of GLM 5.3 Flash (5.6%) and Qwen 3.5 122B-A10B (8.2%).

Offloadable N-gram embeddings

A further 51 billion parameters held in a table that sits outside the compute path and can be offloaded from GPU memory: model capacity grows without a proportional inference cost.

A different training recipe

A four-branch gated residual to hold stability at that sparsity, and the Muon optimiser in place of AdamW. Alibaba reports training at one ninth the cost of Qwen3.7-Plus.

Parameters

125B + 51B

6B active per token

Experts

512

10 routed + 1 shared per token

Context

262K → 1M

native, extended with YaRN

Throughput at 1M tokens

7.6× / 4.9×

prefill / decode

Figures reported by Alibaba Cloud.

Use case Recommended model Why
Hard reasoning, strategic coding, complex agents GLM 5.3 Flash Maximum capability, lowest token price
Fast enterprise workloads, very long context Qwen 3.8 Flash Next Mid-size tier, up to 1M tokens, same price
Production chat, extraction, copilots, high-volume agents Qwen 3.8 27B Compact dense model for high-volume workloads
Broad knowledge, long documents, multimodal workflows Qwen 3.5 122B-A10B Larger knowledge base, same price
Hybrid routing GLM + Qwen GLM by default, Qwen where its profile fits

Use Cases

What Does This Mean
for Your Team?

Everyone in your team can adopt AI — not just your engineers. Marketing, sales, analysts, developers. IG1 AI API meets you where you are.

Chat

Integrate conversational AI into any application. Build smart assistants, customer support bots, and interactive workflows.

Code

Ship features faster with AI-assisted development. Generate, review, refactor and debug code at enterprise scale.

Summarize

Transform thousands of documents into actionable insights. Extract, condense, and analyze enterprise knowledge at scale.

Create

Generate content and analysis at scale. From marketing copy to financial reports, empower every department with AI-powered creation.


Customer testimonial — EasyBourse
Customer testimonial

EasyBourse:
scaling software development with sovereign AI.

FINANCE · SOFTWARE DEVELOPMENT

A filmed testimonial: EasyBourse shares how sovereign AI helps its teams scale software development.

Watch on YouTube

Data Sovereignty

What About
Your Data?

The question every enterprise asks. IG1 AI API is built precisely for this. Your security team says yes. Your legal team says yes. And you can move fast with confidence.

Zero Data Retention

Your prompts are processed, then immediately erased. Nothing is stored, nothing is logged.

No Model Training

We never use your data to train models. Your intellectual property belongs to you.

100% Sovereign

Your data never leaves your enterprise boundaries. Fully compliant with GDPR, EU AI Act, and ISO 27001.

Data processing: Your Prompt → IG1 AI API Processes → Immediately Erased

Your Prompt

IG1 AI Processes

Immediately Erased

No logs • No storage • No training


Advanced Capabilities

Not Just Compliance.
Real Capability.

IG1 AI API is designed to go further. Enterprise-grade AI built for the way you work.

Extended Thinking

Complex, multi-step reasoning for advanced problem solving. Chain-of-thought processing for nuanced enterprise scenarios — powered by the GLM 5.3 Flash efficient frontier model, where smaller models lose coherence across long reasoning chains. Deep reasoning, full sovereignty, no data retention.

Web Search & Tool Use

Integrated real-time web search and connected workflows. Let your AI agents interact with external tools and data sources.

Secure Computer Use

Code execution and secure computer use for true agentic automation. Let AI operate in sandboxed environments safely.


Open-Weight Models

Powered by the
Best Open Models.

Leveraging the most powerful open-weight models available today. No vendor lock-in. Full transparency.

Z.ai

GLM 5.3 Flash Frontier

by Z.ai — Beijing, China

Efficient Frontier LLM

The flagship of our catalogue for reasoning, code and agentic workflows — and its lowest token price.

Qwen

Qwen 3.5

by Alibaba Cloud — Hangzhou, China

122B-A10B LLM

Efficient high-performance model for production chat, code, reasoning, and multi-turn workflows.

Qwen

Qwen Image

by Alibaba Cloud — Hangzhou, China

Vision & Multimodal

Image understanding, visual analysis, and multimodal reasoning for enterprise workflows.

Qwen

Qwen3-VL-Embedding-8B SOTA

by Alibaba Cloud — Hangzhou, China

Multimodal Embedding

State-of-the-art retrieval with multimodal embeddings and dynamic dimensions up to 4096 — far beyond fixed-size text encoders.

BGE

by BAAI — Beijing, China

Fast Embedding & Reranking

Fast, fixed-dimension text embeddings and reranking — a quality reference for high-throughput semantic search and RAG pipelines.

Available September 2026

Two models are joining our sovereign infrastructure, covering the compact tier and the mid-size tier.

Qwen

Qwen 3.8 Flash Next September 2026

by Alibaba Cloud — Hangzhou, China

Mid-Size LLM

The next mid-size tier, set to replace Qwen 3.5 122B-A10B for production chat, code, and multi-turn workflows.

Qwen

Qwen 3.8 27B September 2026

by Alibaba Cloud — Hangzhou, China

27B Dense LLM

Latest-generation compact dense model for production chat, code, extraction, and multi-turn workflows.


Model Pricing

Our Models

Transparent per-token pricing. All models are sovereign, hosted in France, with no hidden fees and no minimum commitment.

Language Models

Model Description 1M Input Tokens 1M Output Tokens
GLM-5.3-FlashSeptember 2026 Natively multimodal efficient frontier model (video, image, text, file) for reasoning, code and agentic workflows. Max context: 1M tokens, 128K output. 0,40 €Cached input: 0,09 € 1,30 €
Qwen3.5-122B-A10B Efficient high-performance model for production chat, code, and reasoning. Default instruct profile. 0,60 € 2,50 €
Qwen3.5-122B-A10B Creative Same as Qwen3.5-122B-A10B with hyperparameters tuned for creative ideation, synthesis, and high-quality content generation. 0,60 € 2,50 €
Qwen3.5-122B-A10B Thinking Same as Qwen3.5-122B-A10B with thinking mode enforced for structured reasoning and analysis. Max context: 256K tokens. 0,60 € 2,50 €
Qwen3.5-122B-A10B Thinking Coder Qwen3.5-122B-A10B configured for advanced coding, architecture, refactoring, debugging, and thinking-enforced development workflows. 0,60 € 2,50 €
Qwen3.8-Flash-NextSeptember 2026 Next mid-size tier, set to replace Qwen3.5-122B-A10B for production chat, code and multi-turn workflows. 0,60 € 2,50 €
Qwen3.8-27BSeptember 2026 Newer-generation compact model for production chat, code, extraction, and multi-turn workflows. Default instruct profile. 0,60 € 2,50 €
Qwen3.8-27B CreativeSeptember 2026 Same as Qwen3.8-27B with hyperparameters tuned for creative ideation, synthesis, and high-quality content generation. 0,60 € 2,50 €
Qwen3.8-27B ThinkingSeptember 2026 Same as Qwen3.8-27B with thinking mode enforced for structured reasoning and analysis. 0,60 € 2,50 €
Qwen3.8-27B Thinking CoderSeptember 2026 Qwen3.8-27B configured for advanced coding, refactoring, debugging, and thinking-enforced development workflows. 0,60 € 2,50 €
BGE-m3 Text embeddings for RAG and semantic search. 0,05 € N/A
BGE-reranker-v2-m3 Reranking for search result optimization. 0,30 € N/A
Document Convert any document in Markdown to be used by LLM. Useful for CAG – Cached RAG including vision RAG use cases. Document images uses IG1 Standard to be described as text. N/A 2,00 €

Image Models

Model Description 1M Input Tokens 1M Output Tokens
Qwen Image Image generation. Token count based on image resolution and quality level. 0,60 € 2,50 €
Qwen Image PE Image generation with Prompt Enhancing by LLM. Prompt enhancement uses Qwen3.5-122B-A10B (charged separately). Token count based on resolution and quality level. 0,60 € 2,50 €
Qwen Image Edit Image editing. Token count based on source and generated image resolution and quality level. 0,60 € 2,50 €
Qwen Image Edit PE Image editing with Prompt Enhancing by LLM. Prompt enhancement uses Qwen3.5-122B-A10B (charged separately). Token count based on source and generated image resolution and quality. 0,60 € 2,50 €

Pricing

World-Class Performance.
3× Lower Cost.

Up to 3x cheaper than US hyperscalers. With IG1 AI API, there is no infrastructure to set up. Applications connect instantly and cost savings start from day one.

Talk to Sales
Self-Hosting IG1 AI API
GPU CAPEX $500K – $1M $0
Specialized Staff ~$200K / year $0
Infrastructure Months of setup Instant
R&D Overhead Significant None
Time to Value 6–12 months Day one

Get Started

Access IG1 AI API
Today.

Get your access keys. Integrate in minutes. Test fast. Scale with confidence.

1. Get Keys

Request your API access keys from our team.

2. Integrate

Connect in minutes with our OpenAI-compatible API.

3. Scale

Test fast, iterate, and scale with confidence.


FAQ

Frequently Asked
Questions.

Everything you need to know about IG1 AI API — sovereignty, models, pricing, and getting started.

Sovereignty & Security

Where is my data hosted?
Exclusively in IG1’s own facilities in France, on Dell Technologies and NVIDIA hardware. Your data never leaves our European infrastructure.
Do you retain or use my data to train models?
No. Zero data retention. Zero data training. Zero data sharing. When your request is processed, it’s gone. Period.
Is IG1 AI compliant with GDPR?
Yes. Also ISO 27001 and the EU AI Act. Our sovereign, EU-hosted architecture with strict no-data-retention guarantees is designed for compliance by default.
How is this different from using ChatGPT, Claude or Gemini?
OpenAI, Anthropic, and Google are US companies subject to US jurisdiction (including the CLOUD Act). Your data transits through and may be stored on US infrastructure. With IG1 AI, your data stays in France, under European law, with zero retention, zero data sharing, zero data training.

Models & Performance

What models does IG1 AI use?
We run the best open-weight models available, led by GLM 5.3 Flash for frontier reasoning, coding, and agentic workflows. The catalog also includes Qwen3.5-122B-A10B and Qwen3.8-27B for efficient high-performance language tasks, Qwen3-VL-Embedding-8B and BGE for embeddings and reranking, and Qwen Image for visual generation.
Can I choose which model to use?
Yes. Choose GLM 5.3 Flash — the most capable and most affordable model in the catalog — for frontier reasoning and coding, Qwen3.5-122B-A10B or Qwen3.8-27B for production workloads that suit their profile, or the dedicated Qwen3-VL-Embedding-8B / BGE embedding and reranking models and Qwen Image for specialized tasks.
Are your models quantized?
Yes — deliberately, and chosen to preserve quality. Where a higher-precision release exists (e.g. Qwen), we serve FP8: the quality impact is negligible while GPU memory and cost are roughly halved, which is the right engineering trade-off. When a frontier model is only released quantized, we serve it at its native precision. What we never do is push aggressive low-bit quantization that would degrade the hardest tasks (multi-step reasoning, complex document analysis, large-scale code generation). The benchmark figures we publish reflect the exact precision we serve in production.

Pricing & Billing

How does pricing work?
Two parts. A flat 100 € / month subscription that includes 100 € of usage across every model in the catalog. Beyond that, usage is metered at the published per-token rates (GLM 5.3 Flash at 0,40 € input, 0,09 € cached input and 1,30 € output; Qwen3.5-122B-A10B and Qwen3.8-27B at 0,60 € / 2,50 €), up to a monthly ceiling of 1 000 € on the 100 € subscription (the ceiling scales with your tier). No hidden fees.
What happens if I hit my cap?
Your monthly ceiling is 1 000 € on the 100 € subscription, and scales with higher tiers. When you reach it, access pauses until the next billing cycle — or until you settle that amount. You are never charged beyond the ceiling you control. If you consistently hit it, we’ll discuss raising your limit.
How is overage billed?
On the 1st of each month we bill the previous month’s overage — usage above your included allowance — together with the coming month’s subscription. Payment is by SEPA direct debit or saved card.
Why a subscription plus metered usage?
You get the predictability of a fixed monthly base and the flexibility to scale when a project demands it — no renegotiation. Light months stay light, heavy months stay capped. You only pay for what you use, up to a ceiling you set.
Can I commit to a higher plan?
Yes. Higher monthly subscription tiers — with a larger included allowance and ceiling — are available for teams that need more. Contact sales to arrange a commitment that fits your volume.
Is there a free trial?
Contact us for pilot program options.

Getting Started

How long does setup take?
Zero infrastructure setup on your side. You can be live the same day.
Do I need a technical team to use IG1 AI?
No. IG1 AI is designed for every team — marketing, sales, HR, legal, analysts, and developers. If you can use a chat interface, you can use IG1 AI.
Can I integrate IG1 AI into my existing tools?
Yes. IG1 AI provides API access compatible with standard AI chat applications (Msty, Open WebUI) and IDE integrations for developers. Connect your existing workflows without changing your tools.
What if a better AI platform comes along?
No lock-in. Monthly billing with no commitment means you can leave anytime. But because we continuously upgrade to the best open-weights models, you’re always on the frontier without switching.

TECHNOLOGY PARTNERS

Powered by Industry Leaders

We partner with the world's leading technology providers to deliver sovereign, enterprise-grade AI infrastructure.

NVIDIA

NVIDIA

GPU infrastructure partner powering our AI inference and training workloads. NVIDIA accelerated computing enables real-time model serving with unmatched performance.

Dell Technologies

Dell Technologies

Enterprise server and storage partner delivering the on-premise hardware backbone for sovereign AI deployments — secure, scalable, and built for production.

Learn More About Our Partnerships

In 2026, AI Is No Longer
a Competitive Advantage.
It's the Baseline.

Everyone has access to models. The real question is: are you building faster? Are you building smarter? Your competitors already are.

IG1 AI API by Iguana Solutions — World-class AI performance, fully sovereign, immediate access.