Aleph Alpha Kolibri: Sovereign Open-Weight AI for Enterprise
On the Day of German Reunification, Aleph Alpha released Kolibri: an open-weight, EU-sovereign language model designed for mission-critical enterprise workloads in regulated industries. The weights are available today on Hugging Face under Apache 2.0, and the model was built entirely in Germany and Finland, under European law, with no foreign entity in the supply chain.
For enterprise teams in public administration, defense, aerospace, and financial services, this matters. Kolibri gives organizations a credible path to deploying frontier-class AI completely on their own infrastructure, without routing data through US or Chinese cloud services, and without legal uncertainty about where training data came from.
What Kolibri Is
Kolibri is a Mixture of Experts (MoE) Transformer with 78.1 billion total parameters and 3.46 billion active parameters per token. The sparsity is the key to its efficiency: each token is routed through six of the model’s 384 experts, keeping per-token compute low while maintaining large model capacity for complex reasoning tasks.
The context window reaches 1 million tokens in extended mode, trained up to 256,000 tokens. Kolibri uses a hybrid attention pattern combining sliding window attention (512 tokens) with full attention every fifth layer, allowing it to scale to long contexts without the quadratic cost of full attention throughout.
Training ran on 20 trillion tokens, selected and deduplicated from more than 200 trillion tokens of raw data. The model’s knowledge cutoff is June 18, 2026, for both English and German content.
Full specifications at launch:
| Attribute | Kolibri |
|---|---|
| Total parameters | 78.1B |
| Active parameters per token | 3.46B |
| Pre-training tokens | 20T |
| Context window (trained max) | 256k |
| Context window (extended) | 1M |
| Total experts | 384 |
| Active experts per token | 6 |
| License | Apache 2.0 |
| Languages | English, German |
| Knowledge cutoff | 18 June 2026 |
| Reasoning effort modes | None, Low, Medium, High |
Aleph Alpha published a technical report alongside the release with full training methodology and benchmark details.
Why Sovereign AI Has Real Enterprise Value
The word “sovereign” carries specific legal and operational meaning for regulated industries, and Kolibri earns the label with concrete engineering decisions rather than marketing positioning.
For a German government ministry, a defense contractor, or a financial institution operating under German banking law, routing AI inference through a US or Chinese hyperscaler creates genuine exposure: GDPR audit requirements, export control considerations, and board-level risk around foreign jurisdiction access to sensitive data. Cloud contract terms rarely provide the guarantees that regulated procurement requires.
Aleph Alpha trained Kolibri in Germany and Finland, under European and German law, with no foreign entity involved at any stage of the supply chain, from data curation through pre-training, post-training, and evaluation. Customers who deploy on-premises receive auditable provenance for every decision in the pipeline.
The EU AI Act, which applies to general-purpose AI models from August 2026, requires operators to document training data, maintain capability evaluations, and preserve technical documentation. Kolibri ships with a legal report addressing EU AI Act and GPAI Code of Practice requirements directly, making compliance documentation available at launch rather than left to operators to reconstruct. The Merlin-Arthur grounding protocol (described below) was also explicitly designed with the EU AI Act’s transparency requirements in mind.
This tracks a pattern Enera sees across enterprise AI deployments: the teams that move fastest are those with infrastructure they can actually trust and govern, not those chasing the highest benchmark scores. AI-native transformation requires that foundation to be solid.
Reasoning at Four Levels
Kolibri introduces four reasoning effort modes: none, low, medium, and high. Operators can dial inference cost and latency against answer quality depending on the task. A document retrieval query can run at low effort with sub-second response times. A complex multi-step legal or regulatory analysis can invoke high-effort reasoning with extended chain-of-thought traces.
The four-level system follows the direction the frontier AI industry has moved broadly in 2026. OpenAI’s reasoning sliders, Anthropic’s effort ladders, and now Kolibri’s four-mode system all reflect the same operational insight: enterprise workloads have highly variable reasoning requirements, and paying for maximum reasoning on every query is economically irrational. Giving developers runtime control over this tradeoff is now a table-stakes feature for serious enterprise models.
The Merlin-Arthur Protocol
One of Kolibri’s most practically valuable features for regulated industries is the Merlin-Arthur protocol. When the context provided to the model does not contain sufficient information to support a confident answer, the model abstains rather than generating a plausible-sounding but unsupported response.
For legal, financial, and government use cases, this is significant. An enterprise AI model that refuses to answer when it lacks sufficient grounding is substantially safer in a compliance context than one that fabricates citations or extrapolates beyond its evidence. Legal teams and compliance officers have consistently ranked hallucination management as a top-three concern when deploying AI in production workflows. The Merlin-Arthur protocol addresses this at the model level, rather than requiring operators to build their own downstream validation layer.
Aleph Alpha combines this with transparent reasoning traces: operators can inspect how the model arrived at its answer, not just what the answer was. That explainability is increasingly required by both regulators and enterprise risk committees.
Kolibri vs. General-Purpose Open-Weight Models
Several strong Apache 2.0 open-weight models have launched in 2026, including Tencent’s HY3 and numerous MoE releases from Chinese and US labs. Kolibri’s differentiation is not raw benchmark performance but a set of enterprise properties that general-purpose releases do not bundle:
German-native capability: Kolibri was specialized for German-language reasoning, argumentation, and document analysis. Its German-English knowledge cutoff is June 2026. English-first open-weight models have strong German performance, but not the same depth of German-specific post-training.
Supply chain integrity: Other Apache 2.0 releases document their architecture and license their weights, but they do not provide the legal report, training-data provenance documentation, and EU-jurisdiction training assurance that Kolibri provides. For procurement in German public administration, that documentation is often a hard requirement, not a nice-to-have.
Size efficiency: At 3.46B active parameters per token, Kolibri is small enough to run on enterprise on-premises infrastructure without hyperscale GPU clusters. That matters for regulated environments that cannot simply add compute from a cloud provider when demand grows.
The Cohere Connection
In late April 2026, Aleph Alpha and Cohere announced a merger described by observers as a de facto acquisition of Aleph Alpha by the Canadian AI company. The resulting joint venture is building sovereign AI offerings across public administration, finance, defense, energy, telecommunications, and healthcare, with a specific deployment path on the STACKIT cloud platform operated by the German Schwarz Group.
For Kolibri, this means enterprise customers who want a managed deployment option can access the model through STACKIT, a European-sovereign cloud. The Apache 2.0 open-weight release runs independently of any Cohere commercial arrangement: organizations can download and self-host now, without any vendor dependency.
German Federal Digital Minister Karsten Wildberger described the Cohere-Aleph Alpha combination as the creation of a “global AI champion” when the merger was announced. The Kolibri launch on German Reunification Day is a deliberate signal about where the joint venture is focused.
What Enterprises Can Do With Kolibri Now
The Apache 2.0 release is available immediately. No waitlist, no access request:
- Download the weights from Hugging Face at
aleph-alpha/Kolibri - Install the
aleph-alpha-inferencepackage, which bundles the required vLLM plugin - Deploy using the provided container image
ghcr.io/aleph-alpha/aleph-alpha-inference - Use the OpenAI-compatible API shape for integration with existing agent frameworks
Enterprise deployments with managed infrastructure, support, and EU AI Act documentation go through Aleph Alpha directly. Commercial interest routes to a demo request.
For teams evaluating sovereign AI infrastructure, Kolibri is now the most complete package available for European regulated contexts: open weights, Apache 2.0, EU-law training, built-in compliance documentation, and a credible managed deployment path. The EU AI Gigafactory program provides compute infrastructure to run it on at scale.
AI-native enterprise builders who need to reason through sovereign infrastructure choices for their own deployments can work with Enera to map compliance requirements to architecture decisions before they become production constraints.