Yandex Launches AliceAI-Foundation-80B-A3B-Base Language Model

Yandex has announced the release and open-sourcing of its new large language model, AliceAI-Foundation-80B-A3B-Base. Developed entirely from the ground up without relying on third-party open-source models, this initiative marks a significant stride towards Yandex’s unified reasoning model (URM), which will power the future agent capabilities of Alice AI.

Superior Performance and Innovative Development

In pre-training benchmarks, AliceAI-Foundation-80B-A3B-Base demonstrates impressive performance, outperforming larger models such as DeepSeek-V4-Flash-Base and Nemotron-3-Super across a wide range of tasks, including factual and expert questions, mathematics, and coding. On complex benchmarks like IMO AnswerBench and LiveCodeBench, the new model leads among comparable open pre-trained models. Furthermore, it surpasses Yandex’s previous closed model, Alice AI LLM 235B, in factual knowledge, mathematics, programming, and long-context handling, despite having nearly three times fewer total parameters and approximately seven times fewer active parameters.

The development of this model took approximately six months, a rapid timeline attributed to the team’s accumulated expertise. During this period, the datasets were reassembled, architectural solutions were refined, hyperparameters were optimized, and data for reasoning and agent interactions was integrated. This release follows Yandex’s recent open-sourcing of the Alice AI Search pre-train, extending their commitment to sharing advancements.

Comprehensive Technical Report and Open Benchmarks

Alongside the model weights, Yandex is publishing an extensive technical report detailing the data preparation pipelines, evaluation protocols, training stabilization, and scaling laws. The report includes ablation experiments for key decisions and discusses approaches that were ultimately discarded. The contributions of the updated corpus, architecture, and hyperparameters were validated through a series of scratch trainings, each involving 2 trillion tokens.

A dedicated section of the report addresses the complexities of measuring pre-train quality. Yandex is also open-sourcing two factual benchmarks—WikiWebFacts and HardMultiQA, with a particular focus on Russian-language context—along with their evaluation protocols. This initiative aims to enable the community to more accurately assess and compare models, especially for tasks requiring deep knowledge and reasoning in Russian.