Advancements in Artificial Intelligence
Sber has announced the release of GigaChat 3.5 Reasoning, a new iteration of its multimodal neural network. This model is positioned as Russia’s first open AI model featuring full reasoning capabilities, trained using online Reinforcement Learning (RL) technology. According to the company, GigaChat 3.5 Reasoning is uniquely designed to deliberate sequentially before providing an answer, manage task chains, and correct its own errors.
Architectural Innovations and Training Process
A pivotal change in the training methodology for GigaChat 3.5 Reasoning involved the entire model undergoing an online RL process after the Supervised Fine-Tuning (SFT) stage. Rather than a single domain or final stage, the training pipeline incorporated six distinct experts, each with specialized knowledge in areas such as mathematics, coding, and agent functionalities. Each expert was equipped with its own reward system, and subsequently, all were integrated back into a unified model. The comprehensive development and implementation of this sophisticated architecture extended over nine months.
Significance for the Russian AI Landscape
The introduction of GigaChat 3.5 Reasoning underscores Sber’s commitment to advancing cutting-edge artificial intelligence technologies. Its declared abilities in reasoning and self-correction represent a significant stride for the Russian open AI community, providing a tool with enhanced cognitive functionalities.
The deliberative capabilities of GigaChat 3.5 sound incredibly promising, especially the error correction feature. I’m curious about how the online Reinforcement Learning truly handles the integration of six ‘experts’ with individual reward systems without creating conflicting biases in the unified model. Also, will this deliberative process be transparent for users, allowing us to see the ‘thought process’ or chain of reasoning, or will it only present the final answer? I’d love to hear more about that!