- Built by the Allen Institute for AI for large mixture-of-experts models.
- Switches from fully sharded data parallelism to distributed data parallelism.
- Delivers 2.7 times higher training throughput on tested hardware.
- Tested up to 1.2 trillion parameters across 512 GPUs.
In this guide
What is Olmo-core 3?
Olmo-core 3 is an open training infrastructure framework created by the Allen Institute for AI to train large mixture-of-experts language models. It serves as one of the foundation systems for upcoming Olmo releases while sharing the underlying tools publicly with researchers.
The framework targets the high communication and coordination costs that usually occur when scaling sparse models across GPU clusters. It is engineered to scale training into the trillion-parameter range without losing computational efficiency.
- Developer
- Allen Institute for AI (Ai2)
- Primary purpose
- Scalable training infrastructure for large MoEs
- Architecture shift
- From FSDP to DDP with resident experts
- Supported precision
- MXFP8 and BF16
- Maximum tested scale
- 2.38 trillion parameters (capacity test)
How does Olmo-core 3 handle MoE training?
Olmo-core 3 replaces the previous fully sharded data parallelism approach with distributed data parallelism. Instead of repeatedly gathering and resharding model weights across GPUs for every mini-batch, it keeps experts resident in GPU memory and routes token data directly to them.
Hardware distribution in Olmo-core 3 relies on three core methods. Expert parallelism divides the expert pool across GPUs, pipeline parallelism splits network layers across GPU groups, and a distributed optimizer divides optimizer states across hardware instead of duplicating them everywhere.
To cut routing overhead, the stack keeps routing metadata on GPUs to let CPUs queue tasks faster. It also relies on grouped GEMM to merge smaller expert computations and rowwise expert parallelism to place input data directly into expert buffers.
Olmo-core 3 training throughput and benchmarks
Olmo-core 3 delivers higher processing speeds compared to earlier framework releases. Allen Institute for AI reported that in an eight-GPU benchmark using NVIDIA B300 hardware, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, up from 19,400 tokens per second with the earlier FSDP implementation.
In capacity scaling tests, the team increased the expert count from 8 to 128 while routing four experts per token, keeping active parameters fixed at roughly 3.2 billion. Even as total model parameters jumped from 4.6 billion to 47 billion, training throughput declined by less than 5%.
MXFP8 support and memory efficiency
Olmo-core 3 integrates support for MXFP8, a low-precision format that uses fewer bits to store values. This reduction decreases the computational burden and lowers the volume of data transferred between GPUs during distributed training.
In a controlled test across four NVIDIA B300 GPUs, activating MXFP8 produced roughly 21% higher throughput than baseline BF16 precision. Peak active memory consumption dropped from 103 GiB down to 95 GiB, with most efficiency gains originating in feed-forward operations and expert data movement.
Allen Institute for AI notes that MXFP8 benefits apply only when numerical savings outweigh the computational costs of converting data between formats.
Scaling limits and testing observations
Olmo-core 3 has demonstrated scaling beyond one trillion parameters in system benchmarks. The framework reached 858 TFLOP/s per GPU across 512 NVIDIA B300 units running a 1.2-trillion-parameter model with 58.36 billion active parameters per token, using random routing to test system limits.
In experimental runs using DeepEP v2 communication, the framework configured up to 2.38 trillion parameters in a short capacity test. Researchers noted this experiment proved hardware scaling boundaries rather than sustained long-term training quality.
The research report also identified a failure mode termed token gerrymandering, where routing balance scores appeared to improve while physical distribution became less uniform. Additionally, reducing learning rates for individual experts because they process fewer tokens failed to improve model results.
Frequently asked questions
What is Olmo-core 3?
Who created Olmo-core 3?
How does Olmo-core 3 improve on earlier Olmo-core versions?
What precision formats does Olmo-core 3 support?
What is the largest scale tested on Olmo-core 3?
Sources: Hugging Face
Want your brand to be the answer AI gives?
See how ready your website is for ChatGPT, Gemini and Perplexity, free, in about ten seconds.