Olmo-core 3 is a Large Language Models (LLMs) tool. Open-source framework for training trillion-parameter mixture-of-experts language models. Key features include Trillion-Parameter Scaling, 2.7x Training Speedup, and Distributed Data Parallelism. Best for scientists and researchers, software developers and engineers and data scientists and analysts.
About Olmo-core 3
OLMo-core 3 is an open-source training framework from the Allen Institute for AI that enables researchers to build and train large mixture-of-experts models at scale. It delivers 2.7x faster throughput and supports models beyond one trillion parameters.
Key Features
Trillion-Parameter Scaling.
2.7x Training Speedup.
Distributed Data Parallelism.
MXFP8 Precision Support.
Fully Open Source.
Production-Ready Infrastructure.
Frequently Asked Questions
OLMo-core 3 is a training framework for building large mixture-of-experts language models. It's designed for AI researchers and developers who want to train their own models at scale, supporting configurations from billions to trillions of parameters.
OLMo-core 3 achieves roughly 2.7 times higher throughput than the earlier FSDP-based implementation. In tests, a 47-billion-parameter model processed 52,000 tokens per second per GPU with the new system, compared to 19,400 tokens per second previously.
Yes, OLMo-core 3 is completely free and open source under the Apache 2.0 license. You can use, modify, and distribute it for both research and commercial purposes. All training code, documentation, and benchmarks are publicly available.
OLMo-core 3 is designed to work across diverse GPU configurations. It has been tested on NVIDIA B300 GPUs and can scale from small clusters to large deployments with hundreds of GPUs. The framework can be adapted to different hardware setups.







