a vocabulary of its own.
Custom tokenizers and carefully prepared training data lay the groundwork for learning.
small by design. open by nature.
An English-first language model with open weights
and native MLX inference.
From the ground up
OpenSML explores how small language models are built: the data, the training, and the decisions along the way.
Custom tokenizers and carefully prepared training data lay the groundwork for learning.
We train models from scratch on Apple Silicon, exploring what accessible hardware can make possible.
We share evaluations, training records and limitations so others can inspect the results and build on the work.
Inside the training setup
Our current cluster brings five Apple Silicon Macs together over a Thunderbolt RDMA ring. Explore the hardware, how the workload is split, and how every worker stays in sync.
Inside distributed trainingKeep exploring
Explore our models, follow the research,
and find a starting point for your own experiments.