small by design. open by nature.

Research preview

introducing
OpenSML-150m

Released · Research preview

OpenSML-150m

An English-first language model with open weights
and native MLX inference.

150.44Mparameters2,048context tokens
Read the release

From the ground up

the model is only
part of the story.

OpenSML explores how small language models are built: the data, the training, and the decisions along the way.

01 / Data & tokenizer

a vocabulary of its own.

Custom tokenizers and carefully prepared training data lay the groundwork for learning.

02 / Training

built on apple silicon.

We train models from scratch on Apple Silicon, exploring what accessible hardware can make possible.

03 / Evaluation

results you can inspect.

We share evaluations, training records and limitations so others can inspect the results and build on the work.

Explore the model research

Inside the training setup

distributed training

Our current cluster brings five Apple Silicon Macs together over a Thunderbolt RDMA ring. Explore the hardware, how the workload is split, and how every worker stays in sync.

Inside distributed training

Keep exploring

small enough to study.
open enough to build on.

Explore our models, follow the research,
and find a starting point for your own experiments.