introducing OpenSML-150m

OpenSML / 001

From scratch. On Apple Silicon.

150 million parameters.
one open experiment.

a small model, built from scratch.

OpenSML-150M is the first public release in the OpenSML family. It is an English-first language model built with Apple’s MLX framework, with a custom 32,000-token byte-level BPE tokenizer paired with educational pretraining data.

The released checkpoint combines two stages of pretraining with three stages of full-parameter supervised fine-tuning. It is shared as a research preview, alongside the code and records needed to inspect how it was built.

150.44MModel parameters
7.8BPretraining token positions
32KTokenizer vocabulary
2,048Context tokens

from pretraining to instruction tuning.

Pretraining used educational web text, DCLM-Edu, English FineWiki and Cosmopedia v2. Stage B continued the same model with an adjusted data mixture, reaching approximately 7.8B lifetime token positions.

PretrainingStage AInitial training
PretrainingStage BContinued training
Fine-tuning / 1Unified384+384 updates
Fine-tuning / 2Repair512+128 updates
Fine-tuning / 3Selected release+256 updates

Three supervised stages then shaped the model’s responses using conversational examples, question answering and instruction constraints. Every model parameter was updated; the selected weights are a direct continuation of the pretrained base.

Training began on four Mac Studios and later expanded to five, using Thunderbolt RDMA for distributed training. The technical report documents the hardware transition, dataset mixtures, training recipes and loss curves.

what changed after fine-tuning?

The same four zero-shot likelihood benchmarks were evaluated on the pretrained base and the selected fine-tuned checkpoint. Fine-tuning improved ARC scores; changes on PIQA and HellaSwag were mixed.

Pretrained baseReleased model

Percentage scores; higher is better. acc uses summed candidate log-likelihood. acc_norm normalizes it by candidate character count. These measure answer-choice scoring, rather than generated-answer correctness.

Full benchmark results, pretrained base compared with the released OpenSML-150M checkpoint
BenchmarkBase acc / normRelease acc / normExamples
ARC-Easy54.67% / 48.53%56.65% / 55.43%2,376
ARC-Challenge23.38% / 26.88%26.02% / 29.52%1,172
PIQA65.23% / 64.36%64.53% / 64.09%1,838
HellaSwag30.21% / 34.44%30.55% / 33.94%10,042
Evaluation settings and instruction following

Full-split zero-shot evaluations cover 15,428 multiple-choice examples: ARC test splits and PIQA/HellaSwag validation splits. The native evaluator uses FP32 scoring, reference MLX attention and raw completion prompts. It follows pinned harness conventions, rather than running the installed lm-evaluation-harness.

On IFEval, the released checkpoint scored 15.16% strict prompt accuracy and 25.30% strict instruction accuracy. Full scoring details and results are on the model card.

These results are specific to the published evaluation setup. The research preview can produce incorrect or repetitive responses. Benchmark results do not establish reliable factual answers, and no coding-success or MT-Bench score is reported for this checkpoint.

designed to be studied.

A decoder-only transformer with grouped-query attention, rotary position embeddings and a SwiGLU feed-forward network. Input and output embeddings share weights.

Transformer layers
20
Hidden width
768
Query / key-value heads
12 / 4
Feed-forward width
2,048
Tokenizer
Byte-level BPE · 32,000 tokens
Position encoding
RoPE · base 10,000
Normalization
RMSNorm + Q/K normalization
Inference
Native MLX · Apple Silicon

try the released model.

The Hugging Face repository contains the selected instruction-tuned weights, tokenizer and a standalone inference script. Use Python 3.11 or later on Apple Silicon:

python -m pip install huggingface_hub
hf download wzebrowski/OpenSML-150M --local-dir ./OpenSML-150M
cd OpenSML-150M
python -m pip install -r requirements.txt
python inference.py --prompt "Say hello in one sentence." --max-new-tokens 32

The native CLI verifies the weight and tokenizer hashes, uses greedy generation and formats prompts as User: … Assistant:. This release uses its own MLX loader; Transformers AutoModel and mlx-lm loading are not supported by this bundle. See the model card for full usage and provenance.

Explore the current five-Mac training setup ↗

explore the work behind the weights.

OpenSML-150M is available for anyone to use, modify, fine-tune and redistribute for personal, research or commercial purposes under Apache-2.0, covering rights held by the author. Third-party training materials retain their own terms; see the model card for data provenance and notices.