
Model Train
Teaching a model to use a Mac
A dataset for the Mac software companies use today, and a Qwen3.5-9B model trained on it.
- Dataset
- 15,854 steps in 450 tasks
- Base model
- Qwen3.5-9B
- Action score
- 0.33 → 0.77
The data
We turned 175 of the most-watched tutorials for macOS, video editing, Blender and Godot into 450 tasks: every click, keystroke and drag, with the reason for it and the exact point on the screen.
Labels are made by models and audited by sampling. In spot checks of 1,015 steps, about 90% of clicks land on the right target and about 70% of actions are fully right.
The model
We chose Qwen3.5-9B because it can read a screen and still runs on a single Mac, then trained a LoRA adapter on the dataset.
On a frozen 477-step offline test, its action score rose from 0.33 to 0.77. This is our own offline test, not an official benchmark, and part of the gain comes from learning the answer format.
What comes next
Live evaluation on real tasks, where success, recovery and wasted actions all count, not only whether the next action matches.
Every commission begins with a conversation.
Tell us what the work is, what data you keep and where it has to run. We reply within two business days.