An oil painting of an old master guiding a young clerk who works on an open MacBook Pro among quills and an Imari cup.

Model Train

Teaching a model to use a Mac

A dataset for the Mac software companies use today, and a Qwen3.5-9B model trained on it.

Dataset
15,854 steps in 450 tasks
Base model
Qwen3.5-9B
Action score
0.33 → 0.77

The data

We turned 175 of the most-watched tutorials for macOS, video editing, Blender and Godot into 450 tasks: every click, keystroke and drag, with the reason for it and the exact point on the screen.

Labels are made by models and audited by sampling. In spot checks of 1,015 steps, about 90% of clicks land on the right target and about 70% of actions are fully right.

The model

We chose Qwen3.5-9B because it can read a screen and still runs on a single Mac, then trained a LoRA adapter on the dataset.

On a frozen 477-step offline test, its action score rose from 0.33 to 0.77. This is our own offline test, not an official benchmark, and part of the gain comes from learning the answer format.

What comes next

Live evaluation on real tasks, where success, recovery and wasted actions all count, not only whether the next action matches.

Every commission begins with a conversation.

Tell us what the work is, what data you keep and where it has to run. We reply within two business days.

Start a projectbiz@random-walk.co.jp