Running a 20B MoE Model at 120 Tokens Per Second on an iPhone
Running a capable AI model on a phone used to mean accepting severe limitations. Small models, slow speeds, and battery drain defined the experience. That story changed dramatically in 2026. A new generation of on-device models, led by innovations like the Maple-Preview MoE architecture, now deliver