FreeToken

Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica | Aug 1, 2026

FreeToken is an edge-native Mixture-of-Experts serving engine for running frontier-scale open-weight models on personal and consumer hardware. It treats GPUs, CPUs, host memory, and interconnects as one elastic inference platform, with bandwidth-adaptive CPU–GPU co-execution, global expert caching, graph-compatible execution, semantic-aware KV and recurrent-state caching for agentic workloads, and runtime VRAM re-allocation between expert caches and KV memory.