Publications

My work spans the stack, from reusable GPU primitives to algorithm–system co-design to complete serving systems, for both language models and multimodal generation. Browse it as a research tree that shows how the projects build on each other, or as a list of all papers.

FreeToken
13.8k stars Edge Serving MoE
FreeToken
Efficient edge-native MoE serving with bandwidth-adaptive execution.
FlashLib
GPU Library Classical ML Operators Agent-Native API
FlashLib
Bringing flash magic to classical machine learning operators.
Flash-KMeans
NeurIPS 2026 Exact K-Means Kernel Optimization
Flash-KMeans
Fast and memory-efficient exact K-Means.
Quant VideoGen
ICML 2026 Long Video KV Cache Quantization
Quant VideoGen
Auto-regressive long video generation via 2-bit KV-cache quantization.
BlendServe
ASPLOS 2026 Offline Inference LLM Serving
BlendServe
Optimizing offline inference for autoregressive large models with resource-aware batching.
StreamDiffusionV2
MLSys 2026 Best Paper Interactive Video Streaming System
StreamDiffusionV2
A streaming system for dynamic and interactive video generation.
vAttention
ICLR 2026 Verified Sparsity Sparse Attention
vAttention
Verified sparse attention.
SLA
ICLR 2026 Sparse-Linear Attention Diffusion Transformers
SLA
Beyond sparsity in diffusion transformers via fine-tunable sparse-linear attention.
Sparse VideoGen2
NeurIPS 2025 Spotlight Semantic Permutation Video Generation
Sparse VideoGen2
Accelerating video generation with sparse attention via semantic-aware permutation.
Radial Attention
NeurIPS 2025 Long Video Sparse Attention
Radial Attention
O(n log n) sparse attention with energy decay for long video generation.
Sparse VideoGen
ICML 2025 Sparse Attention Video Generation
Sparse VideoGen
Accelerating video diffusion transformers with spatial-temporal sparsity.
Prism
OSDI 2026 GPU Sharing Multi-LLM Serving
Prism
Unleashing GPU sharing for cost-efficient multi-LLM serving.
UCCL
OSDI 2026 GPU Networking Transport Layer
UCCL
An extensible software transport layer for GPU networking.
WorldModelBench
NeurIPS 2025 Benchmark World Models
WorldModelBench
Judging video generation models as world models.
Twilight
NeurIPS 2025 Spotlight Adaptive Sparsity Long Context
Twilight
Adaptive attention sparsity with hierarchical top-p pruning.
HashAttention
ICML 2025 Semantic Sparsity Sparse Attention
HashAttention
Semantic sparsity for faster inference.
Post-Training Sparse Attention with Double Sparsity
Sparse Attention KV Cache LLM Inference
Post-Training Sparse Attention with Double Sparsity
Sparse attention for reducing KV-cache bandwidth in LLM inference.
S-LoRA
MLSys 2024 LoRA Serving CUDA Kernels
S-LoRA
Serving thousands of concurrent LoRA adapters.
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Data Quality Benchmark Contamination Evaluation
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Decontamination and benchmark overlap analysis for language models.