Back to Home

All Publications

AutoUVM: Automated Prefetching Framework for LLMs under UVM Oversubscription

ICCD '26

Mao Lin, Hui Feng, Xianzhong Ding, Guilherme Cox, Qian Wang, and Hyeran Jeon

The 44th IEEE International Conference on Computer Design

LLMPROF: Identifying Performance Bottlenecks in LLM Serving Systems with Top-Down Profiling

SC '26

Tianle Zhong, Mao Lin, Hao Wu, Keren Zhou, and Geoffrey Fox

The International Conference for High Performance Computing, Networking, Storage, and Analysis (Supercomputing)

HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing

arXiv 2026

Mao Lin, Xi Wang, Guilherme Cox, Dong Li, and Hyeran Jeon

Preprint

PASTA: A Modular Program Analysis Tool Framework for Accelerators

CGO '26

Mao Lin, Hyeran Jeon, and Keren Zhou

The 23rd ACM/IEEE International Symposium on Code Generation and Optimization

Forest: Access-aware GPU UVM Management

ISCA '25

Mao Lin, Yuan Feng, Guilherme Cox, and Hyeran Jeon

The 52nd Annual International Symposium on Computer Architecture

Understanding Oversubscribed Memory Management for Deep Learning Training

EuroMLSys '25

Mao Lin and Hyeran Jeon

The 5th Workshop on Machine Learning and Systems

DrGPUM: Guiding Memory Optimization for GPU-accelerated Applications

ASPLOS '23

Mao Lin, Keren Zhou, and Pengfei Su

The 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems

Poster: Squeezing GPU Memory Usage in PyTorch

PyTorch Conference '22

Mao Lin, Keren Zhou, and Pengfei Su

A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCs

TCAD '22

Zelin Du, Qianling Zhang, Mao Lin, Shiqing Li, Xin Li, and Lei Ju

IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems