All Publications
AutoUVM: Automated Prefetching Framework for LLMs under UVM Oversubscription
ICCD '26The 44th IEEE International Conference on Computer Design
LLMPROF: Identifying Performance Bottlenecks in LLM Serving Systems with Top-Down Profiling
SC '26The International Conference for High Performance Computing, Networking, Storage, and Analysis (Supercomputing)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
arXiv 2026Preprint
PASTA: A Modular Program Analysis Tool Framework for Accelerators
CGO '26The 23rd ACM/IEEE International Symposium on Code Generation and Optimization
Forest: Access-aware GPU UVM Management
ISCA '25The 52nd Annual International Symposium on Computer Architecture
Understanding Oversubscribed Memory Management for Deep Learning Training
EuroMLSys '25The 5th Workshop on Machine Learning and Systems
DrGPUM: Guiding Memory Optimization for GPU-accelerated Applications
ASPLOS '23The 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
Poster: Squeezing GPU Memory Usage in PyTorch
PyTorch Conference '22A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCs
TCAD '22IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems