First Theory for PipeDream
Randomized PipeDream captures PipeDream's stale-weight behavior in an analyzable block-SGD model, revealing how pipeline depth affects convergence.
Hi, I’m
Machine Learning ResearcherPhD Candidate
I work on efficient optimization and compression methods for large language models, including pruning, sparse fine-tuning, quantization, and pipeline parallelism.
Open to machine learning research internships

Research areas
My research focuses on making large models more efficient through compression, adaptation, and distributed training.
Methods for removing redundant model parameters while preserving model quality.
Training a small, carefully selected subset of model parameters efficiently.
Reducing model precision and memory requirements while controlling quality degradation.
Optimization and convergence analysis for training models across pipeline stages with delayed updates.
Here are some of my favorite projects, from research implementations to open-source tools.
Randomized PipeDream captures PipeDream's stale-weight behavior in an analyzable block-SGD model, revealing how pipeline depth affects convergence.
A block-wise algorithm for pruning large language models using second-order information and coordinated weight compensation.
We introduce Super, which selects a sparse trainable support using activation-aware pruning scores, and Supra, a matched-budget sparse-plus-LoRA adapter.
An open-source, local-first tool that turns retake-heavy narration into a clean audio or video edit while preserving the speaker's real voice.
Research milestones, project updates, thoughts, and talks from recent years.
My proposal connects LLM pruning, sparse fine-tuning, and pipeline-parallel optimization into one research program.
A USD 100,000 Small Translational Research Grant will support TRACE, a project for turning textbooks into verified, interactive courses.
This fall I am supporting Design and Analysis of Algorithms through office hours, assignments, grading, and exams.