Futurewei Technologies
Jun – Sep 2025Built PULSE (oral, CVPR 2026 Sense of Space Workshop) and the self-evolving debugging loop of Deep-Reproducer; contributed to Ill-Defined Math. Outstanding Performance Intern.
I design efficient transformers: sub-quadratic attention, long context, and fast inference.
I'm a Ph.D. candidate at UC San Diego (UCSD) advised by Javier Duarte, working with the CMS experiment at CERN's Large Hadron Collider. I build attention architectures, pretrain and train at scale on GPU clusters, and evaluate how LLMs reason. Previously I was an ML research intern at Futurewei.
I'm looking for research roles building frontier models.


Sliding-window attention makes long context cheaper but costs perplexity. A small zero-initialized convolution before attention cuts perplexity by 2.1% at 8k context with windowed attention, versus 0.6% with dense attention, and the windowed model lands within 0.16% of its dense counterpart. Paired-seed attribution (identical init, data order, and eval batches) resolves these sub-1% effects below seed variance.

Exact attention within small patches, global communication through hierarchical patch tokens, and geometric message passing for local structure. Cuts attention from O(N²) to O(N) and reaches state-of-the-art accuracy at ~6.4k parameters, targeting microsecond-latency inference.

1,300 mathematically ill-defined problems, including a 300-problem expert-verified test set, and a three-stage LLM judge that reaches 1.00 / 0.94 F1 against human labels. Across 27 frontier LLMs, models often recognize the flaw mid-reasoning yet still commit to a definitive answer.

A rich teacher sensor (EDA) is used only during training, then discarded. Per-modality masked autoencoders split each student's embedding into shared and private parts, and the shared subspace is distilled from the frozen teacher at multiple layers, so only cheap deployable sensors run at inference.

Sparse MoE vs. dense transformers on 188-class JetClass-II. No-drop top-1 MoE gains ~1 pp at nearly unchanged FLOPs; more stored experts add little, and routing structure does not track accuracy.

Linear attention with locality-preserving token ordering, partitioned key/value projections, and depthwise convolution over attention maps. Matches full-attention accuracy with lower FLOPs and latency.

How pretraining data scale shapes downstream performance after finetuning, with throughput gains from fewer CPU–GPU syncs, better matmul kernel utilization, and mixed precision.

A multi-agent LLM system that turns research papers into working codebases, reaching a 63.2% replication score on PaperBench. I designed its self-evolving debugging loop: locate, patch, execute, test.

Near-binary attention in ParT comes mainly from the attention mechanism itself, not the physics-inspired interaction matrix, and it singles out key jet substructure.

Self-supervised pretraining for set-structured inputs without hand-crafted augmentations, using redesigned masking and ordered positional encodings.

Audits of what a representation discards mostly reflect the task-generator class, not the representation. Reporting a witness set recovers a planted factor on 3/4 encoders where the top witness misses it.

Gaussian, independent-mode models fail from structured inter-mode dependencies rather than marginal non-Gaussianity, giving a simple test for when richer generative models are needed.
Built PULSE (oral, CVPR 2026 Sense of Space Workshop) and the self-evolving debugging loop of Deep-Reproducer; contributed to Ill-Defined Math. Outstanding Performance Intern.
Efficient transformers, self-supervised pretraining, and ML for the CMS experiment at CERN, including a boosted H → bb̄ measurement with LHC Run 3 data.
GPA 3.97 / 4.0. Dean's List every semester.
Toolkit
PyTorch · multi-GPU distributed training (DDP) · bf16 mixed precision · FlexAttention · CUDA-aware optimization · Slurm, Kubernetes, Docker · Python, C++
Mentoring
Scoped and supervised research for 10+ students. Mentees went on to Ph.D. programs at Princeton and UIUC and co-authored papers at PAI 2026 and ML4PS 2026.
Service
Reviewer for ICLR 2027 and PAI 2026; workshop reviewer for ML4PS 2026, Interp4Discovery and MATH-AI at NeurIPS 2026. Teaching assistant for physics courses at UC San Diego.
Efficient Attention Architectures for Real-Time Jet Tagging at the LHC: SAL-T and PHAT-JeT. Best talk award.
Large-Scale Pretraining and Finetuning for Efficient Jet Classification.