Zihan Zhao · Ph.D. Candidate, UC San Diego & CERN

I design efficient transformers: sub-quadratic attention, long context, and fast inference.

I'm a Ph.D. candidate at UC San Diego (UCSD) advised by Javier Duarte, working with the CMS experiment at CERN's Large Hadron Collider. I build attention architectures, pretrain and train at scale on GPU clusters, and evaluate how LLMs reason. Previously I was an ML research intern at Futurewei.

I'm looking for research roles building frontier models.

ZZZihan Zhao

Selected workArchitecture, training, and evaluation. * denotes equal contribution.

Two panels: the convolution's perplexity gain grows with context length under windowed attention, and tracks the window penalty across window sizes at 8k context.
Sole author

Structure Before Attention: A Zero-Initialized Convolution Recovers Windowed Language-Model Quality

Z. Zhao

Sliding-window attention makes long context cheaper but costs perplexity. A small zero-initialized convolution before attention cuts perplexity by 2.1% at 8k context with windowed attention, versus 0.6% with dense attention, and the windowed model lands within 0.16% of its dense counterpart. Paired-seed attribution (identical init, data order, and eval batches) resolves these sub-1% effects below seed variance.

PHAT-JeT: geometric message passing on the detector grid, and hierarchical attention with local attention inside patches and global attention across patch tokens.
NeurIPS 2026Co-first author

PHAT-JeT: Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging

A. Wang*, Z. Zhao*, A. Xia*, C. Sun, A. Gandrakota, J. Ngadiuba, R. Cavanaugh, J. Duarte

Exact attention within small patches, global communication through hierarchical patch tokens, and geometric message passing for local structure. Cuts attention from O(N²) to O(N) and reaches state-of-the-art accuracy at ~6.4k parameters, targeting microsecond-latency inference.

Example ill-defined problem next to a bar chart showing accuracy drops from well-defined to ill-defined problems for four LLMs.
COLM 2026Co-first author

Ill-Defined Math: Benchmarking LLM Reasoning Beyond Well-Defined Problems

H. Chen*, Y. Lin*, Z. Zhao*, P. Chen*, N. Liu, Y. Hu, Q. Xie, Q. Bai, N. Yan, M. S. Mortazavi, K. Youcef-Toumi

1,300 mathematically ill-defined problems, including a 300-problem expert-verified test set, and a three-stage LLM judge that reaches 1.00 / 0.94 F1 against human labels. Across 27 frontier LLMs, models often recognize the flaw mid-reasoning yet still commit to a definitive answer.

PULSE framework: a privileged EDA teacher is used during pretraining and knowledge transfer, then discarded so only cheap sensor students run at inference.
CVPR 2026 Sense of Space Workshop · OralFirst author

PULSE: Privileged Knowledge Transfer from Rich to Deployable Sensors for Embodied Multi-Sensory Learning

Z. Zhao, K. Pendiyala, M. Mortazavi, N. Yan

A rich teacher sensor (EDA) is used only during training, then discarded. Per-modality masked autoencoders split each student's embedding into shared and private parts, and the shared subspace is distilled from the frozen teacher at multiple layers, so only cheap deployable sensors run at inference.

News

More research

Three panels: accuracy vs number of experts under capacity limits, accuracy vs forward GFLOPs, and QCD rejection vs active experts.
Under review · ML4PS 2026Mentored project

Conditional Capacity and Routing in Mixture-of-Experts Particle Transformers

K. Pendiyala*, H. Zia*, T. Lee*, T. Legge, A. J. De Leon, Z. Zhao, A. Wang, A. Gandrakota, J. Ngadiuba, R. Cavanaugh, J. Duarte

Sparse MoE vs. dense transformers on 188-class JetClass-II. No-drop top-1 MoE gains ~1 pp at nearly unchanged FLOPs; more stored experts add little, and routing structure does not track accuracy.

SAL-T architecture: the LPP-MHA attention module with partitioned key and value projections and a convolution over attention scores, inside the full model stack.
Under review · PRX IntelligenceCo-first author

Spatially Aware Linear Transformer (SAL-T) for Particle Jet Tagging

A. Wang*, Z. Zhao*, S. Katel, V. G. Sahu, E. E. Khoda, A. Gandrakota, J. Ngadiuba, R. Cavanaugh, J. Duarte

Linear attention with locality-preserving token ordering, partitioned key/value projections, and depthwise convolution over attention maps. Matches full-attention accuracy with lower FLOPs and latency.

Deep-Reproducer pipeline: paper breakdown, deep research over text, images, papers and code repos, a memory bank, modular code generation, and self-evolving debugging.
DL4C @ NeurIPS 2025

Deep-Reproducer: From Paper Understanding to Code Generation

P. Chen, N. Yan, Z. Zhao, Y. Lin, H. Chen, Y. Hu, Q. Bai, X. Li, M. S. Mortazavi

A multi-agent LLM system that turns research papers into working codebases, reaching a 63.2% replication score on PaperBench. I designed its self-evolving debugging loop: locate, patch, execute, test.

Histogram of Particle Transformer attention scores concentrated near 0 and 1, and an attention map in the eta-phi plane.
ML4PS @ NeurIPS 2025

Why Is Attention Sparse in Particle Transformer?

T. Legge, A. Wang, J. Ortiz, V. Limouzi, Z. Zhao, A. Gandrakota, E. E. Khoda, J. Ngadiuba, J. Duarte, R. Cavanaugh

Near-binary attention in ParT comes mainly from the attention mechanism itself, not the physics-inspired interaction matrix, and it singles out key jet substructure.

Gap between representation and raw-input error by task-generator class, for four encoders.
Under review · ML4PS 2026Mentored project

Information-Loss Audits of Learned Representations Depend on the Task Generator

Audits of what a representation discards mostly reflect the task-generator class, not the representation. Reporting a witness set recovers a planted factor on 3/4 encoders where the top witness misses it.

Experience & education

Futurewei Technologies

Jun – Sep 2025
ML Research Intern, Intelligent Computing Lab

Built PULSE (oral, CVPR 2026 Sense of Space Workshop) and the self-evolving debugging loop of Deep-Reproducer; contributed to Ill-Defined Math. Outstanding Performance Intern.

University of California San Diego

2022 – present
Ph.D. in Physics, Machine Learning emphasis · Advisor: Javier Duarte

Efficient transformers, self-supervised pretraining, and ML for the CMS experiment at CERN, including a boosted H → bb̄ measurement with LHC Run 3 data.

University of Southern California

2018 – 2022
B.S. Physics, B.A. Mathematics, minor in Computer Programming

GPA 3.97 / 4.0. Dean's List every semester.

Toolkit

PyTorch · multi-GPU distributed training (DDP) · bf16 mixed precision · FlexAttention · CUDA-aware optimization · Slurm, Kubernetes, Docker · Python, C++

Mentoring

Scoped and supervised research for 10+ students. Mentees went on to Ph.D. programs at Princeton and UIUC and co-authored papers at PAI 2026 and ML4PS 2026.

Service

Reviewer for ICLR 2027 and PAI 2026; workshop reviewer for ML4PS 2026, Interp4Discovery and MATH-AI at NeurIPS 2026. Teaching assistant for physics courses at UC San Diego.

Talks