Hi, I am Jian Gu. I completed my PhD program at Monash University, under the supervision of Prof. Aldeida Aleti, Prof. Chunyang Chen, and Prof. Hongyu Zhang. Before that, I earned master’s degree in machine learning at KTH and bachelor’s degree in computer science (elite class) at SDU.
My research investigates the interplay between software engineering and machine learning, especially how engineering principles guide practical analysis and repair techniques to LMs, namely LM Analysis and Repair. I always study LM interpretability as the underlying mechanism, focusing on the inherent semantics, termed as LM Semantics. Proudly share our series of work: Semantic Basis, Semantic Transition, Semantic Alignment, Pseudo-Inverse Tying, Semantic Reference Frame, …
🔥 News
- 2026.07: ✌️ I was granted Chinese Government Award for Outstanding Self-Financed Students Abroad
- 2025.12: ✌️ I defended my PhD thesis on Semantics-Based Analysis and Repair of Neural Language Models
- 2025.09: 🚁 I started research interns at Huawei HKRC, AntGroup CodeFuse
- 2024.10: 🚁 I started visiting trips to NLP Lab @ Tsinghua, SE&AI Lab @ TUM
💻 Featured Work

Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
Pseudo-Inverse Tying is a novel weight-tying approach that improves the semantic coherence of language models by keeping input and output token geometries synchronized during training, enhancing stability, interpretability, and lightweight adaptation.

Semantic-based Optimization for Repairing LLMs: Case Study on Code Generation
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
STAR is a novel semantic-based optimization approach for LM repair that efficiently locates and patches buggy neurons using statistical insights and analytical formulas, outperforming prior methods in effectiveness, efficiency, and minimizing side effects.

Semantic-Aware Layer-Freezing for Computation-Efficient Fine-Tuning of LMs
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
Our semantic-based layer freezing approach improves the efficiency of language model finetuning by determining where to finetune, outperforming existing methods through a detailed semantic analysis of the model’s inference process.

Neuron Patching: Semantic-based Neuron-level LM Repair for Code Generation
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
MINT is an efficient and reliable technique for repairing large language models in software engineering. It can successfully solve model failures by patching merely 1 or 2 neurons, outperforming state-of-the-art methods in coding tasks.

Towards Top-Down Automated Development in Limited Scopes: A Neuro-Symbolic Framework from Expressibles to Executables
Jian Gu, Harald C. Gall
Deep code generation integrates neural models into software engineering for generating code but requires enhancements for project-level tasks, suggesting a taxonomy on code data and introducing a semantic pyramid framework to improve software development processes.
📝 Selected Papers
Language Model Semantics
ArXivSemRF: A Semantic Reference Frame for Residual-Stream Dynamics in LMs
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangACL'26Cross-Scale Knowledge Transfer in Language Models through Latent Semantic Alignment
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangACL'25Semantic-Aware Layer-Freezing for Computation-Efficient Fine-Tuning of LMs
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangArXivVocabulary-Defined Semantics: Latent Space Clustering for Beyond-Context Learning
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
Software Engineering for Machine Learning (SE4AI)
ArXivRethinking Weight Tying: Pseudo-Inverse Tying for Stable LM Training and Updates
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangICSE'26Semantic-based Optimization for Repairing LLMs: Case Study on Code Generation
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangArXivFocus-Aware Neurons: Contextual LM Repair leveraging Selective Gating Attention
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu ZhangTOSEM 2026Neuron Patching: Semantic-based Neuron-level LM Repair for Code Generation
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
Machine Learning for Software Engineering (AI4SE)
FSE'23 (IVR)Towards Top-Down Automated Development in Limited Scopes:
A Neuro-Symbolic Framework from Expressibles to Executables
Jian Gu, Harald C. GallSANER'22Assemble Foundation Models for Automatic Code Summarization
Jian Gu, Pasquale Salza, Harald C. GallICSME'21Multimodal Representation for Neural Code Search
Jian Gu, Zimin Chen, Martin MonperrusASE'26Source-Free Detection and Impact Analysis of Compiler Optimization in Mobile Apps
Han Hu, Xiaoheng Xie, Bo Sun, Jian Gu, Gang Fan, Li LiTSE 2022On the Effectiveness of Transfer Learning for Code Search
Pasquale Salza, Christoph Schwizer, Jian Gu, Harald C. GallTSE 2021Automated Classification of Overfitting Patches with Statically Extracted Code Features
He Ye, Jian Gu, Matias Martinez, Thomas Durieux, Martin Monperrus
🔎 Services
- Software Engineering: ACM TOSEM (Reviewer), IEEE TSE (Reviewer)
- Machine Learning: AAAI (PC Member), ARR (Reviewer), ICLR (Reviewer), NIPS (Reviewer)
- Multidisciplinary: WWW (PC Member), IEEE TAI (Reviewer)
“Machine intelligence is the last invention that humanity will ever need to make. Machines will then be better at inventing than we are, and they’ll be doing so on digital timescales.” – Nick Bostrom