Shawn Im

Shawn Im

Hello! My interest is in developing an understanding of the systems underlying the behavior of machine learning models. Not only is this understanding crucial for safe and beneficial models, but also, as models grow to understand more about the world and human values, a strong understanding of their behavior has the potential to lead to new insights of the world and ourselves. Currently, I am working on understanding how models represent their beliefs and reasoning as well as developing ways for us to better explore the shape of the model's internals.

I am a PhD student at UW-Madison working with Prof. Sharon Li. I am thankful to be supported by the NSF GRFP. Previously, I was a Machine Learning Research intern at Apple. I completed my undergraduate degrees in mathematics and computer science at MIT.

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability
Shawn Im, Changdae Oh, Zhen Fang, Yixuan Li
International Conference on Learning Representations (ICLR) Oral, 2026

[Paper][Code]

Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
James Oldfield, Shawn Im, Yixuan Li, Mihalis A. Nicolaou, Ioannis Patras, Grigorios G Chrysos
Neural Information Processing Systems (NeurIPS), 2025

[Paper][Code]

Understanding the Learning Dynamics of Alignment with Human Feedback
Shawn Im, Yixuan Li
In Proceedings of International Conference on Machine Learning (ICML), 2024

[Paper] [Code]