Related Professors at NYU
A collection of related professors. Inclusion does not imply an
affiliation with the association.
Substantial safety focus
-
Pavel Izmailov.
Reinforcement learning, reasoning, and alignment
. Researcher at Anthropic, formerly OpenAI superalignment.
-
Greg Durrett.
NLP and machine learning. Many papers on alignment
and interpretability.
-
Qi Lei. Weak-to-strong
generalization, distribution shift, and inference-time safeguards
. Google DeepMind faculty.
-
He He. Reward hacking
, control
, and monitoring
. (On sabbatical fall '27)
-
Sam Bowman.
Alignment. Anthropic. (Long-term leave)
-
Sydney Levine. Psychology professor working at "the intersection of cognitive
science of human moral judgement and AI safety". Visiting scientist
at Google Deepmind. Recent papers in alignment, value prioritization evaluation, and meta cognition.
-
Sebastian Benthall. Research
director, Information Law Institute. Work on casual decision
processes
and societal lock-in
.
Occasional safety work
-
Jonathan Mann. Adjunct professor and Good Judgement superforecaster. Writes
about forecasting and AI safety on Substack.
-
Kyunghyun Cho. Mainly focuses
on AI for biology and medicine, but has published on alignment, unlearning, and interpretability.
-
Joan Bruna. Mostly
deep learning theory. One interpretability paper.
-
Eunsol Choi. Partial focus on
continual learning. Papers on LLM behavioral evaluation
and implicit user feedback for training.
-
Arthur Jacot. Deep learning theory
, learning mechanics
.
-
Mengye Ren. Mostly
representation and continual learning. Papers on scaleable
oversight
and forecast calibration.
-
Julia Kempe. Some work on
alignment, mechanistic interpretability, and hallucination detection. Director of research at AmiLabs.
-
Chinmay Hegde .
Red-team jailbreaking
and unlearning
.
-
Siddharth Garg .
Red-team jailbreaking
and interpretability of insecure code generation.
-
Andrew Gordon Wilson.
Mostly theory of deep learning, but has touched out-of-distribution
detection
and hallucination detection.
-
Rajesh Ranganath.
Quantifying tail-risk of LLM outputs
and explanation faithfulness.
-
Ramesh Karri. Security and trust for hardware, lately encompassing LLMs. Recent
papers on jailbreak benchmarking
and unlearning
-
Gautam Kamath. Mostly
differential privacy, but has published on data poisoning
and unlearning.
-
Tal Linzen. Interpretability[22]
and introspection.
-
Haneul Yoo. Natural
language processing, mostly focusing on non-English languages. One
paper on jailbreaking LLMs by switching languages
.
Safety-adjacent work
-
Understanding Reasoning from Pretraining to Post-Training
(Shen et al., 2026)
↩
-
Weak-to-Strong Generalization: Eliciting Strong Capabilities With
Weak Supervision
(Burns et al., 2023)
↩
-
SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models
through Combat
(Jiang et al., 2025)
↩
-
From Distributional to Overton Pluralism: Investigating Large
Language Model Alignment
(Lake et al., 2025)
↩
-
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in
LLM-Based Evaluation
(Tripathi et al., 2025)
↩
-
Understanding Synthetic Context Extension via Retrieval Heads
(Zhao et al., 2025)
↩
-
Does Weak-to-strong Generalization Happen under Spurious
Correlations?
(Liu, et al., 2026)
↩
-
Bridging Distribution Shift and AI Safety: Conceptual and
Methodological Synergies
(Liu, et al., 2026)
↩
-
Think Twice, Generate Once: Safeguarding by Progressive
Self-Reflection
(Phan, et al., 2025)
↩
-
Language Models Learn To Mislead Humans via RLHF
(Wen et al., 2024)
↩
-
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
(Wen et al., 2024)
↩
-
Monitoring Decomposition Attacks in LLMs with Lightweight
Sequential Monitors
(Yueh-Han et al., 2025)
↩
-
Resource Rational Contractualism Should Guide AI Alignment
(Levine et al., 2026)
↩
-
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values
Prioritization with AIRiskDilemmas
(Chiu et al., 2025)
↩
-
Imagining and building wise machines: The centrality of AI
metacognition
(Johnson et al., 2026)
↩
-
The Design and Composition of Structural Causal Decision Processes
(Benthall & Lujan, 2026)
↩
-
AI Wardens
(Diamantis & Benthall, 2026)
↩
-
Abstraction (Substack)
↩
-
Why Alignment Must Precede Distillation: A Minimal Working
Explanation
(Cha & Cho, 2025)
↩
-
Reference-Specific Unlearning Metrics Can Hide the Truth: A
Reality Check
(Cho et al., 2025)
↩
-
The geometry of prompting: Unveiling distinct mechanisms of task
adaptation in language models
(Kirsanov et al., 2025)
↩
-
Emergence of Linear Truth Encodings in Language Models
(Bai et al., 2025)
↩
↩
-
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in
Language Models
(Jhaveri et al., 2026)
↩
-
User Feedback in Human-LLM Dialogues: A Lens to Understand Users
But Noisy as a Learning Signal
(Liu et al., 2025)
↩
-
How DNNs break the Curse of Dimensionality
(Jacot et al., 2025)
↩
-
Which Frequencies do CNNs Need? Emergent Bottleneck Structure in
Feature Learning
(Wen, Jacot, 2026)
↩
-
There Will Be a Scientific Theory of Deep Learning
(Simon et al., 2026)
↩
-
When Does Verification Pay Off? A Closer Look at LLMs as Solution
Verifiers
(Lu et al., 2025)
↩
-
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator
for LLM Forecasting
(Dai et al., 2026)
↩
-
PILAF: Optimal Human Preference Sampling for Reward Modeling
(Feng et al., 2025)
↩
-
From Concepts to Components: Concept-Agnostic Attention Module
Discovery in Transformers
(Su et al., 2026)
↩
-
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in
Long-Form Text Generation
(Bhardwaj, et al., 2026)
↩
-
Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical
Insertion Prompting
(Kulshreshtha, et al., 2026)
↩
-
When Are Concepts Erased From Diffusion Models?
(Lu, et al., 2025)
↩
-
MetaCipher: A Time-Persistent and Universal Multi-Agent Framework
for Cipher-Based Jailbreak Attacks for LLMs
(Chen, et al., 2026)
↩
-
Detecting All-to-One Backdoor Attacks in Black-Box DNNs via
Differential Robustness to Noise
(Fu, et al., 2025)
↩
-
Surgical Repair of Insecure Code Generation in LLMs
(Sandoval, et al., 2026)
↩
-
Out-of-Distribution Detection Methods Answer the Wrong Questions
(Li, et al., 2025)
↩
-
Automatic Layer Selection for Hallucination Detection
(Wang, et al., 2026)
↩
-
Estimating Tail Risks in Language Model Output Distributions
(Angell, et al., 2026)
↩
-
What LLMs explain is not what they believe: Evaluating explanation
sufficiency under models' own input beliefs
(Nguyen, et al., 2026)
↩
-
HarmChip: Evaluating Hardware Security Centric LLM Safety via
Jailbreak Benchmarking
(Wang, et al., 2026)
↩
-
SALAD: Systematic Assessment of Machine Unlearning on LLM-Aided
Hardware Design
(Wang, et al., 2025)
↩
-
Sequential Data Poisoning in LLM Post-Training
(Sanderson, et al., 2026)
↩
-
Machine Unlearning Fails to Remove Data Poisoning Attacks
(Pawelczyk, et al., 2025)
↩
-
Language Models Struggle to Use Representations Learned In-Context
(Lepori et al., 2026)
↩
-
Code-Switching Red-Teaming: LLM Evaluation for Safety and
Multilingual Understanding
(Yoo et al., 2025)
↩