NYU AI Safety Association

Related Professors at NYU

A collection of related professors. Inclusion does not imply an affiliation with the association.

Substantial safety focus Occasional safety work Safety-adjacent work
  1. Understanding Reasoning from Pretraining to Post-Training (Shen et al., 2026)
  2. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision (Burns et al., 2023)
  3. SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat (Jiang et al., 2025)
  4. From Distributional to Overton Pluralism: Investigating Large Language Model Alignment (Lake et al., 2025)
  5. Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation (Tripathi et al., 2025)
  6. Understanding Synthetic Context Extension via Retrieval Heads (Zhao et al., 2025)
  7. Does Weak-to-strong Generalization Happen under Spurious Correlations? (Liu, et al., 2026)
  8. Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies (Liu, et al., 2026)
  9. Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection (Phan, et al., 2025)
  10. Language Models Learn To Mislead Humans via RLHF (Wen et al., 2024)
  11. Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats (Wen et al., 2024)
  12. Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors (Yueh-Han et al., 2025)
  13. Resource Rational Contractualism Should Guide AI Alignment (Levine et al., 2026)
  14. Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas (Chiu et al., 2025)
  15. Imagining and building wise machines: The centrality of AI metacognition (Johnson et al., 2026)
  16. The Design and Composition of Structural Causal Decision Processes (Benthall & Lujan, 2026)
  17. AI Wardens (Diamantis & Benthall, 2026)
  18. Abstraction (Substack)
  19. Why Alignment Must Precede Distillation: A Minimal Working Explanation (Cha & Cho, 2025)
  20. Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check (Cho et al., 2025)
  21. The geometry of prompting: Unveiling distinct mechanisms of task adaptation in language models (Kirsanov et al., 2025)
  22. Emergence of Linear Truth Encodings in Language Models (Bai et al., 2025)
  23. Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models (Jhaveri et al., 2026)
  24. User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal (Liu et al., 2025)
  25. How DNNs break the Curse of Dimensionality (Jacot et al., 2025)
  26. Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature Learning (Wen, Jacot, 2026)
  27. There Will Be a Scientific Theory of Deep Learning (Simon et al., 2026)
  28. When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers (Lu et al., 2025)
  29. Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting (Dai et al., 2026)
  30. PILAF: Optimal Human Preference Sampling for Reward Modeling (Feng et al., 2025)
  31. From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers (Su et al., 2026)
  32. Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation (Bhardwaj, et al., 2026)
  33. Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting (Kulshreshtha, et al., 2026)
  34. When Are Concepts Erased From Diffusion Models? (Lu, et al., 2025)
  35. MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs (Chen, et al., 2026)
  36. Detecting All-to-One Backdoor Attacks in Black-Box DNNs via Differential Robustness to Noise (Fu, et al., 2025)
  37. Surgical Repair of Insecure Code Generation in LLMs (Sandoval, et al., 2026)
  38. Out-of-Distribution Detection Methods Answer the Wrong Questions (Li, et al., 2025)
  39. Automatic Layer Selection for Hallucination Detection (Wang, et al., 2026)
  40. Estimating Tail Risks in Language Model Output Distributions (Angell, et al., 2026)
  41. What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs (Nguyen, et al., 2026)
  42. HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking (Wang, et al., 2026)
  43. SALAD: Systematic Assessment of Machine Unlearning on LLM-Aided Hardware Design (Wang, et al., 2025)
  44. Sequential Data Poisoning in LLM Post-Training (Sanderson, et al., 2026)
  45. Machine Unlearning Fails to Remove Data Poisoning Attacks (Pawelczyk, et al., 2025)
  46. Language Models Struggle to Use Representations Learned In-Context (Lepori et al., 2026)
  47. Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding (Yoo et al., 2025)