NYU AI Safety Association
Menu

Resources

A collection of AI Safety related organizations and people. Inclusion does not imply a relationship with NYU AI Safety. This list is not yet complete.

Related organizations at NYU

Related organizations in NYC

NYU Professors with related research

  1. SafeNudge: Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs (Fonseca et al., 2025)
  2. Why Alignment Must Precede Distillation: A Minimal Working Explanation (Cha & Cho, 2025)
  3. Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check (Cho et al., 2025)
  4. The geometry of prompting: Unveiling distinct mechanisms of task adaptation in language models (Kirsanov et al., 2025)
  5. Emergence of Linear Truth Encodings in Language Models (Bai et al., 2025)
  6. Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models (Jhaveri et al., 2026)
  7. User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal (Liu et al., 2025)
  8. SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat (Jiang et al., 2025)
  9. From Distributional to Overton Pluralism: Investigating Large Language Model Alignment (Lake et al., 2025)
  10. Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation (Tripathi et al., 2025)
  11. Understanding Synthetic Context Extension via Retrieval Heads (Zhao et al., 2025)