light_mode

LLM-Assisted Red Teaming of Diffusion Models through “Failures Are Fated, But Can Be Faded”

Som Sagar, Aditya Taparia, Ransalu Senanayake
NeurIPS Workshop on Red Teaming GenAI: What Can We Learn from Adversaries?, 2024

Red Teaming Diffusion Models teaser figure

In one sentence. An extension of the Failures Are Fated framework to text-to-image diffusion models with LLM-generated rewards and states, action screening inspired by design of experiments, and a comparison of DQN, PPO, and A2C for red teaming.

picture_as_pdf PDF

Summary

This work extends the Failures Are Fated framework to red teaming of text-to-image diffusion models. We introduce LLM-generated rewards and states, action screening inspired by design of experiments to reduce the search space of candidate failure factors, and a comparison of DQN, PPO, and A2C as the underlying reinforcement learning algorithms for discovering failure modes in generative models.

Note. Builds on Failures Are Fated, But Can Be Faded (ICML 2024 spotlight).

BibTeX

@inproceedings{sagar2024llmredteaming,
  title     = {LLM-Assisted Red Teaming of Diffusion Models through ``Failures Are Fated, But Can Be Faded''},
  author    = {Sagar, Som and Taparia, Aditya and Senanayake, Ransalu},
  booktitle = {NeurIPS Workshop on Red Teaming GenAI: What Can We Learn from Adversaries?},
  year      = {2024}
}

Topics

red teaming, diffusion models, text-to-image, generative model safety, reinforcement learning, LLM-assisted evaluation

About the author. Som Sagar is a computer science PhD student at Arizona State University, advised by Ransalu Senanayake in the LENS Lab. He works on reinforcement learning for post-training and agentic LLM systems, and on failure diagnosis, red teaming, and safety evaluation of foundation models, vision-language models, and robot manipulation policies.

home Homepage library_books All publications description CV school Google Scholar