Conference on Empirical Methods in Natural Language Processing (EMNLP) GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
Conference on Uncertainty in Artificial Intelligence (UAI) How Learning Dynamics Drive Adversarially Robust Generalization?
Annual Meeting of the Association for Computational Linguistics (ACL) Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Transactions on Machine Learning Research Diffusion-based Cumulative Adversarial Purification for Vision Language Models
ACM Cyber-Physical System Security Workshop (CPSS) FEVA-ICS: Benchmarking Adversarial Robustness of Machine Learning-based Intrusion Detection Systems in Industrial Control Systems