15:00 (CET)
Abstract:
We consider the problem of learning risk scores to prioritize individuals for scarce resources or interventions, from historical observational data affected by unobserved confounding. In settings such as public health and homelessness prevention, decisions about who receives a scarce resource (e.g., a hospital bed or housing) are often guided by a risk score assigned to each individual based on recorded characteristics, such as responses to a survey. These risk scores are increasingly being learned directly from observational data: historical records of individuals' characteristics, allocation decisions, and outcomes under received allocations. Standard methods such as inverse propensity weighting (IPW), which corrects for the bias introduced by the historical allocation policy, can be used to learn accurate risk scores if the historical decision process is fully explained by the recorded characteristics, an assumption known as unconfoundedness. In practice, however, historical decisions often depend on unrecorded information (e.g., details a caseworker learns in conversation but never logs), causing learned risk scores to systematically under-prioritize exactly the individuals whose unrecorded circumstances drove past prioritization. We propose a method for learning risk scores that are robust to this kind of unobserved confounding, building on IPW. Since propensity weights cannot be reliably estimated under unobserved confounding, we instead treat them as belonging to an uncertainty set determined by the observable data and domain-informed estimates of the degree of confounding, combining sensitivity analysis from causal inference with Wasserstein distributionally robust optimization. The resulting robust risk score learning problem admits a sample-based approximation that we reformulate as an exponential conic program compatible with off-the-shelf solvers. We demonstrate the effectiveness of our approach relative to IPW and other benchmarks on semi-synthetic data derived from datasets in the UCI Machine Learning Repository. Beyond risk scores, this way of representing unobserved confounding as an uncertainty set is not specific to our setting and can be used to robustly learn other causal quantities of interest.
Bio:
Çağıl Koçyiğit is an associate professor at the Department of Engineering and the Department of Economics and Management, University of Luxembourg. Her research focuses on optimization under uncertainty with applications in policy and mechanism design. Her work has been published in prestigious journals including Management Science, Operations Research, and Mathematical Programming, and she has received multiple awards and recognition for her work. She holds a Ph.D. in Management of Technology from École Polytechnique Fédérale de Lausanne (EPFL).
15:00 (CET)
University of Pennsylvania
Alignment of Generative AI Models: Jailbreaking Attacks and Defenses
Abstract:
In recent years, LLM-based agents have been used to solve a multitude of natural language tasks; yet, despite efforts to align them with human intent, popular LLMs remain susceptible to jailbreaking attacks that elicit unsafe content and actions. Early jailbreaks targeted the generation of harmful information, whereas modern attacks seek domain-specific harms (e.g., digital agents violating user privacy and security, or LLM-controlled robots performing harmful actions in the physical world). In the worst case, future attacks may target self-replication or power-seeking behaviors. Therefore, it is critical to study these failure modes and develop effective defense strategies.
A key component of AI safety is model alignment--a broad concept referring to algorithms that optimize LLM outputs to align with human values to improve safety and security. In this talk, I will focus on safety vulnerabilities of frontier LLMs and review the current state of the jailbreaking literature, including robust generalization, open-box and black-box attacks, defenses, and evaluation benchmarks. I will also discuss methodologies that aim to mitigate these vulnerabilities and align language models with human standards.
The talk will require no prior background in AI safety or jailbreaking and will be self-contained.
Bio:
Hamed Hassani is currently a senior research scientist at Google as well as an associate professor of Electrical and Systems Engineering department, the Computer and Information Systems department, and the department of Statistics and Data Science at the University of Pennsylvania. Prior to that, he was a research fellow at Simons Institute for the Theory of Computing (UC Berkeley) affiliated with the program of Foundations of Machine Learning, and a post-doctoral researcher in the Institute of Machine Learning at ETH Zurich. He received a Ph.D. degree in Computer and Communication Sciences from EPFL, Lausanne. He is the recipient of the 2014 IEEE Information Theory Society Thomas M. Cover Dissertation Award, 2015 IEEE International Symposium on Information Theory Paper Award, 2019 National Science Foundation (NSF) CAREER Award, 2020 Air Force Office of Scientific Research (AFOSR) Young Investigator Award, 2020 Intel Rising Star award, the distinguished lecturer of the IEEE Information Society in 2022-23, and the 2023 IEEE Communications Society & Information theory Society Joint Paper Award. Moreover, is the recipient of 2023 IEEE Information Theory Society’s James L. Massey Research and Teaching Award for Young Scholars.
15:00 (CET)
15:00 (CET)
15:00 (CET)
15:00 (CET)
15:00 (CET)
Winter Break