Seminar, Jiwei Zhao, SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions
Refreshments: 10:30 AM
Seminar: 11:00 AM
Title: SADA: Safe and Adaptive Aggregation of Multiple Black-Box Predictions
Abstract: Real-world applications often face scarce labeled data due to the high cost and time requirements of gold-standard experiments, whereas unlabeled data are typically abundant. With the growing adoption of machine learning techniques, it has become increasingly feasible to generate multiple predicted labels using a variety of models and algorithms, including deep learning, large language models, and generative AI. In this talk, I will present a novel approach that safely and adaptively aggregates multiple black-box predictions of uncertain quality for both inference and prediction tasks. Our method provides two key guarantees: (i) it never performs worse than using the labeled data alone, regardless of the quality of the predictions; and (ii) if any one of the predictions (without knowing which one) perfectly fits the ground truth, the algorithm adaptively exploits this to achieve either a faster convergence rate or the semiparametric efficiency bound. We demonstrate the effectiveness of the proposed algorithm through simulations and two real-data analyses with distinct scientific goals.
Research Interests: My overarching research is to use statistical methods and machine learning techniques to analyze data with massive structures in various biomedical studies. He has broad interests in clinical trial, missing data analysis, causal inference, semiparametric, robustness, high-dimensional statistical inference, semi-supervised learning, transfer learning, patient-reported outcomes, electronic health records, health services research, health disparity/equity.