leiyo@dtu.dk
Technical University of Denmark (DTU)
My name is Lei You (由磊). I received my Ph.D. in Computer Science (specializing in Mathematical Optimization) from the Department of Information Technology at Uppsala University in 2019. During the PhD, I interned as a visiting data scientist at The Boston Consulting Group (BCG) Gamma. After the PhD, I worked as a data scientist in Bolt and Wolt (Doordash) in the domain of on-demand logistics optimization. I am a co-founder of https://cspaper.org.
My research develops verifiable AI systems: methods and information interfaces that make AI claims, decisions, and system changes independently checkable.
I study what evidence must be exposed, preserved, or tested when models, tools, memories, permissions, or environments change. My work draws on causal inference, information theory, and optimization to characterize what can actually be verified under limited evidence and limited verification resources.
(See a full list of publications here)
What remains verifiable when models or system components change?
I study AI systems as ecosystems of interacting components. The central question is not only whether one model performs well, but whether its behavior can be reproduced, repaired, or verified by the remaining system under explicit information constraints. This line includes matched in-silico quasi-experimental design, DISCO/PIER audits, and Minimum Viable Replacement: the information cost of replacing a blocked model.
L. You ✉, "Quantifying Model Uniqueness in Heterogeneous AI Ecosystems", preprint. [OpenPrint] [code]
L. You ✉, L. Cao, M. Nilsson, B. Zhao, and L. Lei, "Distributional Counterfactual Explanation With Optimal Transport", International Conference on Artificial Intelligence and Statistics (AISTATS) 2025 (Oral, top 2%). [arXiv] [code]
L. You ✉ and H. V. Cheng ✉, "SWAP: Sparse Entropic Wasserstein Regression for Robust Network Pruning", International Conference on Learning Representations (ICLR) 2024. [arXiv] [code]
What can be trusted when verification itself is scarce?
AI systems can now generate candidates, reviews, claims, code patches, and hypotheses at scale. The bottleneck is verification. This line studies how scarce human or automated checks turn untrusted candidate information into reliable knowledge, with applications to peer review, scientific assessment, RAG systems, and agentic workflows. I co-founded CSPaper (https://cspaper.org) as an open research infrastructure and a testbed for studying verification-limited AI systems.
L. You ✉, "Epistemic Throughput: Fundamental Limits of Attention-Constrained Inference", preprint. [OpenPrint] [code]
L. You ✉, L. Cao, and I. Gurevych, "Preventing the Collapse of Peer Review Requires Verification-First AI", preprint. [OpenPrint]
C. Yu, T. Shi, V. Uotila, S. Deng, L. You, and B. Zhao, "Vista: Verifier-in-the-Loop Agentic RL for Semantic Program Synthesis in Quantum Computing", ACM Conference on AI and Agentic Systems (CAIS) 2026. [link]
What evidence does an AI output actually provide?
This line develops counterfactual explanation methods that respect data geometry, distributional structure, and fairness constraints. It provides the methodological basis for my broader programme: counterfactual claims become scientifically useful only when the admissible transformation and its cost are made explicit.
L. You ✉, "DECAF: Decomposition of Evidence, Contradiction, And Fragility in Perturbation Responses", preprint. [arXiv] [code]
L. You ✉, Y. Bian, and L. Cao, "Joint Distribution–Informed Shapley Values for Sparse Counterfactual Explanations", International Conference on Learning Representations (ICLR) 2026 [arXiv] [code] [software].
L. Zhu, Y. Bian, and L. You ✉, "FairSHAP: Preprocessing for Fairness Through Attribution-Based Data Augmentation", International Conference on Artificial Intelligence and Statistics (AISTATS) 2026. [OpenPrint] [code]
Y. Gu, L. Cao, B. Zhao, L. Lei, and L. You ✉, "DISCOVER: A Solver for Distributional Counterfactual Explanations", European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD) 2026. [arXiv] [code]