“AI in Science: Building Infrastructure for an Explosion of Discovery”
Abstract: AI is rapidly accelerating math and science: just last month an internal OpenAI model solved the Navier-Stokes Millennium Prize problem, and Anthropic’s Claude drove the discovery of a promising enzyme system in a wet lab. Yet AI also allows anyone with an internet connection to generate ostensibly correct pseudoscience at the click of a button. Our scientific infrastructure is buckling under the increased volume.
I propose leveraging AI at key points to focus human oversight, and I demonstrate two applications of AI in peer review. Many peer review venues have deployed prototype AI reviewers to serve alongside human experts, claiming that AI reviewers can help identify errors and weed out obviously poor quality submissions. But do these AI reviewers accurately identify problems in submissions? I present a benchmark called FLAWS (Fault Localization Across Writing in Science), which inserts major flaws into papers to evaluate AI reviewer capabilities. At the time of evaluation (November 2025), we found that the best-performing model (GPT-5) identified inserted errors in <40% of instances. The journal Transactions on Machine Learning Research (TMLR) also used FLAWS to evaluate AI reviewers, finding the top-performing system’s performance satisfactory for experimental deployment alongside human reviewers.
Even with AI reviewers, peer review venues still struggle to keep up with increasing submissions of AI-generated articles – articles that the “authors” may not have even read! We develop a test called greCAPTCHA, which verifies author understanding by asking authors questions about their own papers. In a study with 31 researchers, we show that we are able to distinguish researchers answering questions about their own versus unfamiliar papers with an ROC AUC of 0.9. Semi-structured interviews show broad support for greCAPTCHA, and suggest changes to make before deployment. A public demo is available at grecaptcha.com. These projects examine how AI can support scientific integrity, while preserving human judgment and accountability.
Speaker Bio: Justin Payan is a Postdoctoral Research Associate in the Machine Learning Department at Carnegie Mellon University. He applies machine learning, combinatorial optimization, and the social sciences to improve scientific integrity and efficiency. He has published papers in IJCAI, AAMAS, NeurIPS, WSDM, and EMNLP. His reviewer assignment algorithm, FairSequence, has been used by over 50 venues on OpenReview to achieve fast and fair reviewer assignments. His work assessing author understanding was recently covered by Times Higher Education, and his collaboration with the ALMA observatory in Chile led to a novel telescope scheduling algorithm currently being evaluated for deployment.
The 2026-2027 seminars will be held in person, and are free and open to the public.
