.

Round-Robin Catalysis

Monday, August 3, 2026

To turn abundant carbon dioxide into valuable fuel, we need a fast and efficient way of determining which catalysts work best over the longest time. AI models have the potential to help guide catalyst selection, but as with internet chatbots, AI models are only as good as the data you put into them.

By convening four laboratories from across the nation, including the research group of UC Santa Barbara chemical engineering professor Phillip Christopher, to test an experimental carbon monoxide-producing catalyst, a key first step in turning carbon dioxide into fuels, researchers at SLAC National Accelerator Laboratory have demonstrated the importance of generating highly reproducible experimental data when building AI models for investigations in science. They published the results in Nature Catalysis

Operated at Stanford University for the U.S. Department of Energy’s Office of Science, the SLAC National Accelerator Laboratory explores how the universe works at the biggest, smallest and fastest scales and invents powerful tools used by researchers around the globe, scientific computing and the development of next-generation accelerators.

“Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how the consistency of experimental data can influence the reliability of the outcomes,” said Selin Bac, a postdoctoral researcher in the Christopher lab at UCSB, and first author on the study.

Speeding up catalyst development – with AI

With a good AI model, researchers can enter conditions such as temperature, length of time of the reaction, and catalyst formulation, then run the simulation and see a prediction of how well the catalyst performs. They can then confirm the predictions with a few well-designed experiments, ultimately speeding up catalyst discovery and implementation at a global scale.

In addition to saving time and money, such models can also explore conditions that are difficult to achieve in the lab. Most lab catalysis studies can only look at short time-periods (days), but catalyst deactivation occurs over time (months to years) due to buildup of impurities and repeated exposure to high temperatures.

AI models need large amounts of high-quality data for training. To generate the data, the four labs performed a set of round-robin experiments, in which multiple laboratories conduct the same tests to evaluate reproducibility using previously agreed upon protocols and the same rhodium-based catalyst. 

Squaring the data from round-robin experiments

To the researchers’ surprise, achieving the same results from four labs working independently was harder than anticipated. When they got together to share their results, they realized that they had a problem. 

Each of the four research teams produced results that varied in the amounts of carbon monoxide and methane, an undesirable side product, produced. The computer would not be able to learn from four sets of data that contain different outcomes.

“It was a bit of an eye-opener,” said Adam Hoffman, a senior author of the study and a staff scientist at SLAC, which operates at Stanford University for the U.S. Department of Energy’s Office of Science. “This experience shines light on the practical challenges of including real-world data into machine learning models.”

Painstakingly the teams evaluated their methods. Through rigorous testing they found a handful of sources of mismatch, with one of the biggest contributors to the variability coming down to how hard the mixture was shaken or stirred. 

With further standardization across the four labs – which in addition to SLAC included groups at Pennsylvania State University, Stanford University, and UCSB – the results began to look more consistent. The team outlined several recommendations to strengthen experimental reproducibility, including enhancing the consistency of reactor design, operating protocols, and experimental conditions. 

Hoffman said he hopes the study will help experimentalists and data scientists who are designing AI models to consider how small variations in experimental design across labs can lead to problems with reproducibility and impact on long-term predictions for AI modes.

“We see this work as a guide for the community as to how to think about designing experiments for inclusion in machine learning models,” Hoffman said.

This work was supported in part by the U.S. Department of Energy (DOE) Office of Science. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. The SLAC portion of the research took place at the Stanford Synchrotron Radiation Lightsource (SSRL), a DOE Office of Science user facility.

Illustration showing carbon dioxide and methane molecules moving past a catalyst particle above a graph of reaction rates over time. Colored lines represent differing experimental results, which feed into a network symbolizing a machine-learning model.

Related People: 
Phillip Christopher
Illustration showing carbon dioxide and methane molecules moving past a catalyst particle above a graph of reaction rates over time. Colored lines represent differing experimental results, which feed into a network symbolizing a machine-learning model.

A study conducted at four laboratories across the country demonstrated how experimental protocol and equipment standardization govern result variability. Recognizing this variability is essential when incorporating real-world data into AI machine learning models. (Adam Hoffman and Greg Stewart/SLAC National Accelerator Laboratory)