Testing AI’s Promise for Life Sciences

Testing AI’s Promise for Life Sciences

Liang Zhao lab members in tree-filled park
Liang Zhao, PhD, MAS, MBA, and his team at the UCSF School of Pharmacy

Artificial intelligence can synthesize enormous amounts of scientific knowledge in minutes. But in biomedical research, a fast answer is not necessarily a reliable one. 

That’s the problem Liang Zhao, PhD, MAS, MBA, and his team at the UCSF School of Pharmacy are helping to address. Working in close collaboration with OpenAI colleagues, the researchers are developing a benchmark to evaluate how well AI tools can perform scientific tasks. Included among the tools is GPT-Rosalind, a domain-specific large language model designed for life sciences such as genomics and drug discovery. 

Liang Zhao, PhD, MAS, MBA
Liang Zhao, PhD, MAS, MBA

The UCSF team isn’t developing GPT-Rosalind. Instead, Zhao, a professor in the school’s Department of Bioengineering and Therapeutic Sciences, and his colleagues are building a way to test it — and other AI models — on questions involving clinical data. 

“We’re actually building a benchmark to evaluate the performance of different AI tools, including Rosalind,” Zhao said. 

The work puts the School of Pharmacy in a role that is increasingly important as AI moves deeper into biomedical research: helping to determine not only what these tools can do, but whether researchers can trust their results. 

Putting AI to the test 

Zhao’s team is interested in what happens further along the research pipeline, when a potential therapy has moved into clinical testing and researchers need to work with data from humans. 

“We’re focused on building models to deal with clinical data, not the drug discovery data,” Zhao said. “Rosalind is more for drug discovery at this stage.” 

The team is developing a battery of questions focused on clinical data and mathematical modeling. Those questions can then be posed to different AI systems, and their performance can be evaluated. 

The goal is to find out whether an AI tool can do more than produce an answer that sounds plausible. For example, Zhao’s team is asking whether an AI system can build a model to process clinical data when given the appropriate prompt — and, if it can, how reliable that model is. 

That distinction between capability and reliability is central to the project. In scientific research, an incorrect answer can be more dangerous when it appears convincing. 

Where humans remain essential 

Liang Zhao in lab
Zhao in the Zhao Lab.

Zhao sees AI as a powerful accelerator for research. Tasks that once required researchers to spend months gathering and synthesizing information can increasingly be accomplished in minutes. A literature review, for example, can be rapidly assembled with the help of AI. 

“AI tools are definitely a catalyst accelerator for research,” Zhao said. 

But the same technology that makes research faster could also change how scientists learn to think. Zhao worries that researchers, particularly the next generation of graduate students, could become so accustomed to relying on AI that they lose some of their ability to think independently. 

“Eventually, we may lose the capability to even judge whether an AI-generated result is correct,” he said. 

For Zhao, that is one of the fundamental challenges of AI in science. AI can summarize existing knowledge and synthesize information from enormous bodies of data. But scientific breakthroughs don’t always come from what is already known. They can depend on observation, creativity, unexpected connections, or simply noticing something that nobody thought to look for. 

That makes human judgment an essential part of the equation. 

“AI can accelerate the work, but humans still need to validate its results," said Zhao. “The role of humans is to critically evaluate what AI produces.” 

Why UCSF matters 

Zhao sees UCSF as a natural partner in this emerging field because of its depth in health sciences, including medicine and pharmacy. 

The school brings expertise in biomedical research, therapeutics, and clinical science to a field that is developing rapidly — and where the ability to distinguish a useful AI result from a misleading one will become increasingly important. 

The long-term possibilities are difficult to predict. Zhao imagines a future in which a researcher could describe a disease to an AI system and receive a molecule to investigate, followed by increasingly sophisticated predictions about whether the molecule can bind to a target, how it might behave in the human body, and eventually its efficacy and safety. 

That vision is still far away. “We aren’t there yet,” Zhao said.