Han Liu · Orrington Lunt Professor · Director, Center for Foundation Models & Generative AI · Northwestern University

Research Overview

Han Liu / Research program

Building Scientific Intelligence

We develop the foundations and systems for AI that can reason about scientific problems, model the world, and learn through experimentation.

Our central thesis: scientific experience should improve both what an AI knows and how it discovers. We connect reasoning agents, scientific world models, and autonomous laboratories to make that ambition concrete.

10M+
DNABERT downloadsAs of September 2026
A research architecture for scientific intelligence Reasoning agents propose hypotheses and plans. World models support prediction and intervention design. Autonomous laboratories produce measurements and evidence. Experience feeds back to improve both reasoning and models. ONE CONTINUOUS CYCLE OF DISCOVERY 01 / REASON02 / MODEL03 / EXPERIMENT Scientific agentsWorld modelsAutonomous labs Hypotheses · plansMemory · tool useStructure · dynamicsPrediction · interventionsSimulation · executionMeasurements · evidence NEW EVIDENCE → REUSABLE EXPERIENCE → BETTER DECISIONS
Reasoning chooses what to investigate. Models anticipate possible outcomes. Experiments produce evidence. Learning improves the next cycle.

Research agenda

Three frontiers. One scientific ambition.
Recursive Scientific Intelligence

How can scientific experience become better reasoning?

We study the computational foundations of foundation models: memory, reasoning, adaptation, in-context computation, and tool use. We build scientific agents that organize evidence, formulate hypotheses, and plan with computational and experimental tools. The goal is to turn experience into better strategies and decisions in subsequent research cycles.

Scientific World Models

What must a model understand to guide scientific discovery?

We develop models of scientific objects, processes, and dynamics, from genomes, proteins, and cells to molecular systems and the changing universe. We investigate how these representations can support prediction, uncertainty quantification, and reasoning about interventions. Biology and astronomy provide complementary settings for studying these questions.

Autonomous Laboratories

How can each experiment make the next one more informative?

We connect scientific agents with molecular design, physics-based simulation, robotics, and laboratory instruments. Our focus is the full Design–Build–Test–Learn cycle: choosing experiments, executing them, preserving the context of measurements, and learning from successes and failures to guide the next design.

Selected contributions

Models, decisions, and executable systems.

Representative work with our students and collaborators, spanning scientific representations, protein design policies, research agents, and embodied execution.

Genomic generationModels released
GenomeOcean

Genome foundation models trained on large metagenomic assemblies, with released models up to four billion parameters. The work connects genomic representation learning with sequence generation, supported by public code and model checkpoints.

Cellular representations2026 · Preprint
Cell-JEPA

A joint-embedding predictive approach to single-cell transcriptomics. Predicting latent cellular representations from partial gene-expression observations helps the model learn features that are robust to missing measurements.

Protein evolutionActive research
PEV

Trained on measured protein outcomes, PEV learns a policy for choosing single-residue substitutions or STOP. It scores candidate edits in parallel from one parent encoding, exploring how experimental experience can inform new protein-design decisions.

Research agentsPublished research
SciSciGPT

A prototype AI collaborator for the science of science, developed with Dashun Wang and collaborators. It brings language models into analytical workflows, supporting research iteration and reproducibility.

Embodied execution2026 · Preprint
MagicSim

A unified simulation runtime for embodied agents, connecting task specifications, robot control, planning, evaluation, and trajectory collection. High-level commands execute through robot actions, producing structured multimodal records for learning and evaluation.

Flagship programs

Connecting models, people, and experiments.

Liu helps lead complementary efforts that connect this research agenda to experimental infrastructure and scientific communities.

DREAM
AI-powered protein engineering
$20MNSF initiative

Co-PI · AI lead

DREAM is being developed as a national cloud laboratory for protein engineering. The program connects AI-driven design, cell-free protein synthesis, automated measurements, and iterative learning. Our work on its AI systems and data infrastructure aims to make experimental evidence a foundation for better models and scientific decisions.

Explore the DREAM Cloud Lab ↗
CFMG
Center for Foundation Models
and Generative AI

Director · Northwestern University

The center brings together researchers across disciplines to advance foundation models and generative AI, connecting fundamental research with problems in science and engineering.

Explore CFMG ↗
TEMPI
A world model for
the time-domain sky

Co-lead · NSF–Simons SkAI Institute

TEMPI extends the scientific world-model agenda to astronomy: learning representations of physical dynamics from time-series observations of the changing universe.

Explore TEMPI ↗

Intellectual foundations

Grounded in statistical machine learning.

Our current program grows out of contributions to nonparametric graphical models, nonconvex optimization, and high-dimensional inference. These ideas remain central: learning structure, quantifying uncertainty, and making reliable decisions from incomplete evidence.

Selected recognition
PECASE · Alfred P. Sloan Fellowship · IMS Tweedie New Researcher Award
ASA Noether Young Scholar Award · NSF CAREER Award

Earlier publications →

Flickr gallery

Reading Group

Our weekly reading group explores ideas that may shape the next generation of machine and scientific intelligence — from the foundations of reasoning, memory, and world models to agents that interact with tools, simulations, and the physical world. We use the group not simply to review papers, but to identify emerging paradigms, formulate new research questions for scientific discovery.

Get In Touch

Department of Computer Science
Mudd Hall 3119
Northwestern University
Evanston, IL 60201
Phone: +1 847 491 2793
Email: hanliu@northwestern.edu