Research

Health systems record an extensive amount of data on what happens to patients. A substantial gap remains in harnessing those data to inform better decision making. Our work is directed at closing that gap: turning routinely collected clinical data into evidence rigorous enough to support medical, public health, and regulatory decisions.

  1. 01

    Where randomized evidence is unavailable, we specify the trial that would have answered the question and emulate its protocol in observational data.

  2. 02

    Evidence is of limited value if it cannot be located within the time a decision allows.

  3. 03

    We develop models that learn from clinical data directly rather than from general-purpose web text.

Back to home

01

Target trial emulation

Where randomized evidence is unavailable, we specify the trial that would have answered the question and emulate its protocol in observational data.

Many of the questions that clinicians and regulators face have not been settled by a randomized trial, and for a substantial proportion of them a trial would be prohibitively slow, prohibitively expensive, or unethical. Target trial emulation offers a principled means of addressing such questions. The investigator specifies the protocol of the trial that would have been conducted, including eligibility criteria, treatment strategies, outcomes, and follow-up, and then emulates that protocol in observational data. The resulting analysis inherits the logic of a randomized trial rather than the structure of the database from which it is drawn.

We develop scalable, distributed analytics frameworks that carry out these emulations across institutions, and we apply them to evaluate the effectiveness and safety of chemical, biological, and digital therapeutics. This work draws on large-scale electronic health record data from UCSF and the wider UC Health system, and it supports post-market surveillance, real-world evidence generation, and regulatory science.

02

Evidence retrieval at scale

Evidence is of limited value if it cannot be located within the time a decision allows.

The information required to answer a clinical or regulatory question is frequently present in the medical record already, distributed across millions of free-text notes that cannot feasibly be reviewed by hand. diveEHR is a large-scale clinical text retrieval platform developed to extract that evidence efficiently, at the scale these questions require.

Applied alongside our modelling work, the platform addresses a practical constraint on evidence generation: the effort required to assemble documentation for regulatory submissions, pre-authorization workflows, and CMS quality reporting. Throughout, we maintain the standards of causal validity, privacy, and reproducibility that govern the remainder of our work.

03

Clinical foundation models

We develop models that learn from clinical data directly rather than from general-purpose web text.

General-purpose language models are trained predominantly on text drawn from the open internet. Clinical care generates a distinct form of record, comprising notes, laboratory results, orders, images, and trajectories that unfold over years, with its own structure and its own characteristic failure modes. We are developing UCSF-GPT, a multimodal foundation model trained from scratch on clinical data, so that its learned representations reflect the manner in which care is documented and delivered.

The model is subject to the same requirements as the remainder of our work: causal validity, patient privacy, and reproducibility are treated as design constraints rather than as considerations addressed after the fact.