Overview

Applied sequential decision making on observational health data, where the record is confounded by prior treatment decisions, observations are irregularly sampled, outcomes are delayed, populations differ across sites, and some clinical trajectories cannot be reversed.


Themes

  • Dead-ends and irreversibility in clinical data — identifying states and treatments to avoid from negative outcomes, and whether the approach holds in pediatric cohorts with faster physiology, weight-based dosing, and smaller sample sizes
  • Equity and fairness — distinguishing treatment differences attributable to physiology, clinician bias, or resource constraint, and preventing learned policies from reproducing them
  • Transfer across populations and sites — grounding policy transfer in counterfactual reasoning rather than distributional similarity, and sharing learned structure across sites without sharing patient data
  • Uncertainty over irregular observations — calibrated estimates that widen as time since last measurement increases
  • Sequential allocation beyond the clinic — formulating humanitarian aid allocation as a sequential decision problem with delayed outcomes, subject to data availability


Relevant prior work

Medical Dead-ends and Learning to Identify High-Risk States and Treatments · NeurIPS 2021

Fatemi, Killian, Subramanian, Ghassemi. Uses negative outcomes in data-constrained offline settings to identify behaviors to avoid, guarding against overoptimistic decisions. Introduces separate value estimation over failure and success signals.

An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare · ML4H, NeurIPS 2020

Killian, Zhang, Subramanian, Fatemi, Ghassemi. Systematic comparison of patient state representations for offline RL, quantifying their effect on the resulting policies.

Counterfactually Guided Policy Transfer in Clinical Settings · CHIL 2022

Killian, Ghassemi, Joshi. Combines causal inference with offline RL to support policy transfer between patient populations via counterfactual estimation.

Multiple Sclerosis Severity Classification From Clinical Text · Clinical NLP Workshop, 2020

D'Costa, Denkovski, Malyska, Moon, Rufino, Yang, Killian, Ghassemi. Severity estimation from clinical notes, with a released domain-specific language model.

Risk Sensitive Dead-end Identification in Safety-Critical Offline Reinforcement Learning · TMLR 2023

Killian, Parbhoo, Ghassemi. Applies distributional RL to dead-end discovery, providing earlier identification with risk tolerance as a tunable parameter.

Clinically Motivated Sequential Decision Making Under Uncertainty in Offline Settings · PhD Thesis, University of Toronto, 2024

Killian. Modeling decisions for deriving actionable insight from sequentially observed healthcare data.
Also relevant: Identifying Disparities in Sepsis Treatment using Inverse Reinforcement Learning (Jeong, Killian, Kanjilal, Nayak, Ghassemi; NeurIPS workshops 2022) and Continuous Time Evidential Distributions for Irregular Time Series.