Overview

Methodological work on the structural properties that separate deployed sequential decision problems from standard benchmarks: delayed feedback, partial observability, action spaces too large to enumerate, fixed datasets that record only prior decisions, and state transitions that cannot be reversed.


Themes

  • Dead-ends and irreversibility — identifying states from which a poor outcome is already determined; relating irreversibility, loss of influence, and rare catastrophe; establishing where pessimism toward unobserved actions is justified
  • Credit assignment under delay — whether delay thresholds exist beyond which learning fails, and how delay biases the training signal and distributes credit across decisions
  • Uncertainty and partial observability — value uncertainty from a single network over sampled completions of an incomplete observation, rather than from an ensemble
  • Action spaces and structure — dependencies among sub-actions in combinatorial spaces, learning action validity separately from action value, and the cost of discretizing continuous control
  • Data and sample efficiency — whether offline RL method rankings hold as dataset size decreases below benchmark conventions
  • Transfer and shared experience — which representation of experience survives mismatch between agents or tasks: transitions, successor features, options, or learned structural priors
  • Environment design — automatic curriculum generation using learned models of environment validity in place of hand-specified mutation rules


Relevant prior work

Robust and Efficient Transfer Learning with Hidden Parameter Markov Decision Processes · NeurIPS 2017 (also AAAI 2017)

Killian, Daulton, Konidaris, Doshi-Velez. Models families of related tasks through latent parameters governing their dynamics, enabling policy transfer across instances that differ in dynamics but not in task.

Optimization Methods for Interpretable Differentiable Decision Trees Applied to Reinforcement Learning · AISTATS 2020

Silva, Killian, Rodriguez Jimenez, Son, Gombolay. Gradient-trainable decision trees as RL policies, retaining inspectable structure.

Direct Policy Transfer with Hidden Parameter Markov Decision Processes · Lifelong Learning Workshop, FAIM 2018

Yao, Killian, Konidaris, Doshi-Velez. Transfers policies directly across task instances given estimated latent parameters, without per-instance relearning.

Robust Autonomy Emerges from Self-Play · ICML 2025

Cusumano-Towner et al., incl. Killian. Self-play at large scale with randomization over agent physical and behavioral characteristics, producing robust driving policies in simulation.

BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces · NeurIPS 2025

Landers, Killian, Barnes, Hartvigsen, Doryab. Tree-structured traversal captures sub-action dependencies while evaluating a linear number of joint actions, improving on prior methods by up to 20×.

Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization · ICLR 2026

Landers, Killian, Hartvigsen, Doryab. SPIN pretrains an action structure model on valid action patterns, then trains lightweight control heads: up to 39% higher reward and up to 12.8× faster convergence.
Also relevant: SAINT, modeling joint actions as unordered sets via self-attention, and continuous-time evidential distributions for uncertainty over irregularly sampled observations.