Overview
Methodological work on the structural properties that separate deployed sequential decision problems from standard benchmarks: delayed feedback, partial observability, action spaces too large to enumerate, fixed datasets that record only prior decisions, and state transitions that cannot be reversed.
Themes
- Dead-ends and irreversibility — identifying states from which a poor outcome is already determined; relating irreversibility, loss of influence, and rare catastrophe; establishing where pessimism toward unobserved actions is justified
- Credit assignment under delay — whether delay thresholds exist beyond which learning fails, and how delay biases the training signal and distributes credit across decisions
- Uncertainty and partial observability — value uncertainty from a single network over sampled completions of an incomplete observation, rather than from an ensemble
- Action spaces and structure — dependencies among sub-actions in combinatorial spaces, learning action validity separately from action value, and the cost of discretizing continuous control
- Data and sample efficiency — whether offline RL method rankings hold as dataset size decreases below benchmark conventions
- Transfer and shared experience — which representation of experience survives mismatch between agents or tasks: transitions, successor features, options, or learned structural priors
- Environment design — automatic curriculum generation using learned models of environment validity in place of hand-specified mutation rules
Relevant prior work
Robust and Efficient Transfer Learning with Hidden Parameter Markov Decision Processes · NeurIPS 2017 (also AAAI 2017)
Optimization Methods for Interpretable Differentiable Decision Trees Applied to Reinforcement Learning · AISTATS 2020
Direct Policy Transfer with Hidden Parameter Markov Decision Processes · Lifelong Learning Workshop, FAIM 2018
Robust Autonomy Emerges from Self-Play · ICML 2025
BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces · NeurIPS 2025
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization · ICLR 2026