Abstract
Modeling human decision-making in dynamic environments is fundamental for developing human-aware autonomous systems and cognitive science research. Such behavior typically reflects bounded, state-varying foresight. Yet standard inverse reinforcement learning (IRL) assumes untruncated credit assignment, absorbing this foresight into the inferred reward. Here we introduce SC-AIRL (State-Conditional AIRL), an adversarial IRL framework that learns bounded, state-conditional plan- ning depth from human behavior. By composing per-depth policies through a learned router under a shared reward, SC-AIRL disentangles how deeply an agent plans from what it values. On a naturalistic pedestrian dataset, SC-AIRL recovers participant-level action distributions and behavioral rates, including human-specific waiting events that single-horizon AIRL fails to capture. Inferred depth correlates with independently measured impulsivity and trait anxiety, indicating that inferred state-conditional foresight captures clinically relevant individual differences. By making state-conditional bounded foresight a structural primitive in imitation learn- ing, SC-AIRL enables human behavior modeling for autonomous vehicles and supports behavior-based phenotyping for cognitive and clinical research.