Hierarchical Reinforcement Learning and MPC Architecture for Real-time Mixed-Actuator Plasma Density Profile Control
B. Liu, V. Graber, S. T. Paruchuri, H. A. Khawaldeh, T. Rafiq, E. Schuster
68th Division of Plasma Physics (DPP) Annual Meeting of the American Physical Society (APS)
Chicago, IL, USA, November 2-6, 2026
Real-time plasma density profile control combining discrete pellet injection and continuous gas fueling gives rise to a mixed-integer model predictive control (MI-MPC) problem that is computationally prohibitive for online execution. To overcome this limitation, a hierarchical controller is developed using the physics-based Control Oriented Transport SIMulator (COTSIM). In this architecture, a high-level policy trained offline via proximal policy optimization (PPO) generates a 10-step binary pellet schedule, while a lower-level quadratic MPC optimizes continuous gas fueling in real time. The proposed approach is evaluated against an exact benchmark joint gas-pellet MPC obtained by enumerating all 1024 candidate pellet sequences. Both controllers enforce identical gas constraints and are tested on 100 fixed validation target profiles. Using the integrated normalized profile error as the tracking metric, the hierarchical PPO+MPC strategy achieves a mean error of 11.62%, closely approaching the 11.21% error of the exact joint MPC. Crucially, the median decision time is reduced from 163.6 ms to 1.00 ms, corresponding to a 164-fold speedup, while the worst-case decision time decreases from 179.4 ms to 4.90 ms. These results demonstrate that offline-learned discrete planning can preserve near-optimal MPC tracking performance while enabling real-time hierarchical plasma fueling control.
*Supported by the U.S. DOE under Award DE-SC0010661.