Robust Adversarial Reinforcement Learning for Multi-Variable Burn Control of ITER Plasmas

B. Liu, V. Graber, S. T. Paruchuri, E. Schuster

34th Symposium on Fusion Technology (SOFT)

Aix-en-Provence, France, September 21-25, 2026

Abstract

Burn control of ITER plasmas requires the simultaneous regulation of strongly coupled kinetic properties governed by highly nonlinear and complex dynamics. This work presents a Robust Adversarial Reinforcement Learning (RARL) framework for multi-variable burn control in a nonlinear, zero-dimensional ITER environment. Because ITER is not yet operational, the physics-based models required for controller design lack experimental validation, introducing significant parametric uncertainty. To bridge this sim-to-real gap, an adversarial agent is introduced during the training phase to impose worst-case, physically meaningful perturbations on the uncertain model parameters. The adversary specifically targets parameters governing plasma confinement—which dictates heat and particle transport losses—and deuterium-tritium recycling driven by plasma-wall interactions. Formulated as a min-max game, the training process forces the control policy to learn to maintain performance under adverse operating conditions. The resulting controller regulates the plasma’s ion energy, electron energy, total particle density, and tritium fraction using external deuterium and tritium fueling, alongside auxiliary ion and electron heating. To assess performance, a simulation study is conducted over a set of ITER-relevant scenarios, including nominal operation, enhanced and degraded confinement, fuel-rich and fuel-poor extremes, and a composite worst-case scenario featuring poor confinement, unfavorable fueling, and elevated impurities. In all cases, the RARL algorithm successfully drives the plasma states to their reference targets while preserving stable closed-loop behavior and strictly respecting actuator saturation limits. The controller also exhibits physically intuitive behavior, such as reducing auxiliary heating during improved confinement and increasing fueling under reduced wall-recycling conditions. Across all tested scenarios, including the composite worst-case, relative steady-state errors for all four controlled variables remain below 1%. These results demonstrate that adversarially trained reinforcement learning is a highly promising approach for robust ITER burn control, effectively mitigating the challenges imposed by substantial parametric uncertainty.

*Supported by the US DOE under DE-SC0010661.