Optimizing heat FLUX deposition of diverted plasma configurations using controller neural networks
Controller neural networks trained via reinforcement learning address the inefficiencies of conventional magnetic controllers by autonomously managing plasma configurations for efficient heat and particle exhaust, reducing design effort and optimizing plasma-facing component temperatures.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DEEPMIND TECH LTD
- Filing Date
- 2025-11-20
- Publication Date
- 2026-05-28
AI Technical Summary
Conventional magnetic controllers for magnetic confinement devices face challenges in efficiently managing heat and particle flux from high-temperature plasmas, requiring significant engineering effort and expertise, and are not optimized for maintaining plasma-facing components below safety thresholds.
Implementing controller neural networks trained via reinforcement learning to manage plasma configurations, allowing for simultaneous heat and particle exhaust onto divertors while maintaining plasma-facing component temperatures below thresholds.
The controller neural networks autonomously learn a near-optimal control policy, reducing design effort and replacing complex nested control architectures, providing flexible and general control for magnetic confinement devices.
Smart Images

Figure US2025056390_28052026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 45288-0528WO1OPTIMIZING HEAT FLUX DEPOSITION OF DIVERTED PLASMA CONFIGURATIONS USING CONTROLLER NEURAL NETWORKSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This specification claims priority to U. S. Provisional Application No. 63 / 724,001, filed on November 22, 2024. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.BACKGROUND
[0002] This specification relates to processing data using machine learning models.
[0003] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.
[0004] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a nonlinear transformation to a received input to generate an output.SUMMARY
[0005] This specification describes a control system implemented as computer programs on one or more computers in one or more locations that can generate an adaptive control sequence for a magnetic confinement device using a controller neural network, e g., trained via reinforcement learning on a model of the magnetic confinement device. Examples of reinforcement learning for tokamak control are provided by Degrave, J., Felici, F., Buchli, J. et al. “Magnetic control of tokamak plasmas through deep reinforcement learning.” Nature 602, 414-419 (2022), and Tracey, Brendan D., et al. “Towards practical reinforcement learning for tokamak magnetic control.” Fusion Engineering and Design 200 (2024): 114161, both of which are incorporated by reference herein in their entireties for all purposes.
[0006] At a high-level, the magnetic confinement device includes a chamber (aka a vessel) configured with one or more divertors for removing excess heat from the chamber. Such situations are (or are expected to be) prevalent in state-of-the-art fusion reactors that have a net energy output, i.e., a Q value greater than 1, where heat flux from a plasma is sufficient to melt the plasma-facing components (“PFCs”) in direct contact with the plasma, such as the first wall, divertor, and blanketAttorney Docket No. 45288-0528WO1modules. The magnetic confinement device further includes a set of control coils arranged within and / or about the chamber, and multiple sensors positioned within and / or about the chamber.
[0007] The chamber is configured to support a vacuum for a plasma generated therein, the set of control coils are configured to generate a magnetic field within the chamber to manipulate a configuration of the plasma, and the sensors are configured to collect measurements of a current state of the chamber. The measurements characterize a current configuration of the plasma within the chamber and a current temperature distribution over the PFCs of the chamber, e.g., including the first wall, the divertor(s), and the blanket. The sensors can include magnetic field probes, magnetic flux loops, and electric current sensors for measuring the strength and flux of the magnetic field under a particular electric current in each control coil, thereby characterizing the current configuration of the plasma in terms of its influence on the magnetic field relative to vacuum. The sensors can also include temperature sensors for (directly) measuring the current temperature distribution over the PFCs resulting from heat flux of the hot plasma upon these components.
[0008] In general, the control system is configured to establish a target, diverted configuration of the plasma within the chamber to optimize the time-dependent heat flux of the plasma, such that the temperature distribution over the first wall is kept below maximum allowable thresholds. For example, the control system can be implemented to confine plasmas undergoing sustained thermonuclear fusion, which correspond to plasma temperatures, e.g., electron and / or ion temperatures, in range of about 10 to 100 million degrees Celsius (°C) or higher. Typically, the control system establishes a sequence of diverted configurations over a control loop that strongly confine the hot plasma within a particular region of the chamber at each time step in the control loop, thus allowing any excess heat and particle exhaust to be safely managed during the control loop.
[0009] To do so, the control system is configured to obtain targets (aka target values, references, a target trajectory, etc.) specifying a target state of the chamber. The target state of the chamber defines the target configuration of the plasma within the chamber. In some implementations, the target state of the chamber also defines a desired, target temperature distribution over the first wall of the chamber. In these cases, the control system controls the magnetic field to establish the target configuration and temperature distribution simultaneously, or to establish a state of the chamber that is balanced between these two operational constraints. In other implementations, the targetAttorney Docket No. 45288-0528WO1temperature distribution follows as a result of the target configuration of the plasma and a learned control policy. For example, the control system can be trained to establish the target configuration such that the temperature at each point on the first wall is less than a maximum allowable temperature for the first wall. Note, the maximum allowable temperature for the first wall can depend on a number of factors such as material properties, engineering constraints, and operational requirements of the chamber. For example, the maximum allowable temperature for the first wall can be about 1,200 °C or less, e.g., 1,100 °C or less, 1,000 °C or less, 900 °C or less, 800 °C or less, 700 °C or less, 600 °C or less, 500 °C or less.
[0010] As the target configuration is a diverted configuration, e.g., a single- or multi-null diverted configuration, the target configuration includes one or more target x-points positioned on a target separatrix for the magnetic field generated within the chamber. The separatrix is a boundary that separates closed magnetic field lines confining the plasma from open magnetic field lines which are controlled by the control system to divert excess heat and particle exhaust away from the confined plasma. An x-point is a specific point where the magnetic field lines intersect and change direction, forming a characteristic “X” shape. When positioned on a separatrix, an x-point marks the transition between closed and open magnetic field lines.
[0011] Generally, the target configuration is defined such that each target x-point has a pair of divertor legs terminating on one of the divertor(s) positioned within the chamber, safely directing heat and particle flux exhausted from the plasma away from the first wall. The divertor legs of a target x-point are the extensions of the open magnetic field lines that connect the target x-point to divertor targets of the divertorthat are configured to absorb high heat and particle fluxes. To define the target configuration, the target state of the chamber includes a respective target value for each of multiple configuration parameters of the target configuration of the plasma within the chamber. The configuration parameters include plasma shape parameters, plasma control parameters, and magnetic configuration parameters that specify the geometry and characteristics of the plasma and the magnetic field within the chamber. For example, the configuration parameters can include a plasma current of the plasma, a contour of the target separatrix, a total number of the target x-point(s), a position of each target x-point on the target separatrix, a path of each divertor leg from each target x-point, a position of a center of the plasma, an elongation of the plasma, a triangularity of the plasma, a radius of the plasma, and a position of a limit point of the plasma.Attorney Docket No. 45288-0528WO1
[0012] The desired, target temperature distribution can be established by the control system such that no one point along the first wall experiences a high enough heat load sufficient to cause damage thereof, e.g., melting, erosion, or thermal expansion of the first wall outside safe limits. As mentioned above, in some implementations, the control system establishes the target temperature distribution as a result of the target configuration of the plasma and a learned control policy. In other implementations, the target temperature distribution is an additional constraint imposed on the control policy executed by the control system. In these cases, the target state of the chamber can include a respective target temperature for each of multiple points representing the first wall of the chamber. For example, the target temperature of each point can be less than the maximum allowable temperature for the first wall.
[0013] The control system processes the targets and measurements, using a controller neural network, to generate control inputs for the set of control coils that manipulate the plasma toward the target configuration that establishes the target temperature distribution. To learn an appropriate control policy, the controller neural network is trained on a model of the magnetic confinement device using a reinforcement learning technique, e.g., an actor-critic reinforcement learning technique, a REINFORCE learning technique, or a Dyna-Q reinforcement learning technique. The controller neural network can use a stochastic control policy, a deterministic control policy, or combinations of both during training and / or upon deployment. For example, the controller neural network can explore a large action (i.e., control) space during training by sampling actions from a control policy and thereafter use the expectation value of the control policy at inference.
[0014] The model of the magnetic confinement device typically includes a power supply model for the set of control coils of the magnetic confinement device, an evolution model for the chamber of the magnetic confinement device, and a sensor model for the sensors of the magnetic confinement device.
[0015] The power supply and evolution models are coupled to one another as the set of control coils and the plasma both induce electric current in one another. The power supply model describes the electrical properties of the set of control coils, e.g., resistance, capacitance, and inductance, as well as the inductive coupling with the plasma, to predict the strength and topology of the magnetic field resulting from a particular electric current through each control coil.
[0016] The evolution model is a combined, dynamical model of the plasma describing the evolution and heat transported by the plasma within the chamber, including the interaction of theAttorney Docket No. 45288-0528WO1plasma with the magnetic field and the PFCs of the chamber. The evolution model, thus, predicts the current state of the plasma within the chamber, including the current configuration of the plasma and the current temperature distribution over the PFCs, under the influence of the resultant magnetic field. For example, the evolution model can include a forward Grad-Shafranov evolutive (“FGE”) solver and a heat transport model that are coupled to each other and together describe plasma dynamics, magnetic equilibrium, and heat transport of the plasma in terms of flux coordinates of the magnetic field, e.g., such that the heat flux of the plasma is a function of the (poloidal) magnetic flux.
[0017] An example of a free-boundary FGE solver is provided by Amorisco, N. C., et al. “FreeGSNKE: A Python-based dynamic free-boundary toroidal plasma equilibrium solver,” Physics of Plasmas 31.4 (2024). In the magnetohydrodynamics (“MHD”) picture, heat transport is often modeled diffusively, e g., in terms of an anisotropic Fick’s law, providing the heat flux as a gradient of temperature along magnetic field lines. See for example, Chen, Francis F, “Introduction to Plasma Physics and Controlled Fusion,” Vol. 1. New York: Plenum Press (1984) or Schekochihin, Alexander A., “Lectures on Kinetic Theory and Magnetohydrodynamics of Plasmas,” Lecture Notes for the Oxford MMathPhys Programme (2022) and Eich, Thomas, et al. “Scaling of the tokamak near the scrape-off layer H-mode power width and implications for ITER,” Nuclear fusion 53.9 (2013): 093031.
[0018] The sensor model predicts the measurements that would be collected by the sensors in response to the current state of the plasma within the chamber, e.g., including the magnetic field and flux measurements of the magnetic field, the electric current measurements in the set of control coils, and the temperature measurements of the plasma and the PFCs.
[0019] Using a reinforcement learning technique combined with a reward function and a sufficiently robust model of the magnetic confinement device, the controller neural network can learn a near-optimal control policy for the set of control coils that optimizes heat flux deposition of any desired diverted plasma configuration, providing control systems for magnetic confinement devices that can be configured for sustained thermonuclear fusion, advanced plasma research, safety tests, among other applications.
[0020] These and other features of the control system described herein are summarized below.
[0021] According to a first aspect, a method performed by one or more computers for controlling a magnetic confinement device including a chamber and a set of control coils configured toAttorney Docket No. 45288-0528WO1generate a magnetic field within the chamber in accordance with a respective control input at each of multiple time steps is described. The method includes, at each of the time steps: obtaining a target state of the chamber defining a target, diverted configuration of a plasma within the chamber that establishes a target temperature distribution over a first wall of the chamber, where the target configuration includes one or more target x-points positioned on a target separatrix for the magnetic field, each target x-point having a pair of divertor legs terminating on a respective divertor positioned within the chamber; receiving, from the magnetic confinement device, a measurement of a current state of the chamber characterizing: (i) a current configuration of the plasma within the chamber, and (ii) a current temperature distribution over the first wall of the chamber; preparing, for a controller neural network, a network input including: (i) the target state of the chamber, and (ii) the measurement of the current state of the chamber; processing the network input, using the controller neural network, to generate a network output defining a policy for selecting control inputs for the set of control coils at the time step; selecting the control input at the time step using the network output generated by the controller neural network at the time step; and transmitting, to the magnetic confinement device, the control input at the time step to generate the magnetic field within the chamber in accordance therewith.
[0022] In some implementations of the method, at each of the time steps, the policy is a probability distribution over possible control inputs for the set of control coils at the time step, and the network output includes a set of parameters of the probability distribution at the time step.
[0023] In some implementations of the method, at one or more of the time steps, selecting the control input at the time step using the network output generated by the controller neural network at the time step includes: computing an expected value of the probability distribution from its set of parameters at the time step; and selecting, as the control input at the time step, the mean of the probability distribution.
[0024] In some implementations of the method, at one or more of the time steps, selecting the control input at the time step using the network output generated by the controller neural network at the time step includes: generating the probability distribution in accordance with its set of parameters at the time step; and sampling, from the probability distribution, the control input at the time step.Attorney Docket No. 45288-0528WO1
[0025] In some implementations of the method, at each of the time steps, the probability distribution is a Gaussian distribution, and the set of parameters at the time step includes a mean and covariance of the Gaussian distribution at the time step.
[0026] In some implementations of the method, at each of the time steps, the control input at the time step includes a respective control voltage to be applied to each control coil at the time step.
[0027] In some implementations of the method, the target configuration of the plasma is a singlenull diverted configuration, and the one or more target x-points is a single target x-point.
[0028] In some implementations of the method, the target configuration is a multi-null diverted configuration, and the one or more target x-points is multiple target x-points.
[0029] In some implementations of the method, at each of the time steps, the target state of the chamber at the time step includes, for each of multiple configuration parameters of the target configuration of the plasma within the chamber, a respective target value of the configuration parameter at the time step.
[0030] In some implementations of the method, the configuration parameters include: a plasma current of the plasma, a contour of the target separatrix, a total number of the one or more target x-points, and, for each target x-point, a respective position of the target x-point on the target separatrix.
[0031] In some implementations of the method, the configuration parameters further include, for each divertor leg of each target x-point, a respective path of the divertor leg from the target x-point to the divertor for the target x-point.
[0032] In some implementations of the method, the configuration parameters further include one or more of: a position of a center of the plasma, an elongation of the plasma, a triangularity of the plasma, a radius of the plasma, or a position of a limit point of the plasma.
[0033] In some implementations of the method, at each of the time steps, the target state of the chamber at the time step further includes, for each of multiple points representing the first wall of the chamber, a respective target temperature of the point at the time step.
[0034] In some implementations of the method, at each of the time steps, the respective target temperature of each of the points at the time step is less than a maximum allowable temperature for the first wall.
[0035] In some implementations of the method, the maximum allowable temperature for the first wall is 1,200 Celsius (°C) or less.Attorney Docket No. 45288-0528WO1
[0036] In some implementations of the method, the maximum allowable temperature for the first wall is 7,000°C or less.
[0037] In some implementations of the method, at each of the time steps, the target state of the chamber at the time step further includes, for each control coil in the set, a respective target electric current in the control coil at the time step.
[0038] In some implementations of the method, at each of the time steps, the measurement of the current state of the chamber at the time step includes: a set of magnetic field measurements including, for each of multiple magnetic field probes of the magnetic confinement device, a respective measurement of the magnetic field collected from the magnetic field probe at the time step; a set of magnetic flux measurements including, for each of multiple magnetic flux loops of the magnetic confinement device, a respective measurement of a flux of the magnetic field collected from the magnetic flux loop at the time step; a set of electric current measurements including, for each control coil in the set, a respective measurement of an electric current in the control coil at the time step; and a set of temperature measurements including, for each of multiple temperature sensors of the magnetic confinement device, a respective measurement of the current temperature distribution over the first wall collected from the temperature sensor at the time step.
[0039] In some implementations of the method, the magnetic confinement device is a tokamak, and the chamber is a toroidal chamber having a central axis of revolution.
[0040] In some implementations of the method, the set of control coils includes: a set of toroidal control coils encircling the toroidal chamber, the set of toroidal control coils configured to generate a toroidal field component of the magnetic field; a set of poloidal controls coils positioned about the central axis, the set of poloidal control coils configured to generate a poloidal field component of the magnetic field; and a central solenoidal control coil extending along the central axis, the central solenoidal control coil configured to induce a toroidal plasma current in the plasma.
[0041] In some implementations of the method, the set of control coils is a set of high-temperature superconducting control coils.
[0042] In some implementations of the method, each of the time steps has a length of ten milliseconds or less.
[0043] In some implementations of the method, the time steps have a total duration of ten seconds or more.Attorney Docket No. 45288-0528WO1
[0044] In some implementations of the method, the controller neural network has been trained on a model of the magnetic confinement device using a reinforcement learning technique.
[0045] In some implementations of the method, the model of the magnetic confinement device includes a power supply model for the set of control coils, an evolution model for the chamber, and a sensor model for multiple sensors of the magnetic confinement device.
[0046] In some implementations of the method, the evolution model includes a forward Grad-Shafranov evolutive (“FGE”) solver and a heat transport model.
[0047] In some implementations of the method, the heat transport model is a diffusive heat transport model.
[0048] In some implementations of the method, the magnetic confinement device generates electrical power via thermonuclear fusion of the plasma.
[0049] In some implementations of the method, the controller neural network is a feedforward neural network.
[0050] In some implementations of the method, the controller neural network includes ten neural network layers or less.
[0051] According to a second aspect, a method performed by one or more computers for training a controller neural network on a model of a magnetic confinement device including a chamber and a set of control coils configured to generate a magnetic field within the chamber in accordance with a respective control input at each of multiple time steps is described. The method includes: executing, over the time steps, a simulation including the model of the magnetic confinement device and a reward function; at each of the time steps: obtaining a target state of the chamber defining a target, diverted configuration of a plasma within the chamber that establishes a target temperature distribution over a first wall of the chamber, where the target configuration includes one or more target x-points positioned on a target separatrix for the magnetic field, each target x-point having a pair of divertor legs terminating on a respective divertor positioned within the chamber; receiving, from the simulation, a model output including: a measurement of a current state of the chamber characterizing: (i) a current configuration of the plasma within the chamber, and (ii) a current temperature distribution over the first wall of the chamber; and a reward characterizing an error between: (i) the target state of the chamber, and (ii) the current state of the chamber; preparing, for the controller neural network, a network input including: (i) the target state of the chamber, and (ii) the measurement of the current state of chamber; processing the networkAttorney Docket No. 45288-0528WO1input, using the controller neural network, to generate a network output defining a policy for selecting control inputs for the set of control coils at the time step; selecting the control input at the time step using the network output generated by the controller neural network at the time step; and transmitting, to the simulation, the control input at the time step to generate the magnetic field within the chamber in accordance therewith; and training the controller neural network on the rewards using a reinforcement learning technique.
[0052] In some implementations, the method further includes controlling the magnetic confinement device using the controller neural network trained on the model of the magnetic confinement device.
[0053] In some implementations of the method, the model of the magnetic confinement device includes a power supply model for the set of control coils, an evolution model for the chamber, and a sensor model for multiple sensors of the magnetic confinement device, and executing, over the time steps, the simulation including the model of the magnetic confinement device and the reward function includes, at each of the time steps: receiving a model input including: (i) a current state of the chamber at a preceding time step, and (ii) a control input for the set of control coils at the preceding time step; processing the model input, using the power supply and evolution models, to generate the current state of the chamber at the time step; processing the current state of the chamber at the time step, using the sensor model, to generate the measurement of the current state of the chamber at the time step; and processing the target and current states of the chamber at the time step, using the reward function, to generate the reward for the time step.
[0054] In some implementations of the method, executing, over the time steps, the simulation including the model of the magnetic confinement device and the reward function includes, at each of the time steps: determining, from the current state of the chamber at the time step, whether a physical feasibility constraint of the magnetic confinement device is violated at the time step; and if a physical feasibility constraint of the magnetic confinement device is violated at the time step, terminating the simulation at the time step.
[0055] In some implementations of the method, at each of the time steps, determining, from the current state of the chamber at the time step, whether the physical feasibility constraint of the magnetic confinement device is violated at the time step includes one or more of: determining whether a density of the plasma satisfies a threshold at the time step; determining whether a plasma current of the plasma satisfies a threshold at the time step; determining whether a plasma safetyAttorney Docket No. 45288-0528WO1factor of the plasma satisfies a threshold at the time step; determining whether a respective current through each of one or more control coils in the set satisfies a threshold at the time step; or determining whether the current temperature distribution over the first wall satisfies a threshold at the time step.
[0056] In some implementations of the method, at each of the time steps: the target state of the chamber at the time step includes, for each of multiple configuration parameters of a configuration of the plasma within the chamber, a respective target value of the configuration parameter at the time step, the current state of the chamber at the time step includes, for each of the configuration parameters, a respective current value of the configuration parameter at the time step, and the reward for the time step includes, for each of the configuration parameters, a respective error between: (i) the target value of the configuration parameter at the time step, and (ii) the current value of the configuration parameter at the time step.
[0057] In some implementations of the method, at each of the time steps, the reward for the time step includes a weighted linear combination of the respective errors for each of the configuration parameters at the time step.
[0058] In some implementations of the method, the configuration parameters include: a plasma current of the plasma, a contour of a separatrix, a total number of x-points, and, for each x-point, a respective position of the x-point.
[0059] In some implementations of the method, the configuration parameters further include one or more of: a position of a center of the plasma, an elongation of the plasma, a triangularity of the plasma, a radius of the plasma, or a respective position of each of one or more limit points of the plasma.
[0060] In some implementations of the method, at each of the time steps: the target state of the chamber at the time step further includes, for each of multiple points representing the first wall of the chamber, a respective target temperature of the point at the time step, the current state of the chamber at the time step further includes, for each of the points, a respective current temperature of the point at the time step, and the reward for the time step further includes, for each of the points, a respective error between: (i) the target temperature of the point at the time step, and (ii) the current temperature of the point at the time step.
[0061] In some implementations of the method, at one or more of the time steps, the reward for the time step further includes, for each target x-point, one or more of: a gradient of a flux of theAttorney Docket No. 45288-0528WO1magnetic field at the target x-point, e g., a gradient of the poloidal flux at the target x-point; a difference between: (i) the flux of the magnetic field at the target x-point, and (ii) the flux of the magnetic field at a current separatrix of the magnetic field, e.g. a difference between the poloidal flux at the target x-point and the poloidal flux at the separatrix; and for each divertor leg of the target x-point, a difference between: (i) the flux of the magnetic field along a path of the divertor leg, and (ii) the flux of the magnetic field at the current separatrix. Also or instead the reward for a time step can include a particle flux flowing along one or each divertor leg of the (target) x-point towards a divertor.
[0062] In some implementations of the method, at one or more of the time steps, the reward for the time step further includes, for each current x-point of the magnetic field, a shortest distance of the current x-point to the current separatrix.
[0063] In some implementations of the method, the reinforcement learning technique is an actor-critic reinforcement learning technique, and training the controller neural network on the rewards using the actor-critic reinforcement learning technique includes jointly training the controller neural network and a critic neural network on the rewards.
[0064] In some implementations, the method further includes, for each of the time steps: preparing, for the critic neural network, a critic input including: (i) the current state of the chamber at the time step, and (ii) the reward for the time step; and processing the critic input, using the critic neural network, to generate a critic output estimating a cumulative measure of the rewards received from the simulation at each proceeding time step.
[0065] In some implementations of the method, for each of the time steps, the critic input at the time step further includes one or more of: the target state of the chamber at the time step, the measurement of the current state of the chamber at the time step, or the control input for the set of control coils at the time step.
[0066] In some implementations of the method, jointly training the controller and critic neural networks on the rewards includes: generating an objective function that depends on the rewards and critic outputs; and optimizing the objective function with respect to a respective set of network parameters of each of the controller and critic neural networks.
[0067] In some implementations of the method, the actor-critic reinforcement learning technique is a maximum a posteriori policy optimization (“MPO”) technique.Attorney Docket No. 45288-0528WO1
[0068] In some implementations of the method, the actor-critic reinforcement learning technique is a distributed actor-critic reinforcement learning technique.
[0069] In some implementations of the method, the controller neural network has fewer network parameters than the critic neural network.
[0070] In some implementations of the method, the controller neural network is a feedforward neural network, and the critic neural network is a recurrent neural network.
[0071] In some implementations, the method further includes after training the controller neural network, using the controller neural network to control the real-world magnetic confinement device by, at each of a plurality of time steps: processing the network input comprising: (i) the target state of the real-world chamber, and (ii) the measurement of the current state of the real-world chamber, using the controller neural network, to generate the network output, and transmitting, to the real-world magnetic confinement device, a control input for the set of control coils selected using the network output.
[0072] In some implementations, the controller neural network has been trained by the first or second aspects in any of their abovementioned implementations.
[0073] According to a third aspect, a system including one or more computers and one or more storage devices communicatively coupled to the one or more computers is described. The one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the method of the first or second aspects in any of their abovementioned implementations.
[0074] According to a fourth aspect, one or more non-transitory computer storage media is described. The one or more non-transitory computer storage media store instructions that, when executed by one or more computers, cause the one or more computers to perform the method of the first or second aspects in any of their abovementioned implementations.
[0075] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.
[0076] This specification provides control systems for magnetic confinement devices that utilize controller neural networks trained via reinforcement learning. Particularly, the controller neural networks described herein are trained on robust models of magnetic confinement devices that emulate plasma evolution, magnetic equilibrium, heat transport, power supplies, and sensing instruments. Using such models, the control system can simultaneously establish diverted plasmaAttorney Docket No. 45288-0528WO1configurations, e.g., that exhaust heat and particle fluxes onto divertors at specified, target rates, while maintaining temperatures of other PFCs (e.g., the first wall) below thresholds.
[0077] Moreover, as the control systems described herein utilize a neural network architecture, they can be configured as nonlinear feedback controllers for any magnetic confinement device. That is, the controller neural networks can autonomously learn a near-optimal control policy to efficiently command the unique set of controls of a magnetic confinement device, yielding a notable reduction in design effort compared with conventional magnetic controllers. A single, computationally inexpensive control system can replace a magnetic controller’s complex nested control architecture. This approach can have unprecedented flexibility and generality due to specifying control objectives at a high-level, which shifts the focus towards what the magnetic confinement device should accomplish, rather than how it can be accomplished.
[0078] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0079] FIG. 1A is a schematic diagram depicting an example of a magnetic confinement system including a magnetic confinement device and a control system for controlling the magnetic confinement device.
[0080] FIG. IB is a schematic diagram depicting an example of a divertor for a chamber of the magnetic confinement device shown in FIG. 1A.
[0081] FIG. 1C illustrate examples of target states and cross-sections of the chamber of the magnetic confinement device shown in FIG. 1A.
[0082] FIG. 2 is as schematic diagram depicting an example of a magnetic confinement training system that can train a controller neural network on a model of a magnetic confinement device.
[0083] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION
[0084] Controlled thermonuclear fusion, the fundamental process behind fusion reactors and other thermonuclear systems, is a promising solution for sustainable energy. Fusion reactors can use heat generated from fusion reactions occurring in a hot plasma to produce electrical power withAttorney Docket No. 45288-0528WO1very little radioactive waste. Aneutronic fusion reactors have potential for even greater efficiency as they can produce electrical power directly from charged particles emitted from the plasma. That said, one of the most challenging problems in achieving controlled nuclear fusion is confining the high-temperature, high-pressure plasma within a suitable chamber, such that energy can be safely extracted from the plasma. Due to the extreme temperatures (e.g., tens to hundreds of millions of degrees Celsius), the plasma cannot be in direct contact with any surface of the chamber and must be suspended in a vacuum within it, which is further complicated by inherent instabilities of the plasma.
[0085] However, since a plasma is an ionized gas that conducts electricity, it produces strong magnetic fields, and in turn, can be manipulated by strong magnetic fields. Magnetic confinement devices, such as tokamaks, utilize a time-varying arrangement of magnetic fields to shape and confine plasmas into various plasma configurations. In a tokamak, for example, the plasma is confined into a toroidal axisymmetric configuration (e.g., a donut-like shape) that conforms to the toroidal shape of the chamber. As another example, in a stellarator where the chamber has a helical toroid shape, the plasma is confined into a helical toroid configuration with twisted magnetic field lines. Examples of tokamaks currently deployed or in development include International Thermonuclear Experimental Reactor (“ITER”), Joint European Torus (“JET”), Experimental Advanced Superconducting Tokamak (“EAST”), Tokamak a Configuration Variable (“TCV”), and SPARC in development by Commonwealth Fusion Systems (“CFS”). Examples of other types of magnetic confinement devices currently deployed or in development include spherical tokamaks such as Mega Ampere Spherical Tokamak (“MAST”); stellarators such as Wendelstein 7-X and Helias designs (e g., HELIUS and HELIOS); spheromaks such as MS-Stellarator and SPP (Spheromak Plasma Experiment); Field-Reversed Configurations (“FRC”) such Princeton Field-Reversed Configuration (“PFRC”) and Variable Field FRC (“VF-10”); Magnetic Target Fusion (MTF) devices such as Z-Pinch devices (e.g., the Sandia National Laboratories’ Z Machine) and LTX (LTX-P) devices; Reverse Field Pinch (“RFP”) devices such as RFX-Mod and Pegasus; and Doughnut-Free Configuration devices.
[0086] Conventional magnetic controllers have commonly attacked the high-dimensional, high-frequency, nonlinear problem of plasma confinement using a model predictive control or a set of independent single-input single-output proportional-integral-derivative (PID) controllers that adjust various features of the plasma. The set of PID controllers must be designed to avoid mutualAttorney Docket No. 45288-0528WO1interference and are often further augmented by an outer control loop that implements real-time estimation of the plasma equilibrium. Other types of linear controllers, as well as nonlinear controllers, have also be employed. Although these magnetic controllers have been successful in certain situations, they involve considerable engineering effort and expertise whenever the target plasma configuration is changed. Moreover, these magnetic controllers must be designed for each magnetic confinement device and their unique set of controls, e.g., a unique number, arrangement, and field strength of a particular set of control coils, which can be a painstaking task as successive generations of magnetic confinement devices come online. Further still, such magnetic controllers are not typically optimized for managing the heat and particle flux exhausted from the hot plasma. These challenges arise primarily from the need to protect the materials in contact with the plasma, keep temperatures of plasma-facing components (“PFCs”) below certain safety thresholds, ensure long device lifetimes, and maintain plasma stability, all of which will be important aspects for magnetic confinement devices that continuously generate electrical power via thermonuclear fusion of the plasma.
[0087] To overcome some, or all, of these abovementioned challenges, this specification provides control systems for magnetic confinement devices that utilize controller neural networks trained via reinforcement learning. Particularly, the controller neural networks described herein are trained to simultaneously establish diverted plasma configurations, e.g., that exhaust heat and particle fluxes onto a divertor at specified, target rates, while maintaining temperatures of other PFCs (e.g., the first wall) below thresholds. Moreover, as the control systems described herein utilize a neural network architecture, they can be configured as nonlinear feedback controllers for any magnetic confinement device. That is, the controller neural networks can autonomously learn a near-optimal control policy to efficiently command the unique set of controls of a magnetic confinement device, yielding a notable reduction in design effort compared with conventional magnetic controllers. A single, computationally inexpensive control system can replace a magnetic controller’s complex nested control architecture. This approach can have unprecedented flexibility and generality due to specifying control objectives at a high-level, which shifts the focus towards what the magnetic confinement device should accomplish, rather than how it can be accomplished.
[0088] These and other features relating to the control systems, controller neural networks, and reinforcement learning techniques provided herein are described in more detail below.Attorney Docket No. 45288-0528WO1
[0089] FIG. 1 A is a schematic diagram depicting an example of a magnetic confinement system 10 including a magnetic confinement device 200 and a control system 100 for controlling (or otherwise actuating) the magnetic confinement device 200. The control system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented. For example, the control system 100 can be an integrated control system of the magnetic confinement device 200, or a module of an integrated control system of the magnetic confinement device 200.
[0090] At a high-level, the magnetic confinement device 200 includes a chamber 210 (aka a vessel), a set of control coils 220, and multiple sensors 230.
[0091] The chamber 210 is configured to support a (e.g., extremely) low-pressure environment, referred to as a vacuum, for a plasma 400 generated therein. For example, the chamber 210 can have a base pressure in range from about 10’6Torr to 10'8Torr, or less. The base pressure refers to the pressure within the chamber 210 before any gas is introduced for plasma generation. The base pressure of the chamber 210 can be established using vacuum pumps, e.g., cryogenic, turbo, and / or diffusion pumps. Once the plasma 400 is generated, e.g., via plasma breakdown of a fuel gas, the chamber 210 is typically filled with a small amount of the fuel gas. Examples of fuel gases include hydrogen isotopes (e.g., deuterium and tritium), and helium-3. Protium, the most common isotope of hydrogen, may also be used for initial plasma startup (as well as cleaning). The fuel gas dispersed after plasma breakdown increases the pressure within the chamber 210, but the operating pressure remains sufficiently low to allow for efficient confinement and heating of the plasma 400. The operating pressure refers to the pressure within the chamber 210 during confinement of the plasma 400 therein. For example, the chamber 210 can have an operating pressure in range from about 10-3Torr to 10-4Torr, or less.
[0092] As shown in FIG. 1 A, the chamber 210 is a toroidal chamber having a central axis 211 of revolution, i.e., the chamber 210 has toroidal axisymmetry about the central axis 211. Here, the magnetic confinement device 200 can be a tokamak or other magnetic confinement device utilizing a toroidal chamber. In some implementations, the chamber 210 has a major radius (R) in range from about 1 meter to 10 meters. The major radius refers to the radial distance from the central axis 211 to the center 402 of the plasma 400. The chamber 210 can have a minor radius (a) in range from about 0.5 meters to 2 meters. The minor radius refers to the radius of the plasma 400’s cross-section, measured from the center 402 of the plasma 400 to a separatrix 420 bounding theAttorney Docket No. 45288-0528WO1plasma 400. The separatrix 420, also referred to as the last closed flux surface (“LCFS”), is the outermost magnetic surface that contains the plasma 400. The chamber 210 can have an aspect ratio (A) in a range from about 2 to 5. The aspect ratio A = R / a refers to the ratio of the major radius to the minor radius of the chamber 210.
[0093] The toroidal axisymmetry of the chamber 210 provides a number of advantages such as enhancing plasma confinement, improving plasma stability, optimizing heating and current drive, simplifying engineering and maintenance, enhanced heat and particle exhaust, among others. This is due, at least in part, to the toroidal axisymmetric configuration of the plasma 400 generated within the chamber 210, which can be maintained by uniform and continuous magnetic fields generated within and about the plasma 400. Moreover, as the plasma 400 has a toroidal axisymmetric configuration, it can be modeled relatively efficiently via two-dimensional magnetohydrodynamics (“MHD”), e.g., the Grad-Shafranov equation. Such two-dimensional MHD equations describe the time-dependent or steady-state (e.g., equilibrium) behavior of a crosssection of the plasma 400, as opposed to full, three-dimensional MHD equations which can be computationally intractable. Nevertheless, the chamber 210 can have different topologies and geometries depending on the desired configuration of the plasma 400 confined therein. For example, in some implementations, the magnetic confinement device 200 can be a stellarator, in which case the chamber 210 is a helical toroid chamber, and the plasma 400 has a helical toroid configuration with twisted magnetic field lines that do not form an axisymmetric structure. In this case, magnetic fields are provided by helical control coils that create a complex, three-dimensional magnetic structure.
[0094] As shown in the cross-section of the chamber 210, the chamber 210 includes multiple PFCs positioned therein, including a first wall 212, a baffle 214, and a divertor 300. The cross-section of the chamber 210 also shows a current state 512 of the chamber 210, including a current configuration of the plasma 400 within the chamber 210.
[0095] The first wall 212 conforms to the interior surface of the chamber 210 and is positioned close to the interior surface of the chamber 210. The first wall 212 is the innermost layer of the chamber 210 that directly faces the plasma 400 and defines a boundary between the plasma 400 and the chamber 210, preventing direct contact between the two. During operation, the first wall 212 can be subjected to intense heat fluxes generated from the plasma 400, particularly during transient events like disruptions or edge localized modes (“ELMs”). The first wall 212 can beAttorney Docket No. 45288-0528WO1designed to manage such heat fluxes and withstand steady-state thermal loads. For example, the first wall 212 can include an array of tiles that are composed of and / or coated with high temperature resistant materials, e.g., tungsten, beryllium, carbon-based composites (e.g., graphite and carbon fiber composites), advanced alloys (e.g., steel alloys), and combinations thereof.
[0096] Nevertheless, maintaining temperatures of the first wall 212 below thresholds is an important aspect for long term survivability and optimization of the magnetic confinement device 200, e.g., for continuous thermonuclear fusion. For example, tungsten has a high melting point of about 3422 degrees Celsius (°C), but has a maximum allowable temperature range of about 1200°C to 1300°C as it experiences thermal fatigue and cracking for continuous operation at these temperatures. As another example, beryllium has a low atomic number which reduces plasma contamination and radiation losses, but has a lower melting point of about 1287°C, limiting its maximum allowable temperature range to about 600°C to 800°C. As yet another example, carbonbased composites have high thermal conductivities, but erosion and thermal expansion limits their maximum allowable temperature ranges to about 1000°C to 1200°C. As yet another example, reduced-activation ferritic-martensitic (“RAFM”) steels are sometimes used for portions of the first wall 212, e g., for structural support, but have maximum allowable temperature ranges of about 500°C to 600°C.
[0097] The divertor 300 is positioned near the bottom of the chamber 210 and is configured to absorb excess heat flux that is exhausted from the plasma 400. For example, the divertor 300 can be configured to absorb heat fluxes of about 10 megawatts per meter-squared (MW / m2) to 20 MW / m2, or more. The divertor 300 can also be configured to absorb fluxes of plasma particles and neutrals (neutralized atoms) that are exhausted from the plasma 400. The baffle 214 is positioned above the divertor 300 and is configured to guide the fluxes of heat, plasma particles, and neutrals exhausted from the plasma 400 toward the divertor 300. This generally improves the efficiency of the divertor 300.
[0098] As shown in FIG. IB, the divertor 300 includes a cassette body 310 and multiple divertor targets 320 mounted to the cassette body 310 via a set of mounting interfaces 326. The mounting interfaces 326 position the divertor targets 320 within the chamber 210 at appropriate orientations to receive the heat and particle fluxes from the plasma 400, while providing a (near) vacuum boundary between the divertor targets 320 and the cassette body 310 to reduce conductive heat transfer therebetween. Here, the divertor targets 320 include an inboard divertor target 320-1, anAttorney Docket No. 45288-0528WO1outboard divertor target 320-2, and a dome (aka central divertor target) 320-3 that each have a curved profile to spread the heat load more evenly across the divertor 300. Particularly, the geometry of the divertor targets 320 is designed to account for the rate and energy of particle strikes at particular angles from the plasma 400, thereby distributing the heat flux evenly across the surface of the divertor targets 320. Each of the divertor targets 320-1, 320-2, and 320-3 includes a respective plasma-facing layer 322-1, 322-2, and 322-3 that is incident with the heat and particle fluxes from the plasma 400. The plasma-facing layers 322 withstand and dissipate the high heat loads (tens of megawatts per meter-squared as noted above) that arise from direct contact with the high-energy particles escaping from the plasma 400. Materials used in the plasma-facing layers 322, such as tungsten, are chosen for their high melting point, thermal conductivity, and resistance to erosion under these extreme conditions. Nevertheless, the divertor targets 320 are typically designed to be replaceable, e.g., removeable from the mounting interfaces 326, as these PFCs have the highest susceptibility to degradation.
[0099] As shown in FIG. IB, the divertor targets 320-1, 320-2, and 320-3 can also include one or more gas inlet valves 324-1, 324-2, 324-3, and 324-4. Each gas inlet valve (aka gas puff valve) 324 is configured to inject a gas puff 325 from an incident surface of one of the plasma-facing layers 322. The gas puff 325 then disperses near the surface of the plasma-facing layer 322 and interacts with the particle fluxes from the plasma 400. One main function of the gas inlet valves 324 is to enable plasma detachment in the divertor 300 region. For example, by injecting neutral gases (e.g., deuterium, neon, or argon), the plasma 400 in the divertor 300 region can be radiatively cooled and partially detached from the surface of the divertor targets 320. This helps control the temperature and heat load of the divertor 300, impurities in the plasma 400, and plasma stability. The divertor targets 320 can also include sensors 230, e.g., magnetic field probes, magnetic flux loops, and / or temperature sensors, as well as integrated cooling systems. The mounting interfaces 326 can include pathways for electrical circuitry, gas lines, and cooling lines for receiving measurements from the sensors 230 of the divertor targets 320, actuating the gas inlet valves 324 of the divertor targets 320, supplying gas to the gas inlets 324, and supplying coolant to the cooling systems of the divertor targets 320.
[0100] Returning to FIG. 1A, the set of control coils 220 arranged about the chamber 210 in various axisymmetric geometries that conform to the toroidal shape of the chamber 210. One or more of the control coils 220, e.g., fast G control coils 220G, can also be arranged within theAttorney Docket No. 45288-0528WO1chamber 210 behind the PFCs. The set of control coils 220 are configured to generate a magnetic field within the chamber 210. Here, the set of control coils 220 includes a set of toroidal control coils 220T and a set of poloidal control coils 220P. The set of toroidal control coils 220T encircle the chamber 210 and are configured to generate a toroidal field component of the magnetic field. The set of poloidal controls coils 220P includes inner and outer poloidal coils 220P that are positioned about the central axis 211. The poloidal control coils 220P are configured to generate a poloidal field component of the magnetic field. The set of control coils 220 can also include a central solenoidal control coil 220S extending along the central axis 211. The central solenoidal control coil 220S is configured to induce a toroidal plasma current in the plasma 400 for heating and current drive of the plasma 400. The set of control coils 220 can also include other types of control coils 220 such as ohmic control coils 220Q, for additional ohmic (or resistive) heating and current induction within the plasma 400. The ohmic control coils 2200 are designed to drive the initial current in the plasma 400 by inducing an electric field, which then heats the plasma through electrical resistance.
[0101] In some implementations, the set of a control coils 220 is a set of high-temperature superconducting (“HTS”) control coils, which can create powerful magnetic fields involved to confine the plasma 400, e.g., magnetic field strengths of up to about 20 tesla. For example, the set of control coils 220 can be composed of a high-temperature superconductor such as rare-earth barium copper oxide (“REBCO”), bismuth strontium calcium copper oxide (e.g., Bi-2223 or Bi-2212), or yttrium barium copper oxide (“YBCO”).
[0102] In general, the set of control coils 220 are configured to generate the magnetic field within the chamber 210 in accordance with an adaptive control sequence a(-) computed by the control system 100. The control sequence a(-) = {a0, alta2,... } includes a respective control input (an) 142. n for the set of control coils 220 at each of multiple time steps n = 0, 1, 2... N — 1, where N is the total number of time steps in the control sequence. A control input 142 may also be referred to as an “action” that is selected by the control system 100 to be performed by the magnetic confinement device 200 at a particular time step. For example, a control input an= {Vni} L0at a time step can include a respective control voltage (I to be applied to each control coil 220 at the time step, where I indexes each control coil 220 and M is the total number of control coils 220 of the magnetic confinement device 200. Determining respective control voltages is counterintuitive as the magnetic field of a coil is proportional to the current through the coil not to the appliedAttorney Docket No. 45288-0528WO1voltage, and the voltage which typically needs to change significantly to adjust the current to overcome the inductance (e.g. a high positive / negative voltage could be required to increase / decrease the current according to V = Ldl / dt + IR whereas a smaller voltage could be used to maintain a steady current). However controlling the voltage rather than the current can be advantageous in achieving a very fast response time, which is important for correcting plasma instabilities.
[0103] The sensors 230 are positioned within and / or about the chamber 210 and are configured to collect measurements 134 of the chamber 210 over each of the time steps of the control sequence. For example, the sensors 230 can include magnetic field probes for measuring the magnetic field generated within the chamber 210, magnetic flux loops for measuring a flux of the magnetic field generated within the chamber 210, electric current sensors for measuring a respective electric current in each control coil 220, and temperature sensors for measuring a temperature distribution over the first wall 212 of the chamber 210.
[0104] The control system 100 is communicatively coupled with the magnetic confinement device 200, e.g., via wired communication channels and / or wireless communication channels. The control system 100 includes a controller neural network 110 configured to control (or otherwise regulate) the magnetic field generated within the chamber 210 by actuating the set of control coils 220, e.g., to establish a stable configuration of the plasma 400 having a target plasma current, position, and shape within a target separatrix 420. For example, upon establishing a configuration of the plasma 400 in plasma equilibrium, sustained thermonuclear fusion may proceed. Several aspects of the plasma 400 and the magnetic confinement device 200 itself can also be studied at plasma equilibrium, e.g., stability of the plasma 400, heat and particle exhaust of the plasma 400, degradation of the PFCs such as the first wall 212, baffle 214, and divertor 300, degradation of the set of control coils 220, and / or degradation of the sensors 230, which can be useful information for research and development.
[0105] The control system 100 generates the adaptive control sequence by referencing a target state trajectory 120 at each time step. The target state trajectory £(■)=...} includes a respective target state (tn) 132.n of the chamber 210 for each of the time steps in the control sequence. The target state trajectory 120 serves as a predefined path or set of desired states that the magnetic confinement system 10 aims to follow over time. That is, the target state trajectory 120 represents the intended behavior of the magnetic confinement system 10, providing aAttorney Docket No. 45288-0528WO1benchmark for the control system 100 to achieve by adjusting to measurements 134. That is, the target state can change with time, e.g. to increase, decrease or control the size and shape of the plasma.
[0106] A target state 132 of the chamber 210 defines a target, diverted configuration of the plasma 400 within the chamber 210. As shown in the cross-section of the chamber 210, the target configuration includes a target (active) x-point 412A positioned on the target separatrix 420 that defines the boundary of the portion of the plasma 400 confined within the chamber 210. The region outside the target separatrix 420 is referred to as the scrape-off layer (“SOL”) where the plasma 400 interacts with the first wall 412. The target x-point 412A includes a pair of divertor legs 422 that terminate on the divertor targets 320 of the divertor 300 at respective strike points 423. Note, one or more vacuum (inactive) x-points 412V may also be present within the chamber 210.
[0107] In general, a configuration of the plasma 400 has multiple configuration parameters, including plasma shape parameters, plasma control parameters, and magnetic configuration parameters, that specify the geometry and characteristics of the plasma 400 and magnetic field within the chamber 210. To define the target configuration of the plasma 400, a target state 132 of the chamber 210 includes a respective target value for each configuration parameter.
[0108] For example, to specify a single- or multi-null diverted configuration, the configuration parameters can include one or more of a plasma current ( / p) of the plasma 400, a contour (Csep) of the target separatrix 420, a total number (Nx) of the target x-point(s) 412A, a respective position (qx) of each target x-point 412A on the target separatrix 420, and a respective path (Cleg) of each divertor leg 422 from each target x-point 412A to the divertor 300. Note that for paths, lines, and contours, the target state 132 typically includes a set of discrete points along the path, line, contour, see FIG. 1C for example. For additional specificity, the configuration parameters can also include one or more of a position (qc) of the center 402 of the plasma 400, a so-called elongation (K) of the plasma 400 (e.g. a ratio of the plasma’s height to its width), a so-called triangularity (<5) of the plasma 400 (e.g. dependent on a difference between a radial position of a top or bottom tip of the plasma and a radial position of its center as a fraction of a minor radius of the plasma), a radius (a) of the plasma 400, or a position (<7nm) of a so-called limit point of the plasma 400 (e.g. defining a boundary of the plasma). For example, a limit point for the plasma 400 may be set to prevent the plasma 400 from moving too close to the first wall 212, which can lead to excessive heat load and material erosion. The elongation, triangularity, and radius of the plasma 400 are plasma shapeAttorney Docket No. 45288-0528WO1parameters that help specify the geometry of the plasma 400’ s cross-section, influencing its stability, confinement, and overall performance.
[0109] The target configuration of the plasma 400 establishes a target temperature distribution over the first wall 212 of the chamber 210, when combined with the control policy of the control system 100. The target state 132 of the chamber 210 can also define the target temperature distribution over the first wall 212 as an additional constraint imposed on the control policy. In these cases, the target state 132 of the chamber 210 can include a respective target temperature (T) for each of multiple points representing the first wall 212 of the chamber 210. For example, the target temperature of each point can be less than the maximum allowable temperature for the first wall 212.
[0110] FIG. 1C illustrates examples of target states 132. i and 132.j of the chamber 210 and respective current states 512. i and 512.j of the chamber 210 at different time steps in the control sequence. Note that not all target values of the configuration parameters are shown in FIG. 1C.
[0111] Target state 132. i includes target values for the contour of the target separatrix 420 and the position of the target x-point 412A on the target separatrix 420 at time step n — i. Target state 132i can also include target temperatures for each of the points representing the first wall 212 of the chamber 210 at time step n = i. Current state 12.i shows the magnetic flux contours within the chamber 210 at time step n = i
[0112] Target state 132.j includes target values for the contour of the target separatrix 420 and the position of the target x-point 412A on the target separatrix 420, both of which have now changed at time step n — j. Target state 132.j further includes target values for each path of the divertor legs 422 that terminate at their respective strike points 423 on the divertor 300. Target state 132j can also include target temperatures for each of the points representing the first wall 212 of the chamber 210, which may be the same or different at time step n = j. Current state 512.j shows the magnetic flux contours within the chamber 210 at time step n = j. Heat flux 434 profiles at the strike points 423 are also shown in current state 512.j where heat exhausted from the plasma 400 is dissipated at the divertor 300.
[0113] Returning to FIG. 1A, at each time step, the control system 100 performs the following control loop to generate the control input 142. n at the time step.
[0114] The control system 100 obtains the target state 132. n of the chamber 210 at the time step. The control system 100 then receives, from the magnetic confinement device 200, a measurementAttorney Docket No. 45288-0528WO1(mn) 134.n of the current state (sn) 512.n of the chamber 210 at the time step. The measurement 134.n characterizes a current configuration of the plasma 400 within the chamber 210 and a current temperature distribution over the first wall 212 of the chamber 210. The control system 100 prepares a network input (on) 130. n for the controller neural network 110, where the network input 130. n includes the target state 132. n of the chamber 210 and the measurement 134. n of the current state 512.n of the chamber 210. Here, the network input 130. n can also be referred to as an “observation”on= of the current state 512. n of the chamber 210, biased by the target state 132. n of the chamber 210.
[0115] The control system 100 processes the network input 130. n, using the controller neural network 110, to generate a network output 140. n that defines a policy (7rn) for selecting control inputs for the set of control coils 220 at the time step. The operations of the controller neural network 110 can be concisely described as a function (o; 0), parametrized by a set of network parameters (0). In practice, the policy is a probability distribution 7r„(a) = 7r(n; A„) over possible control inputs 142 at the time step, wherenare a set of parameters of the probability distribution at the time step. The network output 140. n includes the set of parameters (Tn) of the probability distribution which completely specify the probability distribution at the time step given an assumed functional form of the probability distribution. Thus, the controller neural network 110 provides a mapping f: o ■-» n from an observation to a control policy (or the set of parameters of the policy), such that An= f on,' 0). In some implementations, the policy is a Gaussian (normal) distribution?rn(a)=and the network output 140. n includes a mean ( / r ) and covariance (En) of the Gaussian distribution, such that An= { / rn, Ln}. Other examples of probability distributions include, but are not limited to, beta distributions, softmax distributions, exponential family distributions, Laplacian distributions, and the like.
[0116] In general, the controller neural network 110 can have any appropriate neural network architecture that enables it to perform its described function, i.e., processing a network input 130 to generate a network output 140. In particular, the controller neural network 110 can include any appropriate types of neural network layers (e.g., fully-connected layers, recurrent layers, convolutional layers, self-attention layers, etc.) in any appropriate numbers (e.g., 5 layers, 25 layers, or 100 layers) and connected in any appropriate configuration (e.g., as a linear sequence of layers, residual configurations, etc.). In some implementations, the controller neural network 110 is a feedforward neural network including one or more feedforward layers. Feedforward layers areAttorney Docket No. 45288-0528WO1computationally fast as they mainly involve two operations: matrix multiplications and activation functions. To further increase speed, in some implementations, the controller neural network includes ten neural network layers or less, e.g., nine neural network layers or less, eight neural network layers or less, seven neural network layers, or less, six neural network layers or less, or five neural network layers or less.
[0117] The control system 100 then selects the control input 142. n at the time step using the network output 14O.n generated by the controller neural network 110 at the time step. For example, to implement a stochastic control policy, the control system 100 can generate the probability distribution in accordance with its set of parameters at the time step and then sample the control input 142. n from the probability distribution, an~nn(a). As another example, to implement a deterministic control policy, the control system 100 can compute the mean of the probability distribution from its set of parameters at the time step and then select the control input 142.n as the mean of the probability distribution at the time step, an=(a nn(ct)) = e.g., such that an= G inthe case of a Gaussian distribution. In some cases, a deterministic control policy can be advantageous for real-world deployment (aka inference) on the magnetic confinement device 200, while a stochastic control policy can be advantageous for training on a simulation 400 of the magnetic confinement device 200.
[0118] Finally, the control system 100 transmits, to the magnetic confinement device 200, the control input 142. n at the time step to generate the magnetic field within the chamber 210 in accordance therewith. This completes the control loop for the time step. The control system 100 subsequently performs this control loop at the next time step (n + 1), and the next time step (n + 2), and so on.
[0119] FIG. 2 is as schematic diagram depicting an example of a magnetic confinement training system 20 that can train the controller neural network 110 on a model 200-MOD of the magnetic confinement device 200 via a reinforcement learning technique. In this example, the reinforcement learning technique is an actor-critic reinforcement leaning technique (which may be a distributed technique, such as A3C) but other types of reinforcement learning techniques can also be implemented, such as a REINFORCE learning technique or a Dyna-Q reinforcement learning technique, or an MPO (maximum a posteriori policy optimization) technique (e.g. Abdolmaleki et al., arXiv: 1806.06920, 2018).Attorney Docket No. 45288-0528WO1
[0120] The training system 20 implements an episodic training approach to collect training data as the control system 100 interacts with a simulation 500. Each episode corresponds to a single simulation run performed by an actor 22 that terminates either when a termination condition is satisfied or when a fixed simulation time has passed in the episode. The training system 20 then sends an interaction trajectory 622 from the episode to a learner 24 which uses the training data to train the controller neural network 110. The training system 20 typically executes multiple actors 22 in parallel to continuously generate interaction trajectories 622 and update the controller neural network 110.
[0121] Actor:
[0122] Referring first to the actor(s) 22. An actor 22 is an instance of the control system 100 and the simulation 500 that interact with each other over an episode. The episode includes multiple interaction time steps n = 0, 1,2... N — 1, where N is the total number of interaction time steps in the episode corresponding to the fixed simulation time and the episode is terminated at interaction time step n < IV — 1 if a termination condition is satisfied.
[0123] The operations of the control system 100 over an episode are analogous to those described above in FIGs. 1A-1C as the simulation 500 aims to replicate the real-world conditions at inference. Particularly, at each interaction time step in the episode, the control system 100 performs a simulated control loop to generate a control input 142. n at the time step. To reiterate, the control system 100 obtains a target state 132.n of the chamber 210 at the time step. The control system 100 then receives, from the simulation 500, a measurement 134.n of a current state 512.n of the chamber 210 at the time step. The measurement 134.n characterizes a current configuration of the plasma 400 within the chamber 210 and a current temperature distribution over the first wall 212 of the chamber 210. The control system 100 prepares a network input 130. n for the controller neural network 110, where the network input 130. n includes the target state 132. n of the chamber 210 and the measurement 134. n of the current state 512. n of the chamber 210. The control system 100 processes the network input 130.n, using the controller neural network 110, to generate a network output 140. n. The control system 100 then selects the control input 142. n at the time step using the network output 14O.n generated by the controller neural network 110 at the time step. The control system 100 then transmits, to the simulation 500, the control input 142. n at the time step to generate the magnetic field within the chamber 210 in accordance therewith. The controlAttorney Docket No. 45288-0528WO1system 100 subsequently performs this simulated control loop at the next interaction time step (n + 1), and the next interaction time step (n + 2), and so on.
[0124] The simulation 500 includes the model 200-MOD of the magnetic confinement device 200 and a reward function 540. The model 200-MOD of the magnetic confinement device 200 includes an evolution model 510 for the chamber 210, a power supply model 520 for the set of control coils 220, and a sensor model 530 for the sensors 230. The reward function can process a current state and a target state of the chamber to generate a reward that characterizes an error between the current state of the chamber and the target state. The current / target state may be defined, e.g. by one or more configuration parameters defining a configuration of the plasma, and / or by a temperature or temperature distribution of the first wall of the chamber. Optionally the reward function may generate one or more rewards as previously described, e.g. dependent on a magnetic field / flux at a (target) x-point.
[0125] The evolution model 510 is a combined, dynamical model of the plasma 400 describing the evolution and heat transported by the plasma 400 within the chamber 210, including the interaction of the plasma 400 with the magnetic field and the PFCs of the chamber 210, e.g., the first wall 212 and the divertor 300. The evolution model 510, thus, predicts the current state 512 of the plasma within the chamber 210, including the current configuration of the plasma 400 and the current temperature distribution over the PFCs, under the influence of the resultant magnetic field. For example, the evolution model 510 can include a forward Grad-Shafranov evolutive (“FGE”) solver and a heat transport model that are coupled to each other and together describe plasma dynamics, magnetic equilibrium, and heat transport of the plasma in terms of flux coordinates of the magnetic field, e.g., such that the heat flux of the plasma is a function of the (poloidal) magnetic flux. The heat transport model can be a diffusive heat transport model, e.g. a Convection-Diffusion Model (CDM).
[0126] The power supply model 520 describes the electrical properties of the set of control coils 220, e.g., resistance, capacitance, and inductance, as well as the inductive coupling with the plasma 400, to predict the strength and topology of the magnetic field resulting from a particular electric current through each control coil 200. Note, the power supply 520 and evolution 510 models are generally coupled to one another as the set of control coils 220 and the plasma 400 both induce electric current in one another.Attorney Docket No. 45288-0528WO1
[0127] The sensor model 530 predicts the measurements 134 that would be collected by the sensors 230 in response to the current state 512 of the plasma 400 within the chamber 210, e.g., including the magnetic field and flux measurements of the magnetic field, the electric current measurements in the set of control coils 220, and the temperature measurements of the plasma 400 and the PFCs.
[0128] At each interaction time step in the episode, the simulation 500 performs the following operations:
[0129] The simulation 500 receives a model input including: (i) a current state 512. (n-1) of the chamber 210 at a preceding time step, and (ii) a control input 512. (n-1) for the set of control coils 220 at the preceding time step;
[0130] The simulation 500 processes the model input, using the power supply 520 and evolution 510 models, to generate the current state 512.n of the chamber 210 at the time step.
[0131] The simulation 500 processes the current state 512.n of the chamber 210 at the time step, using the sensor model 530, to generate the measurement 134. n of the current state 512.n of the chamber 210 at the time step.
[0132] Finally, the simulation 500 processes the target 132. n and current 512.n states of the chamber 210 at the time step, using the reward function 540, to generate a reward 542. n for the time step.
[0133] The simulation 500 subsequently performs these operations at the next interaction time step (n + 1), and the next interaction time step (n + 2), and so on.
[0134] After the episode, the actor 22 sends the interaction trajectory (d) 622 of the episode to the learner 24. In general, the interaction trajectory 622 includes all the simulation data generated at each interaction time step of the episode d = {£(■), s(-),m(-),a(')>r(‘) including each target state 132, each current state 512, each measurement 134, each control input 142, and each reward 542. The interaction trajectory 622 can then used by the learner 24 for training the controller neural network 110. The actor 22 may then perform another episode to generate another interaction trajectory 622, e.g., with different simulation parameters for the episode.
[0135] Learner:
[0136] Referring now to the learner 24. The learner 24 includes a critic neural network 610, a replay buffer 620, and an objection function 630. The learner 24 stores each interaction trajectory 622 received from the actor(s) 22 in the replay buffer 620. The learner 24 periodically samples anAttorney Docket No. 45288-0528WO1interaction trajectory 622 for an episode, e.g., at random or accordingly to an algorithm, and uses the critic neural network 610 to predict a cumulate measure of the rewards 542 at each interaction time step in the episode.
[0137] Particularly, for each interaction time step in the episode, the learner 24 prepares a critic input including: (i) the current state 512.n of the chamber 210 at the time step, and (ii) the reward 542. n for the time step. The learner 24 then processes the critic input, using the critic neural network 510, to generate a critic output (Qn) 612. n estimating a cumulative measure of the rewards 542 received from the simulation 500 at each proceeding time step. Note, the critic input can also include one or more of: the target state 132. n of the chamber 210 at the time step, the measurement 134. n of the current state 512. n of the chamber 210 at the time step, or the control input 142. n at the time step.
[0138] In general, the critic neural network 510 can have any appropriate neural network architecture that enables it to perform its described function, i.e., processing a critic input to generate a critic output. In particular, the critic neural network 510 can include any appropriate types of neural network layers (e.g., fully-connected layers, recurrent layers, convolutional layers, self-attention layers, etc.) in any appropriate numbers (e.g., 5 layers, 25 layers, or 100 layers) and connected in any appropriate configuration (e.g., as a linear sequence of layers, residual configurations, etc.). For example, the critic neural network 510 can be a feedforward neural network, a recurrent neural network, a convolutional neural network, an attention neural network, or a combination thereof. Moreover, since the critic neural network 510 is only implemented during training, the critic neural network 510 can include a large number of neural network layers, e.g., ten neural network layers or more, fifteen neural network layers or more, twenty neural network layers or more, twenty five neural network layers or more, fifty neural network layers or more, or one hundred neural network layers or more. Hence, the controller neural network 110 typically has more fewer neural network layers and / or fewer network parameters than the critic neural network 510.
[0139] The learner 24 then jointly trains the controller 110 and critic 610 neural networks on the rewards 542 for the episode using the objective function 630. The learner 610 evaluates the objective function 630 on the rewards 542 and critic outputs 612 and subsequently optimizes the objective function 630 with respect to the respective set of network parameters of each of the controller 110 and critic 620 neural networks. For example, the objective function 630 can includeAttorney Docket No. 45288-0528WO1on or more of a policy gradient objective function (aka actor’s objective), an advantage actor-critic (“A2C”) objective function, a value function approximate (aka the temporal difference (TD) error or critic’s objective), a proximal policy optimization (PPO) objective function, or a soft actor-critic (“SAC”) objective function.
[0140] The learner 24 can continuously sample interaction trajectories 622 of episodes from the replay buffer 620 to jointly train the controller 110 and critic 610 neural networks in parallel as the actor(s) 22 are generating the interaction trajectories 622.
[0141] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0142] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.
[0143] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gateAttorney Docket No. 45288-0528WO1array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0144] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0145] In this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0146] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0147] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a randomaccess memory or both. The essential elements of a computer are a central processing unit forAttorney Docket No. 45288-0528WO1performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e g., a universal serial bus (USB) flash drive, to name just a few.
[0148] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0149] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’ s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0150] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.Attorney Docket No. 45288-0528WO1
[0151] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow or JAX framework.
[0152] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0153] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0154] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a sub combination.
[0155] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed inAttorney Docket No. 45288-0528WO1the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0156] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0157] What is claimed is:
Claims
Attorney Docket No. 45288-0528WO1CLAIMS1. A method performed by one or more computers for controlling a magnetic confinement device comprising a chamber and a set of control coils configured to generate a magnetic field within the chamber in accordance with a respective control input at each of a plurality of time steps, the method comprising, at each of the plurality of time steps:obtaining a target state of the chamber defining a target, diverted configuration of a plasma within the chamber that establishes a target temperature distribution over a first wall of the chamber,wherein the target configuration comprises one or more target x-points positioned on a target separatrix for the magnetic field, each target x-point having a pair of divertor legs terminating on a respective divertor positioned within the chamber;receiving, from the magnetic confinement device, a measurement of a current state of the chamber characterizing: (i) a current configuration of the plasma within the chamber, and (ii) a current temperature distribution over the first wall of the chamber;preparing, for a controller neural network, a network input comprising: (i) the target state of the chamber, and (ii) the measurement of the current state of the chamber;processing the network input, using the controller neural network, to generate a network output defining a policy for selecting control inputs for the set of control coils at the time step;selecting the control input at the time step using the network output generated by the controller neural network at the time step; andtransmitting, to the magnetic confinement device, the control input at the time step to generate the magnetic field within the chamber in accordance therewith.
2. The method of claim 1, wherein at each of the plurality of time steps, the policy is a probability distribution over possible control inputs for the set of control coils at the time step, and the network output comprises a set of parameters of the probability distribution at the time step.
3. The method of claim 2, wherein at one or more of the plurality of time steps, selecting the control input at the time step using the network output generated by the controller neural network at the time step comprises:Attorney Docket No. 45288-0528WO1computing an expected value of the probability distribution from its set of parameters at the time step; andselecting, as the control input at the time step, the mean of the probability distribution.
4. The method of claim 2, wherein at one or more of the plurality of time steps, selecting the control input at the time step using the network output generated by the controller neural network at the time step comprises:generating the probability distribution in accordance with its set of parameters at the time step; andsampling, from the probability distribution, the control input at the time step.
5. The method of any of claims 2-4, wherein at each of the plurality of time steps, the probability distribution is a Gaussian distribution, and the set of parameters at the time step comprises a mean and covariance of the Gaussian distribution at the time step.
6. The method of any preceding claim, wherein at each of the plurality of time steps, the control input at the time step comprises a respective control voltage to be applied to each control coil at the time step.
7. The method of any preceding claim, wherein the target configuration of the plasma is a single-null diverted configuration, and the one or more target x-points is a single target x-point.
8. The method of any of claims 1-6, wherein the target configuration is a multi-null diverted configuration, and the one or more target x-points is a plurality of target x-points.
9. The method of any preceding claim, wherein at each of the plurality of time steps, the target state of the chamber at the time step comprises, for each of a plurality of configuration parameters of the target configuration of the plasma within the chamber, a respective target value of the configuration parameter at the time step.
10. The method of claim 9, wherein the plurality of configuration parameters comprises: a plasma current of the plasma, a contour of the target separatrix, a total number of the one or more target x-points, and, for each target x-point, a respective position of the target x-point on the target separatrix.Attorney Docket No. 45288-0528WO111. The method of claim 10, wherein the plurality of configuration parameters further comprises, for each divertor leg of each target x-point, a respective path of the divertor leg from the target x-point to the divertor for the target x-point.
12. The method of any of claims 10-11, wherein the plurality of configuration parameters further comprises one or more of: a position of a center of the plasma, an elongation of the plasma, a triangularity of the plasma, a radius of the plasma, or a position of a limit point of the plasma.
13. The method of any of claims 9-12, wherein at each of the plurality of time steps, the target state of the chamber at the time step further comprises, for each of a plurality of points representing the first wall of the chamber, a respective target temperature of the point at the time step.
14. The method of claim 13, wherein at each of the plurality of time steps, the respective target temperature of each of the plurality points at the time step is less than a maximum allowable temperature for the first wall.
15. The method of claim 14, wherein the maximum allowable temperature for the first wall is 1,200 Celsius (°C) or less.
16. The method of claim 15, wherein the maximum allowable temperature for the first wall is 7,000°C or less.
17. The method of any of claims 9-16, wherein at each of the plurality of time steps, the target state of the chamber at the time step further comprises, for each control coil in the set, a respective target electric current in the control coil at the time step.
18. The method of any preceding claim, wherein at each of the plurality of time steps, the measurement of the current state of the chamber at the time step comprises:a set of magnetic field measurements comprising, for each of a plurality of magnetic field probes of the magnetic confinement device, a respective measurement of the magnetic field collected from the magnetic field probe at the time step;Attorney Docket No. 45288-0528WO1a set of magnetic flux measurements comprising, for each of a plurality of magnetic flux loops of the magnetic confinement device, a respective measurement of a flux of the magnetic field collected from the magnetic flux loop at the time step;a set of electric current measurements comprising, for each control coil in the set, a respective measurement of an electric current in the control coil at the time step; anda set of temperature measurements comprising, for each of a plurality of temperature sensors of the magnetic confinement device, a respective measurement of the current temperature distribution over the first wall collected from the temperature sensor at the time step.
19. The method of any preceding claim, wherein the magnetic confinement device is a tokamak, and the chamber is a toroidal chamber having a central axis of revolution.
20. The method of claim 19, wherein the set of control coils comprises:a set of toroidal control coils encircling the toroidal chamber, the set of toroidal control coils configured to generate a toroidal field component of the magnetic field;a set of poloidal controls coils positioned about the central axis, the set of poloidal control coils configured to generate a poloidal field component of the magnetic field; anda central solenoidal control coil extending along the central axis, the central solenoidal control coil configured to induce a toroidal plasma current in the plasma.
21. The method of any preceding claim, wherein the set of control coils is a set of high-temperature superconducting control coils.
22. The method of any preceding claim, wherein each of the plurality of time steps has a length of ten milliseconds or less.
23. The method of any preceding claim, wherein the plurality of time steps has a total duration of ten seconds or more.
24. The method of any preceding claim, wherein the controller neural network has been trained on a model of the magnetic confinement device using a reinforcement learning technique.Attorney Docket No. 45288-0528WO125. The method of claim 24, wherein the model of the magnetic confinement device comprises a power supply model for the set of control coils, an evolution model for the chamber, and a sensor model for a plurality of sensors of the magnetic confinement device.
26. The method of claim 25, wherein the evolution model comprises a forward Grad-Shafranov evolutive (“FGE”) solver and a heat transport model.
27. The method of claim 26, wherein the heat transport model is a diffusive heat transport model.
28. The method of any preceding claim, wherein the magnetic confinement device generates electrical power via thermonuclear fusion of the plasma.
29. The method of any preceding claim, wherein the controller neural network is a feedforward neural network.
30. The method of any preceding claim, wherein the controller neural network comprises ten neural network layers or less.
31. A method performed by one or more computers for training a controller neural network on a model of a magnetic confinement device comprising a chamber and a set of control coils configured to generate a magnetic field within the chamber in accordance with a respective control input at each of a plurality of time steps, the method comprising:executing, over the plurality of time steps, a simulation comprising the model of the magnetic confinement device and a reward function;at each of the plurality of time steps:obtaining a target state of the chamber defining a target, diverted configuration of a plasma within the chamber that establishes a target temperature distribution over a first wall of the chamber,wherein the target configuration comprises one or more target x-points positioned on a target separatrix for the magnetic field, each target x-point having a pair of divertor legs terminating on a respective divertor positioned within the chamber;receiving, from the simulation, a model output comprising:Attorney Docket No. 45288-0528WO1a measurement of a current state of the chamber characterizing: (i) a current configuration of the plasma within the chamber, and (ii) a current temperature distribution over the first wall of the chamber; anda reward characterizing an error between: (i) the target state of the chamber, and (ii) the current state of the chamber;preparing, for the controller neural network, a network input comprising: (i) the target state of the chamber, and (ii) the measurement of the current state of chamber;processing the network input, using the controller neural network, to generate a network output defining a policy for selecting control inputs for the set of control coils at the time step;selecting the control input at the time step using the network output generated by the controller neural network at the time step; andtransmitting, to the simulation, the control input at the time step to generate the magnetic field within the chamber in accordance therewith; andtraining the controller neural network on the rewards using a reinforcement learning technique.
32. The method of claim 31, further comprising:controlling the magnetic confinement device using the controller neural network trained on the model of the magnetic confinement device.
33. The method of any of claims 31-32, wherein:the model of the magnetic confinement device comprises a power supply model for the set of control coils, an evolution model for the chamber, and a sensor model for a plurality of sensors of the magnetic confinement device, andexecuting, over the plurality of time steps, the simulation comprising the model of the magnetic confinement device and the reward function comprises, at each of the plurality of time steps:receiving a model input comprising: (i) a current state of the chamber at a preceding time step, and (ii) a control input for the set of control coils at the preceding time step;processing the model input, using the power supply and evolution models, to generate the current state of the chamber at the time step;Attorney Docket No. 45288-0528WO1processing the current state of the chamber at the time step, using the sensor model, to generate the measurement of the current state of the chamber at the time step; and processing the target and current states of the chamber at the time step, using the reward function, to generate the reward for the time step.
34. The method of any of claims 31-33, wherein executing, over the plurality of time steps, the simulation comprising the model of the magnetic confinement device and the reward function comprises, at each of the plurality of time steps:determining, from the current state of the chamber at the time step, whether a physical feasibility constraint of the magnetic confinement device is violated at the time step; andif a physical feasibility constraint of the magnetic confinement device is violated at the time step, terminating the simulation at the time step.
35. The method of claim 34, wherein at each of the plurality of time steps, determining, from the current state of the chamber at the time step, whether the physical feasibility constraint of the magnetic confinement device is violated at the time step comprises one or more of:determining whether a density of the plasma satisfies a threshold at the time step; determining whether a plasma current of the plasma satisfies a threshold at the time step; determining whether a plasma safety factor of the plasma satisfies a threshold at the time step;determining whether a respective current through each of one or more control coils in the set satisfies a threshold at the time step; ordetermining whether the current temperature distribution over the first wall satisfies a threshold at the time step.
36. The method of any of claims 31-35, wherein at each of the plurality of time steps:the target state of the chamber at the time step comprises, for each of a plurality of configuration parameters of a configuration of the plasma within the chamber, a respective target value of the configuration parameter at the time step,the current state of the chamber at the time step comprises, for each of the plurality of configuration parameters, a respective current value of the configuration parameter at the time step, andAttorney Docket No. 45288-0528WO1the reward for the time step comprises, for each of the plurality of configuration parameters, a respective error between: (i) the target value of the configuration parameter at the time step, and (ii) the current value of the configuration parameter at the time step.
37. The method of claim 36, wherein at each of the plurality of time steps, the reward for the time step comprises a weighted linear combination of the respective errors for each of the plurality of configuration parameters at the time step.
38. The method of any of claims 36-37, wherein the plurality of configuration parameters comprises: a plasma current of the plasma, a contour of a separatrix, a total number of x-points, and, for each x-point, a respective position of the x-point.
39. The method of claim 38, wherein the plurality of configuration parameters further comprises one or more of: a position of a center of the plasma, an elongation of the plasma, a triangularity of the plasma, a radius of the plasma, or a position of a limit point of the plasma.
40. The method of any of claims 36-39, wherein at each of the plurality of time steps:the target state of the chamber at the time step further comprises, for each of a plurality of points representing the first wall of the chamber, a respective target temperature of the point at the time step,the current state of the chamber at the time step further comprises, for each of the plurality of points, a respective current temperature of the point at the time step, andthe reward for the time step further comprises, for each of the plurality of points, a respective error between: (i) the target temperature of the point at the time step, and (ii) the current temperature of the point at the time step.
41. The method of any of claims 36-40, wherein at one or more of the plurality of time steps, the reward for the time step further comprises, for each target x-point:a gradient of a flux of the magnetic field at the target x-point;a difference between: (i) the flux of the magnetic field at the target x-point, and (ii) the flux of the magnetic field at a current separatrix of the magnetic field; andAttorney Docket No. 45288-0528WO1for each divertor leg of the target x-point, a difference between: (i) the flux of the magnetic field along a path of the divertor leg, and (ii) the flux of the magnetic field at the current separatrix.
42. The method of claim 41, wherein at one or more of the plurality of time steps, the reward for the time step further comprises, for each current x-point of the magnetic field, a shortest distance of the current x-point to the current separatrix.
43. The method of any of claims 31-42, wherein the reinforcement learning technique is an actor-critic reinforcement learning technique, and training the controller neural network on the rewards using the actor-critic reinforcement learning technique comprises jointly training the controller neural network and a critic neural network on the rewards.
44. The method of claim 43, further comprising, for each of the plurality of time steps:preparing, for the critic neural network, a critic input comprising: (i) the current state of the chamber at the time step, and (ii) the reward for the time step; andprocessing the critic input, using the critic neural network, to generate a critic output estimating a cumulative measure of the rewards received from the simulation at each proceeding time step.
45. The method of claim 44, wherein for each of the plurality of time steps, the critic input at the time step further comprises one or more of: the target state of the chamber at the time step, the measurement of the current state of the chamber at the time step, or the control input at the time step.
46. The method of any of claims 44-45, wherein jointly training the controller and critic neural networks on the rewards comprises:generating an objective function that depends on the rewards and critic outputs; and optimizing the objective function with respect to a respective set of network parameters of each of the controller and critic neural networks.
47. The method of any of claims 43-46, wherein the actor-critic reinforcement learning technique is a maximum a posteriori policy optimization (“MPO”) technique.Attorney Docket No. 45288-0528WO148. The method of any of claims 43-47, wherein the actor-critic reinforcement learning technique is a distributed actor-critic reinforcement learning technique.
49. The method of any of claims 43-48, wherein the controller neural network has fewer network parameters than the critic neural network.
50. The method of any of claims 43-49, wherein the controller neural network is a feedforward neural network, and the critic neural network is a recurrent neural network.
51. The method of any of claims 31-50, further comprising after training the controller neural network, using the controller neural network to control the real-world magnetic confinement device by, at each of a plurality of time steps:processing the network input comprising: (i) the target state of the real-world chamber, and (ii) the measurement of the current state of the real-world chamber, using the controller neural network, to generate the network output, and transmitting, to the real-world magnetic confinement device, a control input for the set of control coils selected using the network output.
52. The method of any of claims 1-30, wherein the controller neural network has been trained by the method of any of claims 21-50.
53. A system, comprising:one or more computers; andone or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the method of any of claims 1-52.
54. One or more non-transitory computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any of claims 1-52.
Citation Information
Patent Citations
US20240312657A1