Plasma posture control agent strategy model construction method and device, and medium

By constructing a simulation environment in a tokamak device and using a reinforcement learning agent model to generate current commands, the problem of traditional PID controllers dealing with transient disturbances in nonlinear and complex dynamic environments is solved, thereby improving the stability and safety of the tokamak device.

CN120762290BActive Publication Date: 2025-11-07HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511280083.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-07
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional PID controllers struggle to handle large transient disturbances in the nonlinear and complex dynamic environments of tokamak devices, resulting in poor device stability and safety.

Method used

A tokamak simulation environment is constructed to identify the environmental configuration parameters of runaway scenarios. Current commands are generated through a reinforcement learning agent model, and the model is trained in conjunction with a reward function to update the reinforcement learning agent model to generate strategies to cope with transient disturbances.

Benefits of technology

It improves the stability and safety of tokamak devices in nonlinear and complex dynamic environments, effectively copes with large transient disturbances, and enhances the robustness and adaptability of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762290B_ABST
    Figure CN120762290B_ABST
Patent Text Reader

Abstract

The application discloses a kind of plasma posture control agent strategy model construction method, device and medium, by constructing tokamak simulation environment, the environmental configuration parameter of out-of-control scene under the control of PID controller is identified, and the training environment of configuration is obtained;Using reinforcement learning agent model to learn, generate current command;The state of control point at each time is calculated based on tokamak simulation environment;The action command obtained by the control point state at each time, current command and PID controller are input into reward function, and the environment reward is calculated;According to the environment reward, the control point state at each time and current command, reinforcement learning training is carried out, and the command strategy is updated;When reinforcement learning agent model meets convergence condition, output agent strategy model.The application scheme provides a kind of strategy model in nonlinear and complex dynamic environment to deal with the ability of large disturbance of transient state, guarantees the stability and safety of tokamak device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of fusion plasma control, and in particular to a plasma shape control agent strategy model construction method and device and medium. BACKGROUND

[0002] A tokamak device is an important equipment for researching and realizing nuclear fusion reaction, and its working principle is to constrain high-temperature plasma through a strong magnetic field to make it perform nuclear fusion reaction in a high-temperature and high-pressure state. The control of the plasma shape in the tokamak device is crucial for realizing a stable plasma state and effective nuclear fusion reaction, and it determines whether the tokamak discharge process can be stably operated.

[0003] In the prior art, the control of the plasma shape usually adopts a PID (proportional-integral-derivative) controller. The PID controller generates a corresponding current control command by measuring the difference between the current shape parameter and the target shape parameter, so as to make the shape parameter approach the target value. However, the traditional PID controller has insufficient ability to cope with large transient disturbances in a nonlinear and complex dynamic environment, and the stability and safety of the tokamak device are poor. SUMMARY

[0004] In order to solve the above problems, the present application provides a plasma shape control agent strategy model construction method and device and medium, which can provide a strategy model with the ability to cope with large transient disturbances in a nonlinear and complex dynamic environment, and guarantee the stability and safety of the tokamak device.

[0005] The present application provides a plasma shape control agent strategy model construction method, which comprises the following steps:

[0006] A tokamak simulation environment is constructed, and the environment configuration parameters of the out-of-control scene under the control of the PID controller are identified;

[0007] The environment configuration parameters are configured into the tokamak simulation environment to obtain a configured training environment;

[0008] A preset reinforcement learning agent model is used to learn based on the training environment to generate a current command;

[0009] The current command is input into the tokamak simulation environment, and the control point state at each time is calculated;

[0010] The control point state at each time, the current command and the action command obtained by the PID controller are input into a preset reward function to calculate an environment reward;

[0011] Reinforcement learning training is performed based on the environmental rewards, the control point states at each time point, and the current commands to update the command strategy generated by the reinforcement learning agent model.

[0012] When the reinforcement learning agent model meets the preset convergence condition, the latest reinforcement learning agent model is output as the agent policy model.

[0013] Preferably, the construction of the tokamak simulation environment and the identification of environmental configuration parameters for runaway scenarios under PID controller control specifically include:

[0014] Constructing a tokamak simulation environment;

[0015] Randomly initialize the environment configuration parameters in the tokamak simulation environment to generate different simulation environments;

[0016] The PID controller is used to control each simulation environment;

[0017] When the controlled parameters meet the preset conditions, the environmental configuration parameters of the corresponding runaway scenario are identified.

[0018] Preferably, the tokamak simulation environment includes a power response model with configured parameters, a disturbance scenario model, and a plasma response model.

[0019] Preferably, the reward function includes ;

[0020] in, For the environmental reward at time t, These are the weighting coefficients for each item. The state of the control point at time t The control point error below, The current command output by the reinforcement learning policy at time t. The action command output by the PID controller at time t. The control point state from the start of the round to the current time t. The normalized error integral.

[0021] Preferably, the method further includes:

[0022] During training, monitor the state of the control point at time t. Control point error The amount;

[0023] The state of the control point at time t is monitored. Control point error When the rate of increase of a component reaches a preset trend threshold, a preset adjustment model is used to reduce the weight coefficient. The value;

[0024] wherein the adjustment model is , a weight coefficient of time t a value of a preset initial value a preset coefficient value.

[0025] Preferably, the reinforcement learning training according to the environment reward, the control point state of each time and the current command updates the command strategy generated by the reinforcement learning agent model, comprising:

[0026] initializing training parameters;

[0027] re-randomly configuring model parameters in the tokamak simulation environment at the beginning of each round;

[0028] interacting in the currently configured tokamak simulation environment by the reinforcement learning agent, outputting a current command at each step, the currently configured tokamak simulation environment updating the control point state according to the output current command and feeding back the environment reward and the new control point state;

[0029] after reaching a preset number of cycle update steps, collecting the current command, the control point state and the environment reward, and updating the command strategy of the reinforcement learning agent model based on a preset reinforcement learning algorithm.

[0030] Preferably, the method further comprises:

[0031] acquiring real-time collected environment parameters into the agent strategy model, outputting command values of each controlled object according to the agent strategy model, and performing plasma shape control.

[0032] The embodiment of the application also provides a plasma shape control agent strategy model construction device, the device comprising:

[0033] an environment construction module, configured to construct a tokamak simulation environment and identify environment configuration parameters of an out-of-control scene under PID controller control;

[0034] a parameter configuration module, configured to configure the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment;

[0035] a learning module, configured to learn based on the training environment by using a preset reinforcement learning agent model to generate a current command;

[0036] a state module, configured to input the current command into the tokamak simulation environment to calculate a control point state at each time;

[0037] a reward module configured to input the control point state at each time, the current command and the action command obtained by the PID controller into a preset reward function, and calculate an environment reward;

[0038] a training module configured to perform reinforcement learning training according to the environment reward, the control point state at each time and the current command, and update the command strategy generated by the reinforcement learning agent model;

[0039] an output module configured to output the latest reinforcement learning agent model as an agent strategy model when the reinforcement learning agent model meets a preset convergence condition.

[0040] Another embodiment of the present application provides a plasma configuration control agent strategy model construction device, comprising a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the plasma configuration control agent strategy model construction method of any one of the above embodiments when executing the computer program.

[0041] Another embodiment of the present application provides a computer readable storage medium comprising a stored computer program, wherein the computer readable storage medium is controlled to execute the plasma configuration control agent strategy model construction method of any one of the above embodiments when the computer program is running.

[0042] The present application provides a plasma configuration control agent strategy model construction method, device and medium, which identifies the environment configuration parameters of the out-of-control scene under the control of the PID controller by constructing a tokamak simulation environment, configures the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment, learns based on the training environment using a preset reinforcement learning agent model to generate a current command, inputs the current command into the tokamak simulation environment to calculate the control point state at each time, inputs the control point state at each time, the current command and the action command obtained by the PID controller into a preset reward function to calculate an environment reward, performs reinforcement learning training according to the environment reward, the control point state at each time and the current command, and updates the command strategy generated by the reinforcement learning agent model, and outputs the latest reinforcement learning agent model as an agent strategy model when the reinforcement learning agent model meets a preset convergence condition. The present application can provide a strategy model with the ability to cope with large transient disturbances in a nonlinear and complex dynamic environment, and guarantee the stability and safety of the tokamak device. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 is a flowchart of a plasma configuration control agent strategy model construction method provided by an embodiment of the present application;

[0044] Figure 2 is another flowchart of a plasma shape control agent strategy model construction method provided by an embodiment of the present application;

[0045] Figure 3 is a structural schematic diagram of a control framework of a plasma shape control agent strategy model construction method provided by an embodiment of the present application;

[0046] Figure 4 is a control effect diagram of a traditional PID controller;

[0047] Figure 5 is a control effect diagram of a plasma shape control agent strategy model construction method provided by an embodiment of the present application;

[0048] Figure 6 is a training result diagram of four groups of agent strategy training provided by an embodiment of the present application;

[0049] Figure 7 is a result schematic diagram of control effect and control command of an EAST Tokamak device provided by an embodiment of the present application;

[0050] Figure 8 is a structural schematic diagram of a plasma shape control agent strategy model construction device provided by an embodiment of the present application;

[0051] Figure 9 is another structural schematic diagram of a plasma shape control agent strategy model construction device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0052] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0053] The present application provides a plasma shape control agent strategy model construction method, referring to Figure 1 is a flowchart of a plasma shape control agent strategy model construction method provided by an embodiment of the present application, and the method comprises steps S1-S7:

[0054] Step S1, constructing a Tokamak simulation environment, and identifying environment configuration parameters of an out-of-control scene under PID controller control;

[0055] Step S2, configure the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment;

[0056] Step S3, learning based on the training environment using a preset reinforcement learning agent model to generate a current command;

[0057] Step S4, inputting the current command into the tokamak simulation environment to calculate the control point state at each time;

[0058] Step S5, inputting the control point state at each time, the current command and the action command obtained by the PID controller into a preset reward function to calculate an environment reward;

[0059] Step S6, performing reinforcement learning training according to the environment reward, the control point state at each time and the current command to update the command strategy generated by the reinforcement learning agent model;

[0060] Step S7, outputting the latest reinforcement learning agent model as an agent strategy model when the reinforcement learning agent model meets a preset convergence condition.

[0061] In the specific implementation of the embodiment, the environment of the present application adopts a plasma response model based on physics, then interacts with the response environment through reinforcement learning, and places a traditional control strategy in the environment to collect data for imitation learning in the early stage of training. Referring to Figure 2 is another flowchart of the plasma configuration control agent strategy model construction method provided by the embodiment of the present application. The traditional control strategy is difficult to cover some scenes with large transient disturbance, and these scenes are included in the training environment in the training to enable the reinforcement learning agent to learn an effective path to complete the control task with large transient disturbance. In order to enable the reinforcement learning agent to have a faster convergence speed in the early stage of training, the idea of imitation learning is fused, the output commands of the traditional control strategy and the output commands of the reinforcement learning agent are normalized and compared, and then the comparison is included in the calculation of the reward function, so that the reinforcement learning agent has the guidance of the traditional control strategy in the early stage of training, greatly improving the training efficiency and convergence speed, and saving the training cost. At the same time, through the exploratory nature of reinforcement learning itself, the control ability to cope with scenes with large transient disturbance is explored.

[0062] The specific steps when the scheme is implemented include:

[0063] Model construction and environment preparation, building a plasma response model based on physics, and constructing a tokamak simulation environment.

[0064] Traditional control strategy data collection, for example, by generating diversified simulation environments, the environment configuration parameters corresponding to the out-of-control scenes of the traditional control strategy are recorded.

[0065] In the initialization of each round of reinforcement learning training, the recorded runaway scenario environment configuration parameters are randomly configured into the tokamak simulation environment, forming a training environment of diversified discharge scenarios.

[0066] The current command is input into the tokamak simulation environment, and the control point state at each time is calculated.

[0067] It should be noted that the control point state can include control point error and integral error in this embodiment, and other errors can be used to represent the control point state in other embodiments.

[0068] Imitation learning fusion training, the reinforcement learning agent (Agent) interacts with the configured tokamak simulation environment, and obtains the current command a. The traditional control strategy obtains the traditional strategy command a* in the same environment. Normalize and compare a and a*. The normalized comparison result is included in the reward calculation of the reward model, so that the reinforcement learning agent can get the guidance of the traditional control strategy in the early training.

[0069] Reinforcement learning training and optimization, the reinforcement learning agent updates the strategy according to the reward r and the control point state s fed back by the environment, and combines the reinforcement learning update algorithm.

[0070] Continuously repeat the above interaction, comparison, reward calculation and strategy update process, use the exploratory nature of reinforcement learning to let the agent learn the control ability to deal with large transient disturbance scenes, until the training converges.

[0071] Referring to Figure 3 It is a control framework structure schematic diagram of a plasma shape control agent strategy model construction method provided by an embodiment of the application. The red box part in the figure is the shape controller, that is, the agent strategy model designed by the application. The agent obtained by imitating reinforcement learning training will accept real-time shape data, real-time shape inversion program calculation results, real-time control command calculation and complete the control task.

[0072] The application scheme fuses the control method of imitation learning and reinforcement learning. By introducing expert demonstration data in the training process, the reinforcement learning agent can have a certain strategy basis at the initial stage, thereby reducing the dependence on the sensitivity of the hyperparameters and accelerating the convergence of the strategy. At the same time, by setting a specific training environment, the agent can face the high-error scene that is difficult to handle by traditional methods in the training stage, and realize the autonomous learning of the robust strategy. In order to further improve the high-dimensional state processing capability, the application adopts a deep neural network for feature extraction, so that the controller has the representation and modeling capability for the complex system state, thereby improving the overall performance of the control performance. The application scheme can provide a strategy model capable of coping with large transient disturbance in a nonlinear and complex dynamic environment, thereby guaranteeing the stability and safety of the tokamak device.

[0073] In another embodiment provided by the application, the step S1 specifically comprises:

[0074] A tokamak simulation environment is constructed.

[0075] The environment configuration parameters in the tokamak simulation environment are randomly initialized to generate different simulation environments.

[0076] The PID controller is used to control each simulation environment.

[0077] When the controlled parameters meet the preset conditions, the environment configuration parameters corresponding to the out-of-control scene are identified.

[0078] In the specific implementation of the embodiment, the tokamak is a ring-shaped container for realizing controlled nuclear fusion by magnetic confinement, and the tokamak simulation environment is constructed to reproduce the running state of the tokamak device in the computer simulation.

[0079] On the basis of the constructed tokamak simulation environment, the environment configuration parameters are randomly initialized, such as the power supply parameters in the power supply response model, the disturbance amplitude and frequency in the disturbance scene model, and the like. In this way, a large number of simulation environments under different working conditions can be generated to simulate the running state of the tokamak under various conditions, thereby covering as many actual running scenes as possible.

[0080] The PID controller is used for control: the PID controller is a commonly used closed-loop automatic control device, which adjusts the control amount according to the error of the system by adjusting the proportional, integral and differential parameters. In the tokamak simulation environment, the PID controller is used to control different simulation environments, expecting to make the running state of the tokamak stable or meet a specific control target.

[0081] The preset condition is set according to the safety, stability and other requirements of the tokamak operation, for example, the temperature, density, current and other parameters of the plasma exceed the normal range, or the abnormal fluctuation of the equipment operation state occurs. When the PID controller controls the simulation environment, if some parameters in the environment, such as the plasma parameter operation parameter, meet the preset out-of-control condition, the corresponding simulation environment configuration parameter at this time can be determined, that is, the environment configuration parameter that leads to the out-of-control scene under the control of the PID controller is identified.

[0082] By randomly initializing the environment configuration parameters and identifying the out-of-control scene, the performance of the tokamak under different operating conditions can be comprehensively explored, and potential risk scenarios that may lead to out-of-control under the control of the PID controller can be found.

[0083] In the reinforcement learning training, a response simulation model needs to be configured for data interaction. The input signal of the simulation environment should be the output command of the agent policy, and the output signal of the simulation environment should also be fed back to the reinforcement learning algorithm for iterative updating of the policy. The simulation environment model is used to simulate the evolution process of the plasma configuration, which can be obtained by a deep learning-based data fitting method or by a physics-based mathematical design method. The environment configuration needs to be able to generate environment disturbances with parameters, which helps the reinforcement learning agent to explore in the environment and strengthen the robustness of the agent. A typical case is to randomly generate target quantities in the environment during training, so that the agent can be trained in a random target scenario.

[0084] The reinforcement learning training update algorithm can use mainstream reinforcement learning algorithms such as PPO, TD3, DDPG, SAC, etc. The algorithm needs to be able to support parallel reinforcement learning training to save training costs.

[0085] In another embodiment provided by the present application, the tokamak simulation environment includes a power supply response model, a disturbance scene model and a plasma response model of the configuration parameters.

[0086] Specifically, the power supply response, plasma response and various random disturbance scenes involved in the tokamak simulation environment are established by establishing corresponding physical models such as power supply response models and plasma response models, and the parameters of these models are configurable, providing a basis for subsequent simulation of different operating conditions.

[0087] The constructed tokamak simulation environment includes a power supply response model, a disturbance scene model and a plasma response model of the configuration parameters.

[0088] The training environment requires high-performance CPUs and GPUs. The CPU is primarily used for computation within the simulation environment and for data interaction with the GPU, while the GPU handles matrix operations such as parameter inference and updates for the neural network. The reinforcement learning training algorithm framework can utilize mainstream platforms that support parallel environment generation and execution, such as PyTorch, TensorFlow, and MATLAB, which will facilitate parallel training of reinforcement learning. The simulation environment platform will depend on the simulation model in the project; generally, the platform with the shortest physical time required for a single step of computation is selected. For example, C++, Python, or MATLAB can be used to build the simulation model.

[0089] In another embodiment provided by the present invention, the reward function includes ;

[0090] in, For the environmental reward at time t, These are the weighting coefficients for each item. The state of the control point at time t The control point error below, For the current command output by the reinforcement learning policy at time t, The action command output by the PID controller at time t. Let be the normalized integral error from the start of the round to the current time t.

[0091] In this specific implementation, the design of the reward function is crucial in the reinforcement learning training task, as it determines the agent's training direction and efficiency. This function consists of multiple parts. To ensure the training objective, namely, precise control of the plasma configuration, a reward value is calculated after normalizing the errors at each control point. To improve the convergence efficiency of training, three additional reward forms are introduced to guide the exploration of the strategy, as follows:

[0092] The action penalty item is designed such that the larger the absolute value of the control command, the greater the penalty. The purpose is to enable the agent to complete the control task with as few current commands as possible.

[0093] The error integral item after normalization, the smaller the error integral of the design control point is, the higher the reward value is, due to the mathematical characteristics of the integral value of the error item, which is a small value at the beginning of each round. The normalized error integral item will gradually increase with the increase of the number of steps in each round, so that the reward value calculated by the value will be more smooth, and the integral value will not change greatly due to the occasional behavior of the agent. Incorporating this value into the reward function will help the reinforcement learning agent understand whether the overall trend of the current policy behavior meets the expectations, and will not be affected by the instantaneous calculation deviation caused by the exploration noise. It is helpful to guide the agent to explore the correct direction of the overall control trend.

[0094] The traditional control strategy command comparison item, the traditional control strategy, that is, the PID control strategy based on multiple input and multiple output, is also included in the calculation of the reward function at the beginning of training. The state of the environment is fed back to the traditional control strategy and the reinforcement learning agent strategy during training, and the two calculate the control commands respectively after normalization comparison. The smaller the difference between the two commands is, the higher the reward is. The purpose is to guide the reinforcement learning agent at the beginning of training, which can greatly improve the convergence speed of the reinforcement learning agent at the beginning of training.

[0095] Therefore, the reward function obtained is:

[0096] ;

[0097] Wherein, is the environment reward at time t, are the weight coefficients of each item, is the control point error under the control point state at time t , that is, the target reward value, which can be used to measure the control effect. is the current command output by the reinforcement learning strategy at time t, is the action command output by the PID controller at time t, is the normalized integral error from the start of the round to the current time t.

[0098] In another embodiment provided by the application, the method further comprises:

[0099] During the training process, the component of the control point error under the control point state at time t is monitored;

[0100] When the rate of rise of the component of the control point error under the control point state at time t reaches a preset trend threshold, the value of the weight coefficient is reduced by using a preset adjustment model;

[0101] ​​The adjustment model is , The weight coefficient of time t is The value of is a preset initial value, is a preset coefficient value.

[0102] In the embodiment, whether the control trend is correct is mainly judged by observing the control error contribution in the reward function, that is, the component in the reward function. The round cumulative amount calculated by the component is a gold index for measuring the control task, and the meaning represented by the value is shown in the symbol description.

[0103] Whether the control trend is correct is determined by observing whether there is a significant upward trend in the amount. After it is determined that there is a significant upward trend in the amount, the reinforcement learning agent can be used to interact with the environment in the simulation environment for multiple rounds (which can be a scene with random environment parameters), and the control effect of each controlled object and the steady-state error and other indicators are observed through the simulation results to further determine whether the control trend is correct.

[0104] Generally, if each control object can effectively track each control target in each simulation round, it indicates that the control trend is correct. Further, the weight of the component in the reward function can be reduced. This component is a mimic training component for guiding the reinforcement learning agent to mimic the traditional control strategy. When the control trend is correct, the weight of the component can be reduced to reduce the weight contribution of the mimic training component to the overall reward function. As the training progresses, when the overall control trend of the reinforcement learning agent is correct, the weight of the component can be gradually reduced

[0105] to allow the agent to explore and complete the control task by itself. When the rate of increase of the component is monitored to reach a preset trend threshold, the value of the component is reduced by using a preset adjustment model.

[0106] The adjustment model is , The value of

[0107] of time t is , the preset initial value is , and the preset coefficient value is . The dependence on the traditional controller is gradually reduced, and this form is a typical example, and the specific weight and calculation details can be flexibly adjusted according to different experimental environments.

[0108] The dependence on the traditional controller is gradually reduced, and this form is a typical example, and the specific weight and calculation details can be flexibly adjusted according to different experimental environments.

[0109] ​​In a further embodiment provided by the present application, the step S5 comprises:

[0110] initializing the training parameters;

[0111] re-randomly configuring the model parameters in the tokamak simulation environment at the beginning of each round;

[0112] the reinforcement learning agent interacts with the currently configured tokamak simulation environment, outputs a current command at each step, the currently configured tokamak simulation environment updates the control point state according to the output current command, and feeds back the environment reward and the new control point state;

[0113] after reaching the preset number of cycle updates, the current command, the control point state and the environment reward are collected, and the command strategy of the reinforcement learning agent model is updated based on the preset reinforcement learning algorithm.

[0114] In the implementation of the present embodiment, the reinforcement learning agent interacts with the environment for M rounds in each training of the reinforcement learning training, and the number of rounds is represented by the horizontal coordinate of the meaning of Figure 6 Each round means that the simulation environment runs a complete control task, and the environment parameters need to be randomly configured at the beginning of each training, but the reinforcement learning agent is inherited. The reinforcement learning agent interacts with the environment for N steps in each round, and the control number N is fixed.

[0115] After the reinforcement learning agent samples N iter steps, the sampled data (s, a, r) are collected, and the strategy of the reinforcement learning agent is updated through the reinforcement learning algorithm. N iter steps is an algorithm hyperparameter, which is determined according to experience. The number of times of cycle updating is M*N / N iter The updating process is repeated until the reinforcement learning agent interacts with the environment for M rounds, and the training is finally ended. The value of M is generally determined by experience.

[0116] Specifically, the number of rounds M required for training and the control number N of the reinforcement learning agent interacting with the environment in each round, i.e., the fixed number of control required to complete the control task, are determined, and the algorithm hyperparameter N iter is determined, and the sampled data is collected and the strategy is updated after the reinforcement learning agent samples N iter steps.

[0117] initializing the reinforcement learning agent.

[0118] environment configuration in each round, at the beginning of each round, the parameters in the tokamak simulation environment, such as the power supply response model, the disturbance scene model and the plasma response model, are randomly configured.

[0119] Step interaction, the reinforcement learning agent interacts in the current configured environment, outputs current command a at each step, the environment updates the state according to a and feeds back reward r and new control point state s.

[0120] Data collection and policy update: every N iter steps, collect the N iter step sampling data (s, a, r) and update the policy of the reinforcement learning agent based on these data using reinforcement learning algorithms (such as DDPG, PPO, etc.).

[0121] Inheritance agent: when entering the next round, the reinforcement learning agent inherits the updated policy of the last round, and the environment parameters are randomly configured again.

[0122] When the training of M rounds is completed, the reinforcement learning training cycle is ended. The number of cycle updates of the whole training process is M*N / N iter .

[0123] After reaching the preset cycle update step, the current command, control point state and environment reward are collected, and the command policy of the reinforcement learning agent model is updated based on the preset reinforcement learning algorithm.

[0124] In the initial attempt of reinforcement learning training, different groups of hyperparameter configuration schemes with different spans can be selected, and the training results of different groups of schemes can be compared to narrow the exploration range of hyperparameters. The number of training rounds of a single task should be designed as needed, and in theory, the overall situation of training should be observed according to the cumulative reward value of training rounds. When the cumulative reward of rounds gradually increases and shows a slow growth rate, it may mean that the policy agent has learned a bottleneck, and the design of the reward function can be modified to continue training to obtain a better performance agent. The specific example may vary, which mainly depends on the design of the reward function and the design of the hyperparameters, and requires multiple attempts to gain experience.

[0125] In another embodiment provided by the application, the method further comprises:

[0126] Acquiring the real-time collected environment parameters into the agent policy model, outputting the command values of each controlled object according to the agent policy model, and performing plasma posture control.

[0127] In the specific implementation of this embodiment, when the training is completed, the robustness and disturbance handling ability of the agent need to be tested in the simulation environment, and the control results in the traditional control scheme of imitation learning can be compared. The results should be at least not worse than the traditional control scheme.

[0128] After the simulation test is completed, the agent control strategy is deployed into the control system of the real device to replace the traditional controller and realize the operation test of the real device.

[0129] The scheme of the application introduces an imitation reinforcement learning algorithm, and a reinforcement learning algorithm of imitation learning is introduced into the tokamak plasma shape feedback control to replace a traditional PID controller.

[0130] Through the design of the reward function guide term, in order to realize the shape control task of the tokamak plasma, in addition to designing the purpose index reward value of the task, three types of reward values are designed to guide the learning direction of the reinforcement learning agent, which greatly improves the efficiency of the reinforcement learning and reduces the time cost.

[0131] Through reinforcement learning training, the agent can cope with scenarios that the traditional PID controller cannot cope with, such as scenarios with large transient errors, and embodies more accurate and robust control performance.

[0132] In another embodiment provided by the application, a plasma shape control agent strategy model construction scheme specific implementation process and effect are provided.

[0133] First, a simulation environment is built, and a physical-based plasma equilibrium perturbation model is used as the simulation response model.

[0134] The training environment is configured, and four different reward hyperparameter designs are used.

[0135] Among them, int weight is the weight value of the normalized error integral term in the reward function, the first group represents that the term is not used, and the subsequent three groups use the term with gradually increasing weights, and imitate represents that the imitation learning idea is used, and a traditional control strategy term is used in the reward function.

[0136] The training algorithm selection is performed, the PPO algorithm is used for strategy updating in the training, the algorithm construction platform is a matlab calculation platform, the training environment is a matlab platform environment, and the strategy fitting paradigm is an mlp neural network paradigm.

[0137] The simulation environment is constructed, the simulation response model adopts a physical-based plasma equilibrium disturbance model, the model is written in matlab language, and is implemented on a matlab platform.

[0138] When the simulation environment is compared and tested, the control effects of the strategy trained by the reinforcement learning and the traditional PID strategy are compared, and the comparison environment is a matlab simulation environment.

[0139] Referring to Figure 4 is a control effect diagram of a traditional PID controller, and referring to Figure 5 is a control effect diagram of a plasma shape control agent strategy model construction method provided by the embodiment of the application.

[0140] The blue line represents the response of the controlled physical quantity in the simulation environment, and the red line represents the corresponding control target, that is, Gap out, Gap top and Gap in represent the response of the magnetic flux values of the three control points, respectively, and Xpointbr and Xpointbz represent the magnetic field strengths of the X point of the tokamak plasma in the R axis direction and the Z axis direction, respectively, and the five quantities are the controlled objects in the shape control. Gap outTarget, Gap topTarget and Gap inTarget represent the control targets of the magnetic flux values of the three control points, respectively, XpointbrTarget and XpointbzTarget represent the control targets of the magnetic field strengths of the X point of the tokamak plasma in the R axis direction and the Z axis direction, respectively, Psi_error represents the magnetic flux error, and the unit is weber Wb, and Bx_error represents the magnetic field error, and the unit is tesla T.

[0141] The purpose of the control strategy is to control the controlled physical quantity to the control target, that is, the blue line should finally coincide with the red line. It can be seen from the two figures that the traditional controller cannot effectively control when responding to the disturbance, and the control object has deviated from the control target. The control strategy based on reinforcement learning can well respond to this situation and effectively complete the control, so that the blue line coincides with the red line.

[0142] When the reward function hyperparameters are compared, four different agent strategies are trained according to the parameter configuration of the second term, which are used to compare the effects of the design of each term of the reward function. The first group does not add the normalized error integral term in the reward function design and the traditional control strategy term. The second, third and fourth groups all add the normalized error integral term, and the weight gradually increases. Only the fourth group additionally adds the traditional control strategy term. Except for these hyperparameters, the other hyperparameters of each group are consistent to facilitate comparison.

[0143] Referring to Figure 6 is a training result diagram of the training of the four groups of agent strategies provided by the embodiment of the application, wherein the horizontal axis Episode is the number of training rounds, and the vertical axis is the target reward value, that is, the reward value composed of the control error term. Since the purpose is to control the target well, the target reward value can well evaluate the convergence and convergence speed of the training.

[0144] It can be seen intuitively that the score and convergence speed of the yellow line agent of the third group with the weight configuration of “int weight = 1.5” are better than those of the first group and the second group, that is, the results of “int weight = 0” and “int weight = 0.5”. The purple line of the fourth group, which is trained by combining the idea of imitation learning, is much better than the first three, which shows the superiority of the reward function design of the application.

[0145] Through the real experimental results of the embodiment, the control effect of the controller is realized and tested on the EAST tokamak device. Referring to Figure 7 is a result diagram of the control effect and control command of the EAST tokamak device provided by the embodiment of the application. The first two subgraphs are the corresponding controlled objects, and the control target in the experiment is 0. The third subgraph is the control command value, that is, the output value of the agent strategy. ipf is the current, and ipf1-ipf12 represent the currents of different current channels or coils.

[0146] The steady-state error of each controlled object is less than 2e-3, and a 10s long pulse discharge is completed, which shows that the agent can still meet the control requirements in actual experiments. It is also proved that the embodiment can realize the task transition from the simulation environment to the real environment, that is, the EAST experimental environment.

[0147] The embodiment shows the various results of the plasma shape controller obtained by the embodiment of the application, compares the control effects of the embodiment and the traditional PID controller, and shows the superiority of the design scheme in the embodiment that combines the idea of imitation learning compared with the traditional reinforcement learning training scheme. Finally, through the actual EAST experiment case, it is shown that the embodiment has the ability to transition from the simulation environment to the real world environment. It is shown that the embodiment has more robust and more accurate control effect than the traditional PID control.

[0148] The application provides a tokamak plasma shape control scheme based on imitative reinforcement learning, compared with the prior art:

[0149] The robustness and adaptability of the control strategy can be improved: by introducing a self-learning control strategy based on deep reinforcement learning, the problem of easy loss of control of a traditional multiple-input multiple-output PID controller in a strong disturbance or nonlinear dynamic scene can be solved, and the robustness and adaptability of the control system in the complex plasma shape evolution process are significantly improved.

[0150] Through the fusion of reinforcement learning and imitative learning, the training efficiency is significantly improved: the traditional control strategy is introduced as a reference demonstration into the training process, a imitative reward item is constructed, the control strategy learning direction is effectively guided in the early training stage, the convergence speed of the reinforcement learning strategy is greatly improved, the training time is shortened, the parameter adjustment difficulty is reduced, and the overall engineering cost of the control strategy deployment is saved.

[0151] The reward function design is reasonable, and the stability and learning direction of the strategy are improved: the composite reward function constructed includes a target error item, an action penalty item, an error integral item and an imitative control item, which not only strengthens the sensitivity of the agent to the target control effect, but also enhances the modeling ability of the strategy to the long-term behavior trend, effectively alleviates the problem of unstable strategy caused by short-term fluctuations in training, and improves the overall control quality.

[0152] It can promote the progress of nuclear fusion research: the application first realizes long-pulse discharge based on reinforcement learning shape control on a tokamak device, and promotes the progress of reinforcement learning control in the field of nuclear fusion research. The system scheme process provided by the application will help researchers develop research innovations in this field.

[0153] Another embodiment of the application provides a plasma shape control agent strategy model construction device, as shown in Figure 8 It is a structural schematic diagram of a plasma shape control agent strategy model construction device provided by an embodiment of the application, and the device comprises:

[0154] An environment construction module is configured to construct a tokamak simulation environment and identify environment configuration parameters of a loss-of-control scene controlled by a PID controller.

[0155] A parameter configuration module is configured to configure the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment.

[0156] A learning module is configured to learn based on the training environment using a preset reinforcement learning agent model to generate a current command.

[0157] a state module configured to input the current command into the tokamak simulation environment, and calculate a control point state at each time point;

[0158] a reward module configured to input the control point state at each time point, the current command and the action command obtained by the PID controller into a preset reward function, and calculate an environment reward;

[0159] a training module configured to perform reinforcement learning training according to the environment reward, the control point state at each time point and the current command, and update a command strategy generated by the reinforcement learning agent model;

[0160] an output module configured to output the latest reinforcement learning agent model as an agent strategy model when the reinforcement learning agent model meets a preset convergence condition.

[0161] The plasma shape control agent strategy model construction device provided in the embodiment can perform all steps and functions of the plasma shape control agent strategy model construction method provided in any of the above embodiments, and the specific functions of the device will not be repeated here.

[0162] Referring to Figure 9 is another structural schematic diagram of a plasma shape control agent strategy model construction device provided in an embodiment of the present application. The plasma shape control agent strategy model construction device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a plasma shape control agent strategy model construction program. The processor implements the steps in each of the above plasma shape control agent strategy model construction method embodiments when executing the computer program, such as steps S1-S7 shown in the above embodiment. Figure 1 Alternatively, the processor implements the functions of each module in the above device embodiments when executing the computer program.

[0163] For example, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the plasma shape control agent strategy model construction device. For example, the computer program can be divided into a detection module, an output power control module and a window control module, and the specific functions of each module have been described in detail in the above plasma shape control agent strategy model construction method provided in any of the embodiments, and the specific functions of the device will not be repeated here.

[0164] The plasma shape control agent strategy model construction device can be a desktop computer, a notebook computer, a palm computer, a cloud server, or the like. The plasma shape control agent strategy model construction device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of the plasma shape control agent strategy model construction device, and does not constitute a limitation on the plasma shape control agent strategy model construction device, and can include more or fewer components than the schematic diagram, or combine certain components, or different components, for example, the plasma shape control agent strategy model construction device can also include an input / output device, a network access device, a bus, and the like.

[0165] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The processor is a control center of the plasma shape control agent strategy model construction device, and connects various parts of the plasma shape control agent strategy model construction device through various interfaces and lines.

[0166] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the plasma shape control agent strategy model construction device by running or executing the computer program and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function (such as a sound playing function, an image playing function, or the like), and the like; and the data storage area can store data created according to use of the mobile phone (such as audio data, a phonebook, or the like), and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.

[0167] If the modules integrated by the one plasma shape control agent strategy model construction device are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of each method embodiment can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0168] It should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.

Claims

1. A method for constructing a plasma configuration control proxy strategy model, characterized in that, The method comprises: constructing a tokamak simulation environment, identifying environment configuration parameters of a loss-of-control scenario under PID controller control; configuring the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment; learning based on the training environment using a preset reinforcement learning agent model to generate a current command; inputting the current command into the tokamak simulation environment to calculate the control point state at each time; inputting the control point state at each time, the current command and the action command obtained by the PID controller into a preset reward function to calculate an environment reward; performing reinforcement learning training according to the environment reward, the control point state at each time and the current command to update the command strategy generated by the reinforcement learning agent model; when the reinforcement learning agent model meets a preset convergence condition, outputting the latest reinforcement learning agent model as an agent strategy model.

2. The plasma posture control agent policy model building method of claim 1, wherein, The method further comprises: constructing a tokamak simulation environment; randomly initializing environment configuration parameters in the tokamak simulation environment to generate different simulation environments; controlling each simulation environment using the PID controller; when the controlled parameters meet a preset condition, identifying the environment configuration parameters corresponding to the loss-of-control scenario.

3. The plasma posture control agent policy model building method of claim 1, wherein, The tokamak simulation environment comprises a power supply response model, a disturbance scenario model and a plasma response model of configuration parameters.

4. The plasma posture control agent policy model building method of claim 1, wherein, The reward function comprises ; wherein, is the environment reward at time t, are weight coefficients for each term, is the control point error under the control point state at time t, is the current command output by the reinforcement learning policy at time t, is the action command output by the PID controller at time t, is the normalized error integral under the control point state from the beginning of the episode to the current time t.

5. The plasma posture control agent policy model building method of claim 4, wherein, The method further comprises: During the training process, the component of the control point error at time t is monitored under the control point state at time t; The state of the control point at time t is monitored. Control point error When the rate of increase of a component reaches a preset trend threshold, a preset adjustment model is used to reduce the weight coefficient. The value; The adjustment model is , The weight coefficient of the time t is The value of the weight coefficient of the time t, The preset initial value is The preset coefficient value is 6. The plasma posture control agent policy model building method of claim 1, wherein, The method further comprises: initializing training parameters; at the beginning of each round, randomly configuring model parameters in the tokamak simulation environment again; interacting in the currently configured tokamak simulation environment through the reinforcement learning agent, outputting a current command at each step, and the currently configured tokamak simulation environment updating the control point state according to the output current command and feeding back the environment reward and the new control point state; after a preset number of cycle updates is reached, collecting the current command, the control point state and the environment reward, and updating the command strategy of the reinforcement learning agent model based on a preset reinforcement learning algorithm.

7. The plasma posture control agent policy model building method of claim 1, wherein, The method further comprises: acquiring real-time collected environment parameters into the agent strategy model, outputting command values of each controlled object according to the agent strategy model, and performing plasma shape control.

8. A plasma posture control agent policy model construction apparatus characterized by comprising: The device comprises: an environment construction module configured to construct a tokamak simulation environment and identify environment configuration parameters of a loss-of-control scenario under PID controller control; a parameter configuration module configured to configure the environment configuration parameters into the tokamak simulation environment to obtain a configured training environment; a learning module configured to learn based on the training environment using a preset reinforcement learning agent model to generate a current command; a state module configured to input the current command into the tokamak simulation environment to calculate the control point state at each time; The reward module is configured to input the control point state at each time point and the current command and the action command obtained by the PID controller into a preset reward function to calculate an environment reward. The training module is configured to perform reinforcement learning training according to the environment reward, the control point state at each time point and the current command, and update a command strategy generated by the reinforcement learning agent model. The output module is configured to output the latest reinforcement learning agent model as an agent strategy model when the reinforcement learning agent model meets a preset convergence condition.

9. A plasma posture control agent policy model construction apparatus characterized by comprising: The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to perform the plasma posture control agent strategy model construction method according to any one of claims 1 to 7 when the computer program runs.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls a device in which the computer readable storage medium is located to perform the plasma posture control agent strategy model construction method according to any one of claims 1 to 7 when the computer program runs.

Citation Information

Patent Citations

  • Modeling simulation system for plasma control

    CN115600390A

  • Controlling magnetic field of magnetically confined device using neural network

    CN117616512A