Reinforced learning training method for plasma control simulation platform
By building a plasma response model and designing a reinforcement learning interface in the plasma control simulation verification platform, the problem of the reinforcement learning controller training relying on an accurate simulation environment and the algorithm update of the simulation platform is limited, and an efficient and general reinforcement learning training method is realized, and a high-precision controller is generated.
Patent Information
- Application Number
- CN202510158878.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
The training of reinforcement learning controllers depends on an accurate simulation training environment, and the algorithm update of simulation platforms in the prior art is limited, and the training takes a long time.
The plasma response model is built in the plasma control simulation verification platform (PCSVP), and a reinforcement learning interface is designed, including an independent simulation engine and reinforcement learning signal interaction module (RLBlock). The plasma response model is encapsulated into a standardized Gym environment object to achieve compatibility with the external reinforcement learning framework.
By encapsulating an independent simulation engine and designing a standardized reinforcement learning interface, the time cost during the training process is significantly reduced, the overall training cycle is shortened, training efficiency and versatility is improved, and it supports the rapid construction of a high-fidelity simulation environment and the generation of high-precision controllers.
Smart Images

Figure CN120010262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of plasma control simulation platform, and in particular to a reinforcement learning training method for a plasma control simulation platform. Background Art
[0002] Through international cooperation, we digested and absorbed the simulation technology of ITER's plasma control simulation platform (PCSSP), and studied and developed the plasma control simulation verification platform (PCSVP) to meet the simulation and verification needs of the new generation of control systems and control algorithms. Based on the open source Python technology stack, we implemented the visual modeling and simulation functions similar to MATLAB / Simulink, customized a variety of plasma control simulation interfaces, and broke through the bottleneck technology of foreign simulation software. PCSVP is a visual modeling platform similar to Simulink, and integrates general simulation calculation modules, including transfer function calculation and matrix operations. In addition, we have also customized relevant simulation modules for plasma control, including reading and writing modules for non-relational data in the fusion field, PID control modules with filtering for plasma control, and real-time communication modules for docking with plasma control systems. Therefore, the control algorithm model established based on PCSVP can be quickly tested and verified.
[0003] At present, controllers trained based on reinforcement learning have achieved good results in the field of fusion, but the training of reinforcement learning controllers is very dependent on an accurate simulation training environment. PCSVP has the ability of visual modeling and customized simulation modules related to plasma control, which helps researchers quickly and efficiently build and test training environments related to plasma control. In addition, the current reinforcement learning framework is diverse, and advanced algorithms are constantly being explored, so it is necessary to customize a PCSVP-based and universal reinforcement control simulation training method.
[0004] In summary, the present application proposes a reinforcement learning training method for a plasma control simulation platform. Summary of the invention
[0005] The purpose of the present invention is to propose a reinforcement learning training method for a plasma control simulation platform in order to address the problem in the background technology that the training of a reinforcement learning controller is highly dependent on an accurate simulation training environment.
[0006] The technical solution of the present invention is a reinforcement learning training method for a plasma control simulation platform, comprising the following steps:
[0007] S1: constructing a plasma response model in a plasma control simulation verification platform (PCSVP), wherein the plasma response model integrates a plasma control dedicated module through a visual modeling method;
[0008] S2: Design a reinforcement learning interface, the interface comprising an independent simulation engine and a reinforcement learning signal interaction module (RLBlock);
[0009] The independent simulation engine is encapsulated as a class that can be independently operated from the plasma control simulation verification platform framework, and is used to drive the calculation of the plasma response model, and realize the interaction between the agent model (Agent) and the model through the data sharing area;
[0010] The reinforcement learning signal interaction module is connected to the plasma response model, and is used to receive the action signal output by the agent model, and transmit the observation state, reward signal and termination signal to the agent model;
[0011] S3: Encapsulating the plasma response model into a standardized Gym environment object, defining an environment processing class (envHandler) by inheriting the Gym environment class, and achieving compatibility with an external reinforcement learning framework;
[0012] S4: The plasma response model is driven by the independent simulation engine, and the interaction between the agent model (Agent) and the model is completed in combination with the reinforcement learning signal interaction module (RL Block), and reinforcement learning training is performed without the participation of a graphical user interface (GUI) to generate an optimal control strategy model.
[0013] Optionally, the data sharing area is a memory sharing area in an independent simulation engine, which is used to store the action signal output by the agent model (Agent) and the observation state, reward signal and termination signal collected by the reinforcement learning signal interaction module (RL Block), so as to realize zero communication delay data interaction between the agent model (Agent) and the model.
[0014] Optionally, the input of the reinforcement learning signal interaction module (RL Block) includes an observation state, a reward signal and a termination signal from a plasma response model, and the output is an action signal generated by the Agent; the observation state includes at least one of the vertical position of the plasma, the fast control coil current and its historical data.
[0015] Optionally, the environment handling class (envHandler) provides a standardized reinforcement learning interface method, including a step-by-step simulation method (step), an environment reset method (reset), and a reward calculation logic, to adapt to the training process of different reinforcement learning frameworks.
[0016] Optionally, the independent simulation engine is implemented by a simulation calculation module of a packaged plasma control simulation verification platform (PCSVP), including transfer function calculation, matrix operation and driving logic of a plasma control dedicated module.
[0017] Optionally, the optimal control strategy model is generated through tens of thousands of iterative trainings, during which the agent model strategy is optimized based on the reward signal, and finally a controller for driving the plasma response model is output.
[0018] Optionally, the controller is integrated into a plasma control simulation verification platform (PCSVP) through a machine learning drive module and docked with a plasma response model for real-time control of the plasma vertical position and fast control coil current.
[0019] Optionally, the method further includes a training environment verification step:
[0020] (a) Construct a training environment for the vertical displacement fast controller, and input the plasma vertical position, fast control coil current and its historical data as observation states into the reinforcement learning signal interaction module (RL Block);
[0021] (b) driving the training model through the environment processing class (envHandler) to complete the iterative training of the reinforcement learning algorithm;
[0022] (c) The control model generated by training is connected to the plasma response model to test the control effect and optimize the parameters.
[0023] Optionally, the method is applicable to at least one application scenario of plasma vertical displacement control, fast control coil current regulation and plasma forming optimization in a fusion device.
[0024] Compared with the prior art, the present invention has at least one of the following beneficial technical effects:
[0025] By encapsulating an independent simulation engine and adopting a driver mode without a graphical interface, the additional overhead caused by component communication and visual rendering in traditional simulation platforms is avoided, significantly reducing the time cost during training. The design of the data sharing area realizes zero communication delay interaction between the agent and the model, which is especially suitable for reinforcement learning training scenarios that require tens of thousands of iterations, greatly shortening the overall training cycle.
[0026] The plasma response model is encapsulated using the Gym environment standardized interface, making this method compatible with mainstream reinforcement learning frameworks, supporting researchers to directly call the latest algorithms from the open source community, overcoming the limitation of algorithm updates on closed platforms such as MATLAB / Simulink. At the same time, the standardized method provided by the environment handling class (envHandler) simplifies the adaptation process of different frameworks and improves the scalability of the method.
[0027] Based on PCSVP's visual modeling capabilities and plasma control-specific modules, researchers can quickly build a high-fidelity simulation environment without rewriting the underlying code. The modular design of RL Block allows direct observation of key state quantities and dynamically optimizes strategies through reward signals, accelerating the development and verification of controllers.
[0028] The optimal control strategy model is generated through tens of thousands of iterative training, and the agent decision logic is optimized by combining the reward signal with real-time feedback. The final output controller shows high precision and strong robustness under complex working conditions. In the implementation example, the controller can stabilize the vertical position of the plasma to the target value in a short time, and the current fluctuation of the fast control coil converges rapidly, which verifies the effectiveness of the method.
[0029] The present invention significantly improves the training efficiency, flexibility and versatility by encapsulating an independent simulation engine, designing a standardized reinforcement learning interface and being compatible with the Gym environment. It supports the rapid construction of a high-fidelity simulation environment and the generation of a high-precision controller, solving the problems of limited algorithm updates and time-consuming training on traditional platforms, and is suitable for a variety of control scenarios in fusion devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Design diagram for RL interface;
[0031] Figure 2 This is the schematic diagram of the signal interaction module for reinforcement learning;
[0032] Figure 3 This is a relationship diagram of the reinforcement learning interface class;
[0033] Figure 4 Constructing a graph for the training environment of a fast vertical displacement controller;
[0034] Figure 5 A drivable construction diagram for the visualization model;
[0035] Figure 6 Build graphs for testing models of hardened controllers;
[0036] Figure 7 is the change diagram of the vertical position of plasma;
[0037] Figure 8This is the fast control current change diagram. DETAILED DESCRIPTION
[0038] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and specific embodiments.
[0039] Example 1
[0040] The present invention proposes a reinforcement learning training method for a plasma control simulation platform, and the method flow is described in detail below.
[0041] 1. Construct a plasma response model in the plasma control simulation verification platform (PCSVP), and integrate the plasma control dedicated module into the plasma response model through visual modeling;
[0042] 2. Design a reinforcement learning interface, which includes an independent simulation engine and a reinforcement learning signal interaction module (RLBlock);
[0043] The independent simulation engine is encapsulated as a class that can run independently from the plasma control simulation verification platform framework, which is used to drive the calculation of the plasma response model and realize the interaction between the agent model (Agent) and the model through the data sharing area; the data sharing area is a memory sharing area in the independent simulation engine, which is used to store the action signal output by the agent model (Agent) and the observation state, reward signal and termination signal collected by the reinforcement learning signal interaction module (RL Block), so as to realize zero communication delay data interaction between the agent model (Agent) and the model;
[0044] The reinforcement learning signal interaction module is connected to the plasma response model, and is used to receive the action signal output by the agent model, and transmit the observation state, reward signal and termination signal to the agent model; the input of the reinforcement learning signal interaction module (RLBlock) includes the observation state, reward signal and termination signal from the plasma response model, and the output is the action signal generated by the agent; the observation state includes at least one of the plasma vertical position, the fast control coil current and its historical data;
[0045] 3. Encapsulate the plasma response model into a standardized Gym environment object, and define the environment processing class (envHandler) by inheriting the Gym environment class to achieve compatibility with the external reinforcement learning framework; the environment processing class (envHandler) provides standardized reinforcement learning interface methods, including step-by-step simulation method (step), environment reset method (reset) and reward calculation logic, to adapt to the training process of different reinforcement learning frameworks.
[0046] Fourth, the plasma response model is driven by an independent simulation engine, and the interaction between the agent model (Agent) and the model is completed by combining the reinforcement learning signal interaction module (RLBlock), and reinforcement learning training is performed without the participation of a graphical interface (GUI) to generate an optimal control strategy model. The independent simulation engine is implemented through the simulation calculation module of the packaged plasma control simulation verification platform (PCSVP), including transfer function calculation, matrix operation and driving logic of the plasma control dedicated module.
[0047] Among them, the optimal control strategy model is generated through tens of thousands of iterative training. During the training process, the agent model (Agent) strategy is optimized based on the reward signal, and finally the controller used to drive the plasma response model is output.
[0048] In addition, the controller is integrated into the plasma control simulation verification platform (PCSVP) through a machine learning driver module and docked with the plasma response model for real-time control of the plasma vertical position and fast control of the coil current.
[0049] In this embodiment, the training environment verification step is also included:
[0050] (a) Construct a training environment for the vertical displacement fast controller, and input the plasma vertical position, fast control coil current and its historical data as observation states into the reinforcement learning signal interaction module (RL Block);
[0051] (b) Drive the training model through the environment processing class (envHandler) to complete the iterative training of the reinforcement learning algorithm;
[0052] (c) The control model generated by training is connected to the plasma response model to test the control effect and optimize the parameters.
[0053] It is worth mentioning that the reinforcement learning training method for the plasma control simulation platform is suitable for at least one application scenario in the plasma vertical displacement control, fast control coil current regulation and plasma forming optimization in the fusion device.
[0054] Example 2
[0055] like Figure 1As shown in the figure, the RL interface is designed into two parts, an independent engine and a reinforcement learning signal interaction module (RLBlock). In order to facilitate the agent to drive the model and reduce the time cost caused by communication, an independent engine that can run independently without the framework is encapsulated on the basis of the PCSVP simulation engine, and a data sharing area is set up in the independent engine as a data bridge between the agent and the environment. The agent can write command signals to the sharing area through the independent engine and obtain observation signals. At the same time, the RL interface also includes an RL Block specifically used to collect observation information of the environment. It can be connected to the model, write the data to be observed into the sharing area, and transmit the agent's action signal to the model.
[0056] Gym is a popular and diverse reference environment collection with a well-unified environment interface that can well represent general RL training environment problems. Therefore, Gym is chosen to encapsulate the PCSVP model creation environment, which can be applied to most RL frameworks.
[0057] The construction of the reinforcement learning interface is divided into three parts:
[0058] (1) First, define a reinforcement learning interface module (RL Block) through the framework's Block specification. Figure 2 As shown, it makes the agent model (Agent)'s behavioral decision (Action) for the response model, and receives the model state (State) that the agent needs to observe, the decision reward (Reward), and the signal whether the simulation is ended (Terminated).
[0059] (2) In order to realize the interaction between Agent and RL Block, it is necessary to drive the response model connected to RL Block during training. In the framework, the simulation engine is the main object that controls the simulation process and drives the model calculation, so we encapsulate the simulation engine as a class that can run independently from the framework so that the Agent can drive the plasma response model.
[0060] (3) Gym is an open source toolkit for developing and comparing reinforcement learning algorithms. It provides a series of standardized environments that enable researchers and developers to experiment, test, and compare the performance of various reinforcement learning algorithms on different problems. Many reinforcement learning frameworks provide methods to call the Gym tool to test algorithm performance. Therefore, by packaging the model on PCSVP as a Gym environment object, it can be adapted to most reinforcement learning frameworks. Figure 3As shown, a PCSVP envHandler class is defined to inherit the Gym.environment class, and the independent simulation engine is called through the class to drive the plasma response model including the interface Block, and transmit the signal through the interface Block.
[0061] This method of integrating the model with the reinforcement learning framework takes advantage of the PCSVP visual modeling and can help the agent model of reinforcement training to quickly build the model environment. At the same time, through an independent simulation engine, there is no GUI involved in the training process and no need for component communication, which greatly improves the training efficiency. The interface inherits the environmental standards of Gym, so that both custom training codes and training algorithms provided by the RL framework can be used, which has high versatility.
[0062] Example 3
[0063] The RL interface test is divided into two parts. The first is to use PCSVP and RL interface to build a training environment for RL framework training test. The second is to test the AI driver interface of the controller with better performance (higher reward) after training, and observe the control effect of the RL controller. On the one hand, this solution verifies that the RL interface of PCSVP can complete the fusion simulation of PCSVP model and reinforcement learning framework, and on the other hand, it tests the practicality of AI model driver interface. The following will take the training of vertical displacement fast controller as an example for testing.
[0064] First, if Figure 4 As shown, we need to build a fast control response model, connect the State state quantity that needs to be observed to the RL Block, and accordingly connect the decision Action output by the RL Block to the response model (discrete state-space). There are three state quantities that need to be observed: the vertical position of the plasma (Z), the current current of the fast control coil (IC), and the IC current at the previous moment. The calculation method of Reward will be constructed in the envHandler class, and the Terminated signal does not need to be specially customized, so the Reward and Terminated ports are set to empty.
[0065] Then, use Figure 5The envHandler class shown above constructs a drivable training model for building a vertical displacement fast controller, that is, constructs the graphical model into code that can be directly driven by an external framework, and provides methods such as step (a method for step-by-step simulation) and reset (a method for resetting the training environment) to support reinforcement training. Each round of reinforcement learning training will produce a control model, and a training session will often go through tens of thousands of iterations, and a more ideal control model will be obtained through continuous trial and error and correction.
[0066] Finally, if Figure 6 As shown, by writing the control model file into the machine learning driver module (zx_RL_ctrl), the graphical calculation drive of the controller can be completed to form a vertical displacement controller. Then, the controller is connected to the fast control response model to test the controller's control ability over the response model.
[0067] The main purpose of the rapid control of vertical displacement is to quickly pull the vertical position of the plasma to the target position (usually 0) through the IC coil to ensure that the plasma can be quickly formed. Figure 7 and Figure 8 To enhance the control effect of the controller, it can be seen from the comparison that the vertical position Z of the controlled plasma is quickly pulled to near 0 in a short time, while Figure 8 The IC current in the circuit also stabilizes after rapid fluctuations. Here, only a simple interface test is performed without parameter adjustment and tuning. It can be seen that the interface can be used for enhanced training of the model and a better controller can be obtained after more training.
[0068] The above specific embodiments are only several optional embodiments of the present invention. Based on the technical solutions of the present invention and the relevant inspirations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A reinforcement learning training method for a plasma control simulation platform, characterized in that: The following steps are involved: S1: constructing a plasma response model in a plasma control simulation verification platform, wherein the plasma response model integrates a plasma control dedicated module through a visual modeling method; S2: Design a reinforcement learning interface, the interface comprising an independent simulation engine and a reinforcement learning signal interaction module; The independent simulation engine is encapsulated as a class that can be independently operated from the plasma control simulation verification platform framework, and is used to drive the calculation of the plasma response model, and realize the interaction between the proxy model and the model through the data sharing area; The reinforcement learning signal interaction module is connected to the plasma response model, and is used to receive the action signal output by the agent model, and transmit the observation state, reward signal and termination signal to the agent model; S3: Encapsulating the plasma response model into a standardized Gym environment object, defining an environment processing class by inheriting the Gym environment class, and achieving compatibility with an external reinforcement learning framework; S4: driving the plasma response model through the independent simulation engine, completing the interaction between the agent model and the model in combination with the reinforcement learning signal interaction module, and performing reinforcement learning training without the participation of a graphical interface to generate an optimal control strategy model.
2. A reinforcement learning training method for a plasma control simulation platform according to claim 1, characterized in that: The data sharing area is a memory sharing area in an independent simulation engine, which is used to store the action signal output by the agent model and the observation state, reward signal and termination signal collected by the reinforcement learning signal interaction module, so as to realize zero communication delay data interaction between the agent model and the model.
3. The reinforcement learning training method for a plasma control simulation platform according to claim 1, characterized in that: The input of the reinforcement learning signal interaction module includes the observation state, reward signal and termination signal from the plasma response model, and the output is the action signal generated by the agent; the observation state includes at least one of the plasma vertical position, the fast control coil current and its historical data.
4. The reinforcement learning training method for a plasma control simulation platform according to claim 1, characterized in that: The environment processing class provides a standardized reinforcement learning interface method, including a step-by-step simulation method, an environment reset method, and a reward calculation logic, to adapt to the training processes of different reinforcement learning frameworks.
5. The reinforcement learning training method for a plasma control simulation platform according to claim 1, characterized in that: The independent simulation engine is realized by encapsulating the simulation calculation module of the plasma control simulation verification platform, including transfer function calculation, matrix operation and driving logic of the plasma control dedicated module.
6. The reinforcement learning training method for a plasma control simulation platform according to claim 1, characterized in that: The optimal control strategy model is generated through tens of thousands of iterative training. During the training process, the agent model strategy is optimized based on the reward signal, and finally a controller for driving the plasma response model is output.
7. A reinforcement learning training method for a plasma control simulation platform according to claim 6, characterized in that: The controller is integrated into the plasma control simulation verification platform through a machine learning drive module and connected to the plasma response model for real-time control of the plasma vertical position and fast control coil current.
8. A reinforcement learning training method for a plasma control simulation platform according to claim 1, the method further comprising a training environment verification step: (a) Construct a training environment for the vertical displacement fast controller, and input the plasma vertical position, fast control coil current and its historical data as observation states into the reinforcement learning signal interaction module; (b) driving the training model through the environment processing class to complete iterative training of the reinforcement learning algorithm; (c) The control model generated by training is connected to the plasma response model to test the control effect and optimize the parameters.
9. A reinforcement learning training method for a plasma control simulation platform according to any one of claims 1 to 8, characterized in that: The method is applicable to at least one application scenario of plasma vertical displacement control, fast control coil current regulation and plasma forming optimization in a fusion device.