Nuclear power plant operation regulation optimization method based on Sarsa algorithm
By applying reinforcement learning methods based on Sarsa algorithm in nuclear power plants, the operation procedures of nuclear power plants are optimized, and the problem of complex operation procedures of nuclear power plants is solved and it is difficult to quickly deal with sudden failures is achieved, and a more efficient and safe decision-making process is achieved.
Patent Information
- Application Number
- CN202311723556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-13
AI Technical Summary
Nuclear power plants have complex operating procedures and are difficult to deal with sudden failures quickly. Traditional methods have problems such as high trial and error costs and long experimental cycles.
The reinforcement learning method based on the Sarsa algorithm is adopted, and the reinforcement learning model is designed, model training and optimization is carried out to optimize the operating procedures of the nuclear power plant.
It realizes automatic adaptation and self-learning of nuclear power plant operating procedures, improves the decision-making efficiency and safety of nuclear power plants in the event of sudden failures, and reduces resource waste.
Smart Images

Figure CN120145797A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of nuclear power simulation, and particularly relates to a method for optimizing the operation regulations of a nuclear power plant based on the Sarsa algorithm. Background Art
[0002] Reinforcement learning is a data-driven decision-making technology with characteristics such as autonomous learning and high nonlinearity. By continuously interacting the training object with the environment, obtaining feedback information from the environment and adjusting its own strategy, a specific goal is finally achieved or the benefit of a certain behavior is maximized. It can effectively address a series of problems faced by industrial control, such as high trial-and-error costs and long experimental cycles.
[0003] Nuclear power control has characteristics such as multi-variables, strong coupling, nonlinearity, and large time delay. Controlling the operation regulations of a nuclear power plant under different working conditions is a typical decision-making problem. Applying reinforcement learning to this has high research value and can quickly implement it in actual operations, solving common problems faced in traditional nuclear power industrial control. For example, the operation procedure steps for a steam generator heat transfer tube rupture are as many as 56 steps, and each step faces actual and operational choices. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for optimizing the operation regulations of a nuclear power plant based on the Sarsa algorithm. Using the existing nuclear power simulator as the environmental model and the existing operation procedures as the starting point and imitation object of reinforcement learning, relying on expert resources such as mechanism model engineers and senior operators in the simulation center, training the reinforcement learning model to achieve the automatic adaptation of the reinforcement learning model to environmental changes and continuous self-learning and evolution. Iteratively learning each step of operation under typical nuclear power conditions according to the feedback of the environmental model to obtain the optimal operation regulations under this condition. Provide research ideas and guiding opinions for optimizing the operation regulations of nuclear power, and thus contribute to improving the safety and economy of nuclear power operation.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] A method for optimizing the operation regulations of a nuclear power plant based on the Sarsa algorithm,
[0007] (1) State space: Taking the breakage fault of the steam pipeline inside the containment as the research object, dividing the states in this application scenario to obtain the state space, including temperature state, pressure state, flow state, liquid level state, sensor state, and equipment switch state; the reinforcement learning model makes decisions based on the perceived environmental state;
[0008] (2) Action Space: Taking the break failure of the steam pipeline inside the containment as the research object, the actions in this application scenario are classified to obtain the action space, including closing valves, starting pumps, discharging gases, isolating areas, and triggering alarms; the reinforcement learning model takes corresponding actions according to the selected decisions.
[0009] (3) Steps of Reinforcement Learning Research:
[0010] Building a Simulation Model and Data Preparation: Build a simulation model that can simulate the break failure of the steam pipeline inside the containment. This model includes environmental states, action space, and reward function; make the simulation model simulate the break failure situation under different environmental states to generate training data; each data sample includes the current state, the actions taken, environmental feedback, and whether the termination state is reached.
[0011] Design of Reinforcement Learning Model: The reinforcement learning model algorithm used is the Sarsa algorithm. By continuously trying and updating the policy in the environment, the Sarsa algorithm gradually learns which actions to take in specific states to obtain the maximum cumulative reward.
[0012] Model Training and Optimization: Use the experience replay mechanism to randomly sample samples from the training data for training to reduce the correlation between data; use a higher exploration probability at the beginning of training and gradually reduce the exploration probability to enable the model to balance between exploration and exploitation; use a target network to stabilize training and regularly update the parameters of the target network.
[0013] Model Evaluation and Iterative Optimization: Use the simulator to evaluate the performance of the trained model in an untrained test environment, observe its decision-making performance in different situations, and optimize the model according to the evaluation results.
[0014] Based on various state categories, the state space is divided into multi-dimensional states, and each dimension represents an environmental state. In this state representation, each state dimension contains different values or states, representing various states that may occur in the case of a break failure.
[0015] The reinforcement learning model selects one or more actions to respond to the break failure according to the current environmental state and the reward function.
[0016] Conduct reinforcement learning based on the specific application of the simulation model, the relevant mechanism principles of the nuclear power simulator, and the business logic of the nuclear power scenario; build and train the model according to the obtained data, and iteratively optimize the intelligent agent of reinforcement learning through the feedback of the simulator.
[0017] Use the MAAP5 software to make the simulation model simulate the break failure situation under different environmental states to generate training data.
[0018] Environmental feedback refers to the next state and the reward.
[0019] The termination state refers to the resolution of the fault or an emergency shutdown.
[0020] The model is tuned according to the evaluation results. The content to be adjusted includes, but is not limited to, the neural network structure, training hyperparameters, and reward function.
[0021] The trained reinforcement learning model is deployed to the actual nuclear power scenario. During operation, data is continuously collected, and the model is optimized online based on the actual feedback to adapt to the changes and uncertainties in the real environment.
[0022] The beneficial effects achieved by the present invention are as follows:
[0023] By combining the reinforcement learning model with the specific application scenario of nuclear power, more effective decisions can be made through the counterintuitive trade-off between samples and computational efficiency. In the link of simulating accident handling in a nuclear power plant, it can assist in improving the decision-making efficiency of personnel, reducing resource waste at critical moments, providing technical support for digital nuclear power, and facilitating the construction of digital twins of nuclear power plants. Description of the Drawings
[0024] Figure 1 It is a flowchart of the present invention;
[0025] Figure 2 It is a relationship diagram of the present invention with the platform. Detailed Embodiments
[0026] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0027] (1) State Space
[0028] Taking the break fault of the steam pipeline inside the containment as the research object, the states in this application scenario are divided to obtain the state space. The reinforcement learning model makes decisions based on the perceived environmental states.
[0029] For the research object of the break fault of the steam pipeline inside the containment, the state space is divided into the following to enable the reinforcement learning model to make decisions based on the perceived environmental states: 1. Temperature state; 2. Pressure state; 3. Flow state; 4. Liquid level state; 5. Sensor state; 6. Equipment switch state, etc.
[0030] Based on the above state categories, the state space is divided into multi-dimensional states, and each dimension represents an environmental state. In this state representation, each state dimension contains different values or states, representing various states that may occur in the case of a break fault.
[0031] (2) Action Space
[0032] Taking the break failure of the steam pipeline inside the containment as the research object, the actions in this application scenario are divided to obtain the action space. The reinforcement learning model takes corresponding actions according to the selected decision.
[0033] For the research object of the break failure of the steam pipeline inside the containment, the action space is divided into the following to enable the reinforcement learning model to take corresponding actions according to the selected decision: 1. Close the valve; 2. Start the pump; 3. Discharge the gas; 4. Isolate the area; 5. Trigger the alarm, etc.
[0034] The reinforcement learning model selects one or more actions to cope with the break failure according to the current environmental state and the reward function.
[0035] (3) Reinforcement learning research steps
[0036] Based on the basic research conditions such as the specific application of the simulation model, the relevant mechanism principles of the nuclear power simulator, and the business logic of the nuclear power scenario, reinforcement learning is carried out. Build and train the model according to the obtained data, and iteratively optimize the agent of reinforcement learning through the feedback of the simulator.
[0037] 1) Simulation model and data preparation
[0038] Build a simulation model. Based on the relevant mechanism principles of the nuclear power simulator and the business logic of the nuclear power scenario, build a simulation model that can simulate the break failure of the steam pipeline inside the containment. This model needs to include the environmental state, action space, reward function, etc.
[0039] Generate training data. Use the MAAP5 software to make the simulation model simulate the break failure situation under different environmental states and generate training data. Each data sample should include the current state, the actions taken, the environmental feedback (the next state and the reward), and whether the termination state is reached (such as the failure is resolved or an emergency shutdown).
[0040] 2) Reinforcement learning model design
[0041] The reinforcement learning model algorithm used in the research is the Sarsa algorithm. The Sarsa algorithm is an algorithm for learning the policy of the Markov decision process. Its core idea is to learn the optimal policy by continuously updating the state-action value function Q(s,a).
[0042] By continuously trying and updating the policy in the environment, the Sarsa algorithm can gradually learn which actions to take in a specific state to obtain the maximum cumulative reward, thus realizing the reinforcement learning of the agent.
[0043] 3) Model training and optimization
[0044] Use the experience replay mechanism to randomly sample samples from the training data for training to reduce the correlation between data.
[0045] At the initial stage of training, use a higher exploration probability (such as the ε-greedy strategy), and gradually reduce the exploration probability to enable the model to balance between exploration and exploitation.
[0046] Use a target network to stabilize training and regularly update the parameters of the target network.
[0047] 4) Model evaluation and iterative optimization
[0048] Use a simulator to evaluate the performance of the trained model in an untrained test environment and observe its decision-making performance under different conditions. According to the evaluation results, optimize the model. The content that needs to be adjusted includes but is not limited to the neural network structure, training hyperparameters, reward function, etc.
[0049] Deploy the trained reinforcement learning model to the actual nuclear power scenario. During operation, continuously collect data and perform online optimization of the model according to the actual feedback to adapt to the changes and uncertainties in the real environment.
Claims
1. A method for optimizing the operating procedures of nuclear power plants based on the Sarsa algorithm, characterized in that: (1) State space: Taking the breakage fault of the steam pipeline inside the containment as the research object, the states in this application scenario are divided to obtain the state space, including temperature state, pressure state, flow state, liquid level state, sensor state, and equipment switch state; The reinforcement learning model makes decisions based on the perceived environmental state; (2) Action space: Taking the breakage fault of the steam pipeline inside the containment as the research object, the actions in this application scenario are divided to obtain the action space, including closing the valve, starting the pump, discharging gas, isolating the area, and triggering an alarm; The reinforcement learning model takes corresponding actions according to the selected decision; (3) Steps of reinforcement learning research: Building a simulation model and data preparation: Build a simulation model that can simulate the breakage fault of the steam pipeline inside the containment. This model includes environmental state, action space, and reward function; Make the simulation model simulate the breakage fault situation under different environmental states to generate training data; Each data sample includes the current state, the actions taken, environmental feedback, and whether the termination state is reached; Reinforcement learning model design: The reinforcement learning model algorithm used is the Sarsa algorithm. By continuously trying and updating the policy in the environment, the Sarsa algorithm gradually learns which actions to take in specific states to obtain the maximum cumulative reward; Model training and optimization: Use the experience replay mechanism to randomly sample samples from the training data for training to reduce the correlation between data; Use a higher exploration probability at the initial stage of training and gradually reduce the exploration probability to enable the model to balance between exploration and exploitation; Use a target network to stabilize the training and regularly update the parameters of the target network; Model evaluation and iterative optimization: Use the simulator to evaluate the performance of the trained model in an untrained test environment, observe its decision-making performance in different situations, and optimize the model according to the evaluation results.
2. The method for optimizing the operating procedures of nuclear power plants based on the Sarsa algorithm according to claim 1, characterized in that: Based on various state categories, the state space is divided into multi-dimensional states. Each dimension represents an environmental state. In this state representation, each state dimension contains different values or states, representing various states that may occur in the case of a breakage fault.
3. The method for optimizing the operating procedures of nuclear power plants based on the Sarsa algorithm according to claim 1, characterized in that: The reinforcement learning model selects one or more actions according to the current environmental state and the reward function to deal with the breakage fault.
4. The method for optimizing the operating procedures of nuclear power plants based on the Sarsa algorithm according to claim 1, characterized in that: Reinforcement learning is carried out based on the specific application of the simulation model, the relevant mechanism principles of the nuclear power simulator, and the business logic of the nuclear power scenario; Build and train the model according to the obtained data, and iteratively optimize the intelligent agent of the reinforcement learning through the feedback of the simulator.
5. The method for optimizing the operating procedures of nuclear power plants based on the Sarsa algorithm according to claim 1, characterized in that: Using the MAAP5 software, the simulation model simulates the break failure situation under different environmental states to generate training data.
6. The method for optimizing the operating procedures of a nuclear power plant based on the Sarsa algorithm according to claim 1, characterized in that: Environmental feedback refers to the next state and reward.
7. The method for optimizing the operating procedures of a nuclear power plant based on the Sarsa algorithm according to claim 1, characterized in that: The termination state refers to the resolution of the failure or an emergency shutdown.
8. The method for optimizing the operating procedures of a nuclear power plant based on the Sarsa algorithm according to claim 1, characterized in that: The model is optimized according to the evaluation results. The content that needs to be adjusted includes but is not limited to the neural network structure, training hyperparameters, and reward function.
9. The method for optimizing the operating procedures of a nuclear power plant based on the Sarsa algorithm according to claim 1, characterized in that: The trained reinforcement learning model is deployed to the actual nuclear power scenario. During operation, data is continuously collected, and the model is optimized online according to the actual feedback to adapt to the changes and uncertainties in the real environment.