Nuclear power plant reactor core optimization power control method based on imitation learning

By imitating learning and integrated learning methods, using real data of nuclear power plants to train agents, the problems of limited computing resources and insufficient applicability in multiple operating conditions in the existing technology are solved, and the efficiency, safety and accuracy of core power control of nuclear power plants are achieved.

CN120428563AActive Publication Date: 2025-08-05SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510566496.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The core power optimization control technology of existing nuclear reactors is difficult to implement with limited computing resources, and the existing models rely on simulation data, resulting in insufficient generalization capabilities of the model and cannot be applied to multi-condition control. The direct control of immature agents may lead to serious accidents.

Method used

Using a method based on imitation learning, the agent is trained using real data from the nuclear power plant, and the optimal hyperparameters are found through integrated learning and Bayesian optimization, and a multi-agent model is built, combining the data distribution of multiple working conditions to avoid immature agents directly interacting with the power plant, improving the accuracy and safety of the model.

Benefits of technology

The accuracy and consistency of core power control in multiple operating conditions is achieved, the reliability and accuracy of control is improved, the safety risks caused by immature agents are avoided, and the generalization ability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428563A_ABST
    Figure CN120428563A_ABST
Patent Text Reader

Abstract

The invention discloses a nuclear power plant reactor core optimization power control method based on imitation learning. A reactor core optimization power control model comprising a plurality of intelligent agents and an average layer is constructed based on the imitation learning method; and averaging the output results of the trained agents as a final control result. According to the invention, nuclear power plant reactor core optimization power control is developed based on an imitation learning method, an intelligent agent is trained in a supervised manner to learn a nuclear power plant power control system and decision logic of an operator, and serious consequences caused by direct interaction between an immature intelligent agent and a power plant are avoided while real data of the power plant are effectively utilized. A plurality of agents are integrated by adopting an integrated learning method, so that the model can be suitable for multiple working conditions of peak regulation, normal start-stop and emergency shut-down of a plurality of nuclear power plants. Each agent is optimized based on the Bayesian method in the training process, the optimal hyper-parameter is searched, and the network training efficiency and the fault diagnosis precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of nuclear power plant control, specifically a nuclear power plant core optimization power control method based on imitation learning. Background Art

[0002] Existing nuclear reactor core power optimization control technologies face several challenges in practical application. For example, the computational burden of deep learning networks makes them difficult to implement with existing computing resources. Most models rely on simulation data or environments. For example, reinforcement learning requires agents to interact directly with their environment during training, learning through trial and error. Direct control of nuclear power plants by immature agents could lead to serious accidents, so these technologies must be implemented within power plant simulations. However, the accuracy of simulation environments may not fully reflect the actual operating environment, thus affecting the generalization capabilities of the models. Furthermore, most research focuses on specific power plant operating conditions, making the results difficult to apply to the multi-condition control of actual power plants. Summary of the Invention

[0003] In response to the above-mentioned shortcomings of the existing technology, the present invention proposes a method for optimizing nuclear power plant core power control based on imitation learning. Based on the imitation learning method, the method develops optimized nuclear power plant core power control, and supervised training of intelligent agents to learn the nuclear power plant power control system and the operator's decision logic. This effectively utilizes real power plant data while avoiding the serious consequences of immature agents directly interacting with the power plant. An ensemble learning method is used to integrate multiple intelligent agents, making the model applicable to multiple operating conditions such as peak load regulation, normal startup and shutdown, and emergency shutdown of multiple nuclear power plants. During the training process, each intelligent agent is optimized based on the Bayesian method to find the optimal hyperparameters, thereby improving the efficiency of network training and the accuracy of fault diagnosis.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a nuclear power plant core optimization power control method based on imitation learning. After constructing a core optimization power control model based on the imitation learning method, the output results of the trained intelligent agents are averaged as the final control result. The core optimization power control model includes: multiple intelligent agents and an averaging layer, wherein: each intelligent agent learns the historical operating data of the real nuclear power plant through random sampling.

[0006] The training described above uses Bayesian optimization to find the optimal hyperparameters, reducing the complexity of manual parameter adjustment and improving the accuracy of the model.

[0007] The present invention relates to a system for implementing the above-mentioned method, comprising: a data preprocessing unit, an integrated sampling unit, an agent training unit and an integrated control unit, wherein: the data preprocessing unit collects historical operating data of the nuclear power plant, performs data sparseness, data cleaning, state action reconstruction and standardization processing, and obtains float32 type tensor data that conforms to deep learning data; the integrated sampling unit performs self-service sampling with replacement processing on the data preprocessed by the data preprocessing unit to obtain a training subset containing approximately 63.2% of the original data volume; the agent training unit performs multi-agent training on multiple sub-data set results obtained by self-service sampling of the integrated sampling unit to obtain multiple mature trained agents, and performs control; the integrated control unit performs average processing based on the control results of multiple agents in the agent training unit to obtain the final control result. Technical Effects

[0008] This invention utilizes an integrated learning architecture for multiple nuclear power plant operating conditions. It generates training subsets with differentiated operating condition distributions through self-service sampling. Using Bayesian optimization in a continuous action space, it embeds control rod stepping constraints within a Gaussian process proxy model, enabling the coordinated optimization of hyperparameter search and nuclear power safety rules. Compared to existing technologies, this invention overcomes the limitations of single-condition models, achieving 98.4% control consistency under peak-shaving and startup / shutdown conditions, a 41.7% improvement over traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 Flowchart of the present invention;

[0010] Figure 2 This is the data preprocessing flow chart;

[0011] Figure 3 Schematic diagram of the intelligent agent structure;

[0012] Figure 4 Flowchart for Bayesian optimization;

[0013] Figure 5 Schematic diagram for multi-agent integration;

[0014] Figure 6-Figure 7 This is a schematic diagram of the peak load regulation control results of Units 1-2;

[0015] Figure 8-Figure 9 This is a schematic diagram of the control results of the startup conditions of Units 1-2. DETAILED DESCRIPTION

[0016] like Figure 1 As shown, this embodiment relates to a nuclear power plant core optimization power control method based on imitation learning, including:

[0017] Step 1: Collect and preprocess raw data, including:

[0018] 1.1 Collect operating data and historical data of the distributed control system (DCS) of the nuclear power plant, including: operating data including: parameters of the primary and secondary circuit systems of the nuclear power plant, such as primary circuit coolant temperature, steam generator water temperature, main steam pressure transmitter and control rod position; historical data including: normal operating conditions, start-up and shutdown conditions and peak load conditions.

[0019] 1.2 Data sparse processing: such as Figure 2 As shown, the original sampling sequence is defined as: t k =t0+kΔt raw , Δt raw is the original data sampling interval, and 60-second interval resampling is achieved through the following function mapping: t′ m =t0+nΔT,ΔT=60s. In case of missing sampling point data, linear interpolation method is used to fill the missing data: Where: t a =max{t′ m |t′ m <t},t b =min{t′ m |t′ m >t}.

[0020] In the processing process of this embodiment, 60 seconds is used as the resampling interval.

[0021] 1.3 Data cleaning, including two types of failure data cleaning: discrete failure and continuous failure. For discrete failure data, linear interpolation is used to fill in the missing data. For continuous failure data, the failure time period and the number of data points are first recorded, and the data characteristics of the time period are manually analyzed to fill in or discard the data.

[0022] When the data values before and after the failure point do not change much or their rate of change is small, the average value of the data before and after is used to fill the data, specifically: Otherwise, it indicates that the nuclear power plant is in a non-steady-state operating condition, and this period of invalid data is discarded.

[0023] 1.4 Calculate the initial state (State) and initial decision (Action): Take the 67 reactor parameters and the current position of the 18 rod groups at time T as State, and the difference x between the position of the 18 rod groups at time T+1 and the position at time T T+1 -x T As an Action to recalculate the entire dataset.

[0024] 1.5 Data standardization: The Z-Score method is used to standardize the data set, and the original data is processed into standard data with a mean of 0 and a standard deviation of 1. Specifically: in: x ol is the data calculated in 1.4, μ is x old The mean of δ is x old The variance of x is n, the number of samples, new The data are obtained after standardization.

[0025] 1.6 Format conversion: Convert the data type to float32 type to obtain the original data set D containing n samples.

[0026] Step 2: Sample the original data set obtained in step 1 and generate a training set. Specifically, for an original data set D containing n samples, a new data set D is generated by the bootstrap sampling method. ′ , that is, each time a sample is randomly extracted from the original data set and copied to the new data set D ′ In this way, the sample may still be drawn next time. After repeating n times, a new data set D can be obtained. ′ , the number of samples is n, which is the same as the number of samples in the original dataset.

[0027] The above sampling method will result in some samples being drawn multiple times, while others will not be drawn. For a sample, the probability of not being drawn in a single sampling is (1-1 / n), and the probability of not being drawn after n samplings is (1-1 / n). n , taking the limit we get That is, about 36.8% of the samples in the original dataset D were not sampled.

[0028] Step 3, agent training: Use gradient descent and back propagation methods to train the n datasets generated by n self-sampling. Figure 3 The agent shown.

[0029] The intelligent agent includes: two linear modules consisting of a fully connected layer, a Relu activation function and a Dropout layer, a linear layer and an integer layer, wherein: the input state first passes through the two linear modules, and the input state is multiplied by the weight w of the linear layer to obtain a continuous output. The integer layer converts the continuous output of the linear layer into a discrete integer output, and finally obtains the number of movement steps of the control stick group.

[0030] like Figure 4 As shown, the training uses the Bayesian optimization method to find the optimal hyperparameters, specifically including:

[0031] 1) Define the objective function to be optimized: The objective function to be optimized in this method is the forward propagation process of the agent, and the final output of the function is the loss between the output data and the label. MSE ;

[0032] 2) Initialization of observation points: Select some initial data points through some specific strategies, and preliminarily estimate the basis of the model by observing the output values of these data points;

[0033] 3) Component model: A Gaussian process model is built based on the initial sampling points, taking uncertainty into account and providing a confidence interval for each prediction;

[0034] 4) Select the next parameter: Use the maximum confidence limit strategy or the expected improvement strategy to select the next parameter to be evaluated. This parameter should be the one that the current model believes is most likely to be the global optimal solution.

[0035] 5) Get output: Based on the parameters selected in the previous step, get the output of the objective function at that point;

[0036] 6) Update the model: Add the relationship between the newly observed output and input points to the existing model and re-estimate the model parameters;

[0037] Step 4: Average the outputs of multiple trained agents and output them as the final control rod position to the reactor control system.

[0038] Through specific practical experiments, the generalization ability and accuracy of the model were tested using real operating data from a power plant. The training data used included 10 training sets for three operating conditions: peak shaving, normal startup, and normal shutdown. Two sets of peak shaving data and two sets of normal startup data were used as test sets. The test results are as follows: the red curve represents the actual rod position change process of the nuclear power plant, and the blue curve represents the model control result.

[0039] contrast Figure 6-Figure 9 As can be seen, the multi-agent ensemble learning model achieves high accuracy under both operating conditions. This is because ensemble learning samples the original dataset, potentially generating three different distributions: a large amount of peak-shaving data, a large amount of data from reactor startups and shutdowns, and data with similar amounts of data from the two operating conditions. Based on this, multiple agents are trained. By combining the learning results of multiple agents, the limitation of a single model in covering multiple data distributions is overcome, ultimately effectively improving control accuracy.

[0040] Compared with the existing technology, this method is an offline reinforcement learning method for nuclear power plants: traditional reinforcement learning methods require direct interaction with the environment, and immature intelligent agents may cause accidents in nuclear power plants during control. Training in a simulation environment cannot be transferred from the simulation environment to the real nuclear power plant due to the calculation error between the simulation environment and the real nuclear power plant. The present invention uses the imitation learning method to train the control body using the real historical operation data of the nuclear power plant. While avoiding the control accidents that may be caused by immature intelligent agents, it makes full use of the real historical operation data to improve the reliability of the intelligent agent control; compared with existing control methods that mostly only target a single working condition, the control error surges when the working condition switches. This method adopts an integrated learning framework and randomly samples multi-condition training data. The consistency of multi-condition control reaches 98.4%.

[0041] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principles and purpose of the present invention. The scope of protection of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. All implementation schemes within its scope shall be subject to the constraints of the present invention.

Claims

1. A nuclear power plant core optimization power control method based on imitation learning, characterized in that: include: Step 1: Construct a core optimization power control model based on the imitation learning method; Step 2: Take the average of the trained agent output results as the final control result; The core optimization power control model includes: multiple intelligent agents and an averaging layer, wherein: each intelligent agent learns the historical operating data of the real nuclear power plant through random sampling; The training described above uses Bayesian optimization to find the optimal hyperparameters, reducing the complexity of manual parameter adjustment and improving the accuracy of the model.

2. The method for optimizing power control of a nuclear power plant core based on imitation learning according to claim 1 is characterized in that: The training set used in the training is obtained by sampling the original data set; The data sampling mentioned above refers to: for an original data set D containing n samples, a new data set D is generated by the self-service sampling method. ′ , that is, each time a sample is randomly extracted from the original data set and copied to the new data set D ′ In this way, the sample may still be drawn next time, and a new data set D is obtained after repeating n times. ′ , the number of samples is n, which is the same as the number of samples in the original dataset.

3. The method for optimizing power control of a nuclear power plant core based on imitation learning according to claim 2, wherein: The original data set is obtained in the following way: 1.1 Collect operating data and historical data of the nuclear power plant's distributed control system (DCS). Operating data includes parameters of the primary and secondary circuits of the nuclear power plant, such as primary circuit coolant temperature, steam generator water temperature, main steam pressure transmitter, and control rod position. Historical data includes normal operating conditions, start-up and shutdown conditions, and peak load conditions. 1.2 Data sparse processing: Define the original sampling sequence as: is the original data sampling interval, and 60-second interval resampling is achieved through the following function mapping: In the case of missing sampling point data, linear interpolation method is used to fill the missing data: Where: t a =max{t′ m |t′ m <t},t b =min{t′ m |t′ m >t}; 1.3 Data cleaning, including two types of failure data cleaning: discrete failure and continuous failure. For discrete failure data, linear interpolation is used to fill in missing data. For continuous failure data, the failure time period and number of data points are first recorded, and the data characteristics of the time period are manually analyzed to fill in or discard the missing data. When the data values before and after the failure point do not change much or their rate of change is small, the average value of the data before and after is used to fill the data, specifically: Otherwise, it indicates that the nuclear power plant is in an unsteady-state operating condition, and this period of invalid data is discarded; 1.4 Calculate the initial state (State) and initial decision (Action): Take the 67 reactor parameters and the current position of the 18 rod groups at time T as State, and the difference x between the position of the 18 rod groups at time T+1 and the position at time T T+1 -x T As an Action, to recalculate the entire dataset; 1.5 Data standardization: The Z-Score method is used to standardize the data set, and the original data is processed into standard data with a mean of 0 and a standard deviation of 1. Specifically: in: x old is the data calculated in 1.4, μ is x old The mean of δ is x old The variance of x is n, the number of samples, new The data are obtained after standardization; 1.6 Format conversion: Convert the data type to float32 type to obtain the original data set D containing n samples.

4. The method for optimizing power control of a nuclear power plant core based on imitation learning according to claim 1, wherein: The intelligent agent includes: two linear modules consisting of a fully connected layer, a Relu activation function and a Dropout layer, a linear layer and an integer layer, wherein: the input state first passes through the two linear modules, and the input state is multiplied by the weight w of the linear layer to obtain a continuous output. The integer layer converts the continuous output of the linear layer into a discrete integer output, and finally obtains the number of movement steps of the control stick group.

5. The method for optimizing power control of a nuclear power plant core based on imitation learning according to any one of claims 1 to 4, characterized in that: The training described above uses the Bayesian optimization method to find the optimal hyperparameters, specifically including: 1) Define the objective function to be optimized: The objective function to be optimized in this method is the forward propagation process of the agent, and the final output of the function is the loss between the output data and the label. MSE ; 2) Initialization of observation points: Select some initial data points through some specific strategies, and preliminarily estimate the basis of the model by observing the output values of these data points; 3) Component model: A Gaussian process model is built based on the initial sampling points, taking uncertainty into account and providing a confidence interval for each prediction; 4) Select the next parameter: Use the maximum confidence limit strategy or the expected improvement strategy to select the next parameter to be evaluated. This parameter should be the one that the current model believes is most likely to be the global optimal solution. 5) Get output: Based on the parameters selected in the previous step, get the output of the objective function at that point; 6) Update the model: Add the relationship between the newly observed output and input points to the existing model and re-estimate the model parameters.

6. A nuclear power plant core optimization power control system based on imitation learning that implements the method according to any one of claims 1 to 5, characterized in that: include: A data preprocessing unit, an integrated sampling unit, an agent training unit, and an integrated control unit, wherein: the data preprocessing unit collects historical operating data of the nuclear power plant, performs data sparsification, data cleaning, state action reconstruction, and standardization processing to obtain float32-type tensor data that conforms to deep learning data; the integrated sampling unit performs self-service sampling with replacement on the data preprocessed by the data preprocessing unit to obtain a training subset containing approximately 63.2% of the original data volume; the agent training unit performs multi-agent training on multiple sub-data set results obtained by self-service sampling in the integrated sampling unit to obtain multiple maturely trained agents and perform control; the integrated control unit performs average processing based on the control results of multiple agents in the agent training unit to obtain the final control result.

Citation Information

Patent Citations

  • Automatic searching method for variable-power operation strategy optimization scheme of core unit

    CN107065556A

  • Nuclear power station accident intelligent control method and system

    CN116189945A

  • Integrated optimization system for core design and core control of reactor using reinforcement learning and its method

    JP2005274547A

  • Method and system, having incremental adjustment function, for adjusting control rod of nuclear power unit

    WO2022105356A1

Cited By

  • Nuclear power unit power regulation scheme generation method and system based on deep reinforcement learning

    CN121541495A

  • Nuclear power plant diagnostic model development methods, diagnostic methods, systems, media and equipment

    CN122365038A