An automatic suppression and auxiliary warning prompt method for xenon oscillation based on the 3AO method

By installing high-precision sensors and intelligent algorithms in the nuclear reactor, the control rod and boron concentrations are monitored and automatically adjusted in real time, the reactor instability caused by xenon oscillation is solved, and a safe and stable nuclear reactor operation is achieved.

CN119400472BActive Publication Date: 2025-07-08CNNC NUCLEAR POWER OPERATION MANAGEMENT CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510000121.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-07-08
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

In the prior art, nuclear reactors lack real-time monitoring and automatic early warning when xenon oscillation occurs, and rely on manual adjustment of control rods and boron concentrations, resulting in unstable power of the reactor, posing safety hazards and high operational complexity.

Method used

The xenon oscillation automatic suppression auxiliary early warning method based on the 3AO method is adopted. The core parameters are monitored in real time by installing high-precision sensors, and the xenon oscillation is recognized by using a convolutional neural network. The optimal control rod and boron concentration adjustment are calculated in combination with the deep Q network algorithm, and the operator is automatically prompted to adjust, and the adjustment is performed through the PID controller to form a closed-loop feedback system.

Benefits of technology

It realizes timely identification and automatic suppression of xenon oscillations, reduces human errors, improves the stability and operating efficiency of the reactor, ensures safety and flexibility, and extends the unit life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400472B_ABST
    Figure CN119400472B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically suppressing xenon oscillations and providing auxiliary warning prompts based on the 3AO method, belonging to the technical field of nuclear reactor core monitoring. First, a variety of high-precision sensors are arranged at key positions of the nuclear reactor to timely capture signs of xenon oscillations. By analyzing sensor data, the system can accurately identify early signals of xenon oscillations, providing a basis for formulating subsequent control strategies. The core of the control strategy is the integration of the 3AO method and the deep Q-network algorithm. This algorithm not only comprehensively considers the immediate state of the core but also optimizes control decisions by learning historical data, thereby calculating the optimal control rod position and boron concentration adjustment plan. Subsequently, according to the control instructions generated by the intelligent algorithm, the system automatically prompts the operator to adjust the control rods and boron concentration to respond in real time to the changing requirements of the nuclear reactor. By providing real-time feedback on the adjustment effects, an efficient closed-loop control system is formed, which can continuously improve the accuracy and response speed of the control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent monitoring of nuclear power reactor cores, and specifically discloses a method for automatically suppressing xenon oscillation and providing auxiliary warning based on the 3AO method. Background Art

[0002] Currently, domestic nuclear power units are increasingly frequently adjusting their loads in response to grid requirements. Such frequent load changes pose certain challenges to reactor operation and core safety management. At the same time, load changes at the end of the reactor life may cause xenon oscillations, presenting certain challenges to the control capabilities of operators and carrying the risk of local power distortion in the core due to improper control. For second-generation and second-generation + units that have been in operation for many years, it is not possible to have in-core self-powered detectors or conduct similar technical transformations. To enable the core online monitoring function of third-generation pressurized water reactor units and enhance the unit's control capabilities in responding to load tracking changes, existing xenon oscillation suppression methods mostly rely on manual adjustment of control rods and boron concentration by operators, with slow response speeds, being easily affected by human factors, and having certain potential safety hazards. Summary of the Invention

[0003] The present invention aims to solve the problems of monitoring and warning of xenon oscillations in nuclear reactor operations. In traditional nuclear reactor operations, the occurrence of xenon oscillations often results in unstable reactor power due to the untimely response of manual or semi-automatic control systems, and may trigger safety accidents in severe cases. In addition, in traditional methods, operators need to continuously monitor the reactor status, which not only increases the operation complexity but also raises the risk caused by human errors. Therefore, the present invention needs to solve the following technical problems:

[0004] 1. How to monitor and predict the occurrence of xenon oscillations in real time and accurately so as to make timely responses.

[0005] 2. How to adjust the positions of control rods and boron concentration with auxiliary warning to reduce manual operations and improve the accuracy and efficiency of control.

[0006] 3. How to ensure the safety and stability of the nuclear reactor under different operating conditions and avoid reactor accidents caused by improper control.

[0007] To overcome the deficiencies in the prior art, the present invention provides a method for automatically suppressing xenon oscillation and providing auxiliary warning based on the 3AO method, which can quickly respond when xenon oscillations occur and automatically prompt the operator to adjust the control rods and boron concentration to suppress the occurrence of xenon oscillations and ensure the safe and stable operation of the core. The technical solution of the present invention is as follows:

[0008] A method for automatically suppressing xenon oscillation and providing auxiliary warning based on the 3AO method, comprising the following steps:

[0009] Step 1: Monitor the operating parameters of the reactor core in real time, and use a convolutional neural network to identify the occurrence of xenon oscillations;

[0010] Step 2: Based on the 3AO algorithm, calculate the optimal control rod and boron concentration adjustment scheme to adapt to the dynamic changes under different operating conditions.

[0011] Step 3: According to the instructions of the control module, automatically prompt the operator to adjust the control rod position and boron concentration.

[0012] Step 4: Provide real-time feedback on the adjustment effect, analyze the difference between the adjustment result and the expected goal, and dynamically optimize the subsequent adjustment strategy.

[0013] Furthermore, the specific process of Step 1 is as follows:

[0014] First, install high-precision sensors at key parts of the reactor core to comprehensively monitor the operating state of the reactor core.

[0015] Temperature sensors: Uniformly arrange temperature sensors around the central axis to monitor the temperature change of the reactor core and identify hot spots and temperature anomalies.

[0016] Pressure sensors: Install pressure sensors at the top and bottom of the reactor core, distributed in four quadrants, to monitor the coolant pressure and identify pressure fluctuations and anomalies.

[0017] Neutron flux (power): The RPN power range detector of the off-core nuclear measurement system.

[0018] Coolant flow sensors: Install flow sensors at the inlet and outlet of the coolant to monitor the coolant flow and ensure the stable operation of the cooling system.

[0019] Denoise, normalize, and extract features from the collected original sensor data; use a moving average filter to remove high-frequency noise, perform normalization to eliminate the influence of different dimensions, and then extract features. Extract the time-domain features of each sensor: mean, variance, kurtosis, skewness, maximum value, minimum value; frequency-domain features: spectral energy, main frequency, and time-frequency domain features: wavelet transform coefficients to form a comprehensive feature vector. Then use the sliding window technique to divide the continuous data into multiple overlapping windows, and each window is used as a sample and input into the convolutional neural network model.

[0020]

[0021] Among them, is the time step of the window length, is the th sample, is the data of the time step

[0022] ​The architecture of the convolutional neural network is as follows:

[0023] Input layer: Shape , that is, each sample contains data of time steps, and each time step has

[0024] Convolutional layer 1: 64 (3,3) convolutional kernels, activation function ReLU;

[0025] Pooling layer 1: Max pooling, window size (2,2);

[0026] Convolutional layer 2: 128 (3,3) convolutional kernels, activation function ReLU;

[0027] Pooling layer 2: Max pooling, window size (2,2);

[0028] Fully connected layer 1: 256 nodes, activation function ReLU;

[0029] Output layer: 1 node, activation function Sigmoid, outputting the probability of xenon oscillation occurrence ;

[0030] Loss function and optimization algorithm

[0031] Binary cross-entropy is used as the loss function to measure the difference between the model's predicted probability and the actual label:

[0032]

[0033] Among them, is used as the binary cross-entropy loss function to guide the model optimization, is the number of samples, is the true label of the sample, 0 indicates no xenon oscillation, 1 indicates xenon oscillation occurrence, represents the predicted probability of the model for the th sample.

[0034] The Adam optimizer is adopted, combined with the gradient descent method to adaptively adjust the learning rate:

[0035]

[0036]

[0037]

[0038]

[0039]

[0040] Among them, and are the estimates of the first - order and second - order moments respectively, and are the bias - corrected forms of the first - order and second - order moments respectively, represents the gradient, and are the decay rates of the first - order and second - order moments respectively, is the learning rate, and ϵ is a small constant to prevent division by zero, is the value before parameter update, is the value after parameter update.

[0041] After updating the model parameters, evaluate the model performance on the validation set, adjust the model learning rate according to the validation results, improve the model convergence speed and stability, set a threshold for the real - time monitored data, and predict the probability of xenon oscillation occurrence through a convolutional neural network to predict whether xenon oscillation occurs.

[0042] Furthermore, the specific process of step 2 is as follows:

[0043] Select the fusion of the deep Q - network algorithm and the 3AO method. Using the initial control scheme provided by the 3AO method, the DQN algorithm conducts autonomous learning and policy optimization on this basis, calculates the optimal control rod and boron concentration adjustment scheme, adapts to the dynamic changes under different operating conditions, effectively improves the suppression effect of xenon oscillation, and ensures the safety and stability of core operation.

[0044] The fusion process is as follows:

[0045] Action space initialization

[0046] Use the basic control actions defined by the 3AO method as the initial action set of the reinforcement learning agent. On the basis of the 3AO method, expand the action space of the agent, introduce more refined control actions, and slightly adjust the control rod position and boron concentration to improve the control accuracy.

[0047]

[0048]

[0049]

[0050] Among them, represents the basic control actions defined by the 3AO method, represents the actions expanded by reinforcement learning, is the comprehensive action space, are respectively the control rod rising, holding, and falling, They are respectively the newly added actions: the control rod rises by 0.5%, the control rod drops by 0.5%, the boron concentration increases by 0.1 ppm, the boron concentration remains unchanged, and the boron concentration decreases by 0.1 ppm.

[0051] Policy initialization

[0052] The initial policy of the agent preferentially selects the actions provided by the 3AO method , and gradually explores and expands actions by introducing the ε-greedy policy ; The hybrid control strategy is:

[0053]

[0054] Among them, represents the specific control action to be selected and executed, represents a decision rule, is an action selected from the action space A, such that at the current state the Q-function value is maximized. That is, at the state select the action that can obtain the maximum expected return according to the current value function parameters ;

[0055] Policy optimization process

[0056] Action selection, using -greedy policy, randomly select actions for some time to promote exploration.

[0057]

[0058] Q-network update, using the loss function to measure the difference between the target network and the main Q-network:

[0059]

[0060] Among them, is the loss function of the DQN algorithm, is used to represent the expected value of the variable, is the experience replay buffer, represents a set of transition samples randomly sampled from the experience replay buffer , , is the expectation on the experience replay buffer, is the immediate reward, is the discount factor, is the current network parameter, is the parameter of the target Q-network, represents the next state the maximum Q value under, Indicates the current state Take an action The Q value of;

[0061] Target network synchronization, every certain number of steps, copy the main Q network parameters To the target network: ,

[0062] Experience replay, store the state transition Into the experience replay buffer , and randomly sample a small batch of samples for training. Through multiple rounds of training, the agent gradually optimizes the control strategy.

[0063] The system receives sensor data in real time, including the temperature, pressure, neutron power, and coolant flow rate of the reactor core. The agent learns to select the optimal action in different states through interaction with the environment to maximize the cumulative reward, and the system adjusts the control rod position and boron concentration according to the optimal action. The actions include raising, lowering, or maintaining the current position of the control rod and adjusting the specific value of the boron concentration; the adjustment effect is monitored through the feedback module to dynamically optimize the control strategy.

[0064] Parameters, the algorithm can ensure the smoothness and continuity of the entire adjustment process.

[0065] Furthermore, the specific process of step 3 is as follows:

[0066] Convert the control rod position adjustment amount and boron concentration change amount calculated by the control module into specific execution instructions through the intelligent control algorithm, and determine the specific control rod position adjustment and boron concentration change requirements after parsing the instructions; the intelligent algorithm is responsible for strategy optimization and decision-making, while the PID controller is responsible for specific execution and fine adjustment to ensure the efficient combination of the intelligent algorithm and the execution device, and work together to achieve efficient xenon oscillation suppression.

[0067] Execution instruction generation: After the algorithm receives the optimal execution plan, calculate the specific execution steps and necessary control parameters, and ensure the smoothness and continuity of the entire adjustment process by continuously adjusting Parameters; calculate the real-time control instruction :

[0068]

[0069] Among them, Is the real-time control instruction, Is the system error, that is, the difference between the target position or boron concentration and the current reading; They are the proportional, integral, and derivative coefficients respectively; the PID controller judges and adjusts the position of the control rod and the boron concentration by inputting the current system error e(t), the integral of the error, and the derivative of the error, that is, the PID controller judges and adjusts the position of the control rod and the boron concentration by inputting the current position of the control rod and the current value of the boron concentration.

[0070] Control rod drive: A high-precision servo motor is used to drive the control rod to ensure high-precision and fast response for the adjustment of the control rod position; the servo motor can achieve fine adjustment according to the instructions of the control module to reduce the position error; in addition, the servo motor is also equipped with a position encoder to real-time feedback the position data of the control rod for the feedback module to monitor and optimize;

[0071] Boron concentration adjustment: An intelligent flow control valve is used to control the mixing and distribution of the boron solution to ensure precise adjustment of the boron concentration. The intelligent flow control valve can achieve fast and precise flow control according to the algorithm instructions; the actuator valve is equipped with a flow sensor to real-time monitor the flow rate and concentration of the boron solution and feedback to the control module for dynamic adjustment;

[0072] The PID controller plays a key role in this method of converting the adjustment amount output by the intelligent algorithm, that is, the change amount of the control rod position and the boron concentration, into specific execution instructions. Specifically, the PID controller converts the input data into adjustment amounts through the following workflow:

[0073] Data acquisition: The PID controller receives data from sensors, including the current position of the control rod and the current value of the boron concentration. These data are real-time acquired through a high-precision position encoder and a boron concentration sensor installed on the control rod to ensure the accuracy and timeliness of the data.

[0074] Error calculation: The PID controller compares the current system state with the preset target state to calculate the system error. The system error includes the deviation of the control rod position and the boron concentration.

[0075] Error processing: The PID controller processes the current value, cumulative value, and change rate of the system error according to the three control links of proportional, integral, and derivative: Proportional P makes corresponding adjustments immediately according to the magnitude of the current error to reduce the error; Integral I accumulates past errors to eliminate long-term steady-state errors; Derivative D predicts the change trend of future errors to increase the stability and response speed of the system.

[0076] Adjustment amount calculation: Through the combined action of the three links of data acquisition, error calculation, and error processing, the PID controller calculates precise control instructions , that is, the position adjustment amount of the control rod and the change amount of the boron concentration. These adjustment amounts ensure that the adjustment of the control rod and the boron concentration can accurately respond to the change of the system state and maintain the core operation parameters within the preset target range.

[0077] Execution and feedback: The PID controller sends the calculated control instructions to the servo motor and the intelligent flow control valve to perform specific control rod position adjustment and boron concentration regulation. The real-time feedback data of the execution device is input into the PID controller again to form a closed-loop control, ensuring the continuous optimization and stability of the adjustment process.

[0078] Further, the specific process of step 4 is as follows:

[0079] After the control adjustment is completed, the system will continue to collect the operating parameters of the reactor core, temperature, pressure, neutron flux, and coolant flow rate, as well as the current position of the control rod and the current value of the boron concentration; these data provide the basis for the system to judge the adjustment effect.

[0080] The xenon oscillation monitoring interface provides clear feedback of xenon oscillation information to the operator by displaying the operating parameters and control status of the core in real time. The operator can monitor the real-time status of the core through this interface, identify the occurrence of xenon oscillation in a timely manner, and manually intervene or adjust the control strategy according to the feedback information to ensure the stable and safe operation of the reactor.

[0081] Dynamic optimization adjustment

[0082] Data acquisition: After the control rod position and boron concentration are adjusted, the system continuously monitors and records the operating parameters of the core, namely temperature, pressure, neutron flux, coolant flow rate, as well as the current position of the control rod and the current value of the boron concentration.

[0083] Data comparison: Compare the real-time collected feedback data with the preset target state and calculate the deviation of each parameter.

[0084] Error identification: Identify the source of the deviation through comparative analysis and determine whether the xenon oscillation suppression is sufficient.

[0085] Adjustment suggestion: If it is found that the xenon oscillation suppression is insufficient, the system will suggest further lowering the control rod or increasing the boron injection amount to achieve more stable nuclear chain reaction control.

[0086] System performance evaluation and optimization

[0087] The system calculates the overall performance index of the system through multiple measurement errors during the evaluation period to measure the quality of the adjustment effect.

[0088] Optimization process: According to the latest operating data and feedback results, the system regularly performs the following steps:

[0089] Data update: Input the latest collected operating parameter and control status data into the system.

[0090] Process adjustment: According to the performance evaluation results, adjust the control strategy and operation steps, including the convolutional neural network model and the DQN model, to improve the adaptability and stability of the system.

[0091] Continuous improvement: Through multiple rounds of cycles, continuously optimize the control process to ensure the continuous improvement of the xenon oscillation suppression effect.

[0092] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0093] 1. By installing high-precision sensors at key parts of the nuclear reactor, the system of the present invention can monitor in real time and accurately capture the early signals of xenon oscillation. This instant monitoring function enables the system to quickly identify and predict the occurrence of xenon oscillation before it develops to a level that may affect the stability of the nuclear reactor.

[0094] 2. By analyzing the sensor data collected in real time, the system of the present invention can not only accurately predict xenon oscillation, but also timely prompt the operator to adjust the control strategy, reducing the monitoring and adjustment burden of the operator and minimizing human operation errors caused by fatigue or monitoring mistakes. The operator can focus more on system monitoring and decision support rather than daily parameter adjustment, thus improving the overall operation efficiency and safety.

[0095] 3. By integrating the 3AO method and the deep Q-network algorithm, the present invention realizes the automatic adjustment prompt for the control rod position and boron concentration. Based on real-time data and historical operation experience, this intelligent algorithm dynamically calculates the optimal control strategy and automatically adjusts the operating parameters of the nuclear reactor to effectively suppress xenon oscillation. The closed-loop feedback system in this process further enhances the accuracy and response speed of the control strategy, ensuring the stability of the nuclear reactor under different operating conditions. This stable operating state helps protect the core components of the nuclear reactor, reduces wear and aging, and thus extends the life of the unit and improves economic benefits. In addition, in the face of changing load conditions, the system can automatically adjust parameters to maintain the best operating efficiency and safety, enhancing the reliability and flexibility of the unit under unstable grid conditions. Description of the Drawings

[0096] Figure 1 It is a schematic diagram of the xenon oscillation monitoring interface;

[0097] Figure 2 It is a schematic diagram of the early warning prompt;

[0098] Figure 3 It is a flow chart of the xenon oscillation suppression execution;

[0099] Figure 4 It is a flow chart of the automatic xenon oscillation suppression method. Detailed Embodiments

[0100] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0101] As Figures 1-4 shown, a method for automatically suppressing xenon oscillation and giving auxiliary warning prompts based on the 3AO method includes the following steps:

[0102] Step 1: Monitor the operating parameters of the reactor core in real time and identify the occurrence of xenon oscillation.

[0103] Step 2: Calculate the optimal control rod and boron concentration adjustment scheme based on the 3AO method algorithm to adapt to the dynamic changes under different operating conditions.

[0104] Step 3: Automatically prompt the operator to adjust the control rod position and boron concentration according to the instructions of the control module.

[0105] Step 4: Feedback the adjustment effect in real time, analyze the difference between the adjustment result and the expected target, and dynamically optimize the subsequent adjustment strategy.

[0106] The specific process of the above Step 1 is as follows:

[0107] First, install high-precision sensors at key parts of the reactor core to comprehensively monitor the operating state of the reactor core.

[0108] Temperature sensors: Uniformly arrange temperature sensors around the central axis to monitor the temperature change of the reactor core and identify the hot spot area and temperature anomalies.

[0109] Pressure sensors: Install pressure sensors at the top and bottom of the reactor core, distributed in four quadrants, to monitor the coolant pressure and identify pressure fluctuations and anomalies.

[0110] Neutron flux: The RPN power range detector of the out-of-core nuclear measurement system measures the distribution of neutrons in the reactor core and the power output, and identifies power fluctuations.

[0111] Coolant flow sensors: Install flow sensors at the inlet and outlet of the coolant to monitor the coolant flow and ensure the stable operation of the cooling system.

[0112] Denoise, normalize, and extract features from the collected raw sensor data; use a moving average filter to remove high-frequency noise, perform normalization to eliminate the influence of different dimensions, and then extract features. Extract the time-domain features of each sensor: mean, variance, kurtosis, skewness, maximum value, and minimum value; frequency-domain features: spectral energy, main frequency, and time-frequency domain features: wavelet transform coefficients to form a comprehensive feature vector. Then, use the sliding window technique to divide the continuous data into multiple overlapping windows, and each window is used as a sample and input into the convolutional neural network model.

[0113]

[0114] Among them, is the time step of the window length, is the th sample, is the data of the time step .

[0115] The convolutional neural network architecture is as follows:

[0116] Input layer: shape , that is, each sample contains data of time steps, and each time step has sensor data.

[0117] Convolutional layer 1: 64 (3,3) convolutional kernels, activation function ReLU;

[0118] Pooling layer 1: max pooling, window size (2,2);

[0119] Convolutional layer 2: 128 (3,3) convolutional kernels, activation function ReLU;

[0120] Pooling layer 2: max pooling, window size (2,2);

[0121] Fully connected layer 1: 256 nodes, activation function ReLU;

[0122] Output layer: 1 node, activation function Sigmoid, output the probability of xenon oscillation occurrence ;

[0123] Loss function and optimization algorithm

[0124] Use binary cross-entropy as the loss function to measure the difference between the model prediction probability and the actual label:

[0125]

[0126] Among them, is used as the binary cross-entropy loss function to guide the model optimization, is the number of samples, is the true label of the sample, where 0 indicates no xenon oscillation and 1 indicates xenon oscillation occurs, represents the predicted probability of the model for the th sample.

[0127] The Adam optimizer is adopted, and the learning rate is adaptively adjusted in combination with the gradient descent method:

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Among them, and are the estimates of the first-order and second-order moments respectively, and are the bias-corrected forms of the first-order and second-order moments respectively, represents the gradient, and are the decay rates of the first-order and second-order moments respectively, is the learning rate, and ϵ is a small constant to prevent division by zero, is the value before parameter update, is the value after parameter update.

[0134] After updating the model parameters, evaluate the model performance on the validation set, adjust the model learning rate according to the validation results, improve the model convergence speed and stability, output the probability of xenon oscillation occurrence, and predict whether xenon oscillation occurs.

[0135] The specific process of the above step 2 is as follows:

[0136] Select the fusion of the deep Q-network algorithm and the 3AO method. Using the initial control scheme provided by the 3AO method, the DQN algorithm conducts autonomous learning and policy optimization on this basis, calculates the optimal control rod and boron concentration adjustment scheme, adapts to the dynamic changes under different operating conditions, effectively improves the suppression effect of xenon oscillation, and ensures the safety and stability of the reactor core operation.

[0137] The fusion process is as follows:

[0138] Action space initialization

[0139] The basic control actions defined by the 3AO method are used as the initial action set of the reinforcement learning agent. Based on the 3AO method, the action space of the agent is expanded, and more refined control actions are introduced to finely adjust the control rod position and boron concentration to improve the control accuracy.

[0140]

[0141]

[0142]

[0143] Among them, represents the basic control actions defined by the 3AO method, represents the actions extended by the reinforcement learning, is the comprehensive action space, are respectively the control rod rising, holding, and dropping, are respectively the newly added actions: the control rod rising by 0.5%, the control rod dropping by 0.5%, the boron concentration increasing by 0.1 ppm, the boron concentration remaining unchanged, and the boron concentration decreasing by 0.1 ppm.

[0144] Policy initialization

[0145] The initial policy of the agent preferentially selects the actions provided by the 3AO method , and gradually explores the extended actions by introducing the ε-greedy policy ; the hybrid control policy is:

[0146]

[0147] Among them, represents the specific control action to be selected and executed, represents a decision rule, is an action selected from the action space A, such that under the current state , the Q-function value is maximized. That is, under the state , select the action that can obtain the maximum expected return according to the current value function parameter evaluation;

[0148] Policy optimization process

[0149] Action selection, using -greedy policy, select random actions for some time to promote exploration.

[0150]

[0151] Q-network update, use the loss function to measure the difference between the target network and the main Q-network:

[0152]

[0153] Among them, is the loss function of the DQN algorithm, used to represent the expected value of the variable, is the experience replay buffer, represents a set of transition samples randomly sampled from the experience replay buffer and , is the expectation on the experience replay buffer, is the immediate reward, is the discount factor, is the current network parameter, is the parameter of the target Q network, represents the next state under the maximum Q value, represents the current state under the action taken;

[0154] Target network synchronization, every certain number of steps, copy the main Q network parameter to the target network: ,

[0155] Experience replay, store the state transition into the experience replay buffer , and randomly draw a small batch of samples for training. Through multiple rounds of training, the agent gradually optimizes the control strategy.

[0156] The system receives sensor data in real time, including the temperature, pressure, neutron power, and coolant flow rate of the reactor core. The agent learns to select the optimal action in different states through interaction with the environment to maximize the cumulative reward, and the system adjusts the control rod position and boron concentration according to the optimal action. The actions include raising, lowering, or maintaining the current position of the control rod and adjusting the specific value of the boron concentration; the adjustment effect is monitored through the feedback module to dynamically optimize the control strategy.

[0157] With these parameters, the algorithm can ensure the smoothness and continuity of the entire adjustment process.

[0158] The specific process of step 3 above is as follows:

[0159] The position adjustment amount of the control rod and the change amount of the boron concentration calculated by the control module are converted into specific execution instructions through an intelligent control algorithm. After parsing the instructions, the specific requirements for the control rod position adjustment and boron concentration change are determined. The intelligent algorithm is responsible for strategy optimization and decision-making, while the PID controller is responsible for specific execution and fine adjustment to ensure the efficient combination of the intelligent algorithm and the execution device, and they work together to achieve efficient xenon oscillation suppression.

[0160] Execution instruction generation: After the algorithm receives the optimal execution plan, it calculates the specific execution steps and necessary control parameters. By continuously adjusting parameters, the algorithm can ensure the smoothness and continuity of the entire adjustment process; calculate real-time control instructions :

[0161]

[0162] Among them, is the real-time control instruction, is the system error, that is, the difference between the target position or boron concentration and the current reading; are the proportional, integral, and differential coefficients respectively; the PID controller judges and adjusts the position of the control rod and the boron concentration by inputting the current system error e(t), the integral of the error, and the differential of the error. That is, the PID controller judges and adjusts the position of the control rod and the boron concentration by inputting the current position of the control rod and the current value of the boron concentration.

[0163] Control rod drive: A high-precision servo motor is used to drive the control rod to ensure high-precision and fast response for the control rod position adjustment; the servo motor can achieve fine adjustment according to the instructions of the control module to reduce the position error; in addition, the servo motor is also equipped with a position encoder to real-time feedback the position data of the control rod for the feedback module to monitor and optimize;

[0164] Boron concentration adjustment: An intelligent flow control valve is used to control the mixing and distribution of the boron solution to ensure precise adjustment of the boron concentration. The intelligent flow control valve can achieve fast and precise flow control according to the algorithm instructions; the execution valve is equipped with a flow sensor to real-time monitor the flow rate and concentration of the boron solution and feedback to the control module for dynamic adjustment;

[0165] The PID controller plays a key role in converting the adjustment amount output by the intelligent algorithm, that is, the change amount of the control rod position and the boron concentration, into specific execution instructions in this method. Specifically, the PID controller converts the input data into the adjustment amount through the following workflow:

[0166] Data acquisition: The PID controller receives data from sensors, including the current position of the control rod and the current value of boron concentration. This data is obtained in real time through a high-precision position encoder and a boron concentration sensor installed on the control rod to ensure the accuracy and timeliness of the data.

[0167] Error calculation: The PID controller compares the current system state with the preset target state and calculates the system error. The system error includes the deviation of the control rod position and boron concentration.

[0168] Error processing: The PID controller processes the current value, cumulative value, and rate of change of the system error according to the three control links of proportional, integral, and derivative: The proportional P makes corresponding adjustments immediately according to the magnitude of the current error to reduce the error; the integral I accumulates past errors to eliminate long-term steady-state errors; the derivative D predicts the change trend of future errors to increase the stability and response speed of the system.

[0169] Calculation of adjustment amount: Through the combined action of data acquisition, error calculation, and error processing, the PID controller calculates precise control instructions , that is, the position adjustment amount of the control rod and the change amount of boron concentration. These adjustment amounts ensure that the adjustment of the control rod and boron concentration can accurately respond to the change of the system state and maintain the core operation parameters within the preset target range.

[0170] Execution and feedback: The PID controller sends the calculated control instructions to the servo motor and the intelligent flow control valve to perform specific control rod position adjustment and boron concentration regulation. The real-time feedback data of the execution device is input into the PID controller again to form a closed-loop control, ensuring the continuous optimization and stability of the adjustment process.

[0171] The specific process of step 4 above is:

[0172] The specific process of step 4 is:

[0173] After the control adjustment is completed, the system will continue to collect the operation parameters of the reactor core, temperature, pressure, neutron flux, and coolant flow rate, as well as the current position of the control rod and the current value of boron concentration; these data provide a basis for the system to judge the adjustment effect.

[0174] The xenon oscillation monitoring interface provides clear feedback of xenon oscillation information to the operator by displaying the operation parameters and control status of the core in real time. The operator can monitor the real-time state of the core through this interface, identify the occurrence of xenon oscillation in a timely manner, and manually intervene or adjust the control strategy according to the feedback information to ensure the stable and safe operation of the reactor.

[0175] Dynamic optimization adjustment

[0176] Data Acquisition: After the control rod position and boron concentration are adjusted, the system continuously monitors and records the operating parameters of the core, namely temperature, pressure, neutron flux, coolant flow rate, and the current position of the control rods and the current value of the boron concentration.

[0177] Data Comparison: Compare the real-time acquired feedback data with the preset target state and calculate the deviation of each parameter.

[0178] Error Identification: Identify the source of deviation through comparative analysis and determine whether the xenon oscillation suppression is sufficient.

[0179] Adjustment Suggestion: If it is found that the xenon oscillation suppression is insufficient, the system will suggest further lowering the control rods or increasing the boron injection amount to achieve more stable control of the nuclear chain reaction.

[0180] System Performance Evaluation and Optimization

[0181] The system calculates the overall performance index of the system by evaluating the multiple measurement errors within the evaluation period to measure the quality of the adjustment effect.

[0182] Optimization Process: According to the latest operating data and feedback results, the system regularly performs the following steps:

[0183] Data Update: Input the latest acquired operating parameters and control status data into the system.

[0184] Process Adjustment: Adjust the control strategy and operation steps according to the performance evaluation results, including the convolutional neural network model and the DQN model, to improve the adaptability and stability of the system.

[0185] Continuous Improvement: Through multiple rounds of cycles, continuously optimize the control process to ensure continuous improvement of the xenon oscillation suppression effect.

[0186] Numerical Example:

[0187] Sensor Installation

[0188] Temperature Sensor: Install platinum resistance temperature detectors at 6 positions around the central axis of the core, one every 60 degrees, to monitor the temperature distribution inside the core.

[0189] Pressure Sensor: Install 4 piezoelectric pressure sensors at the top and bottom of the core, distributed in four quadrants, to monitor the change of coolant pressure in real time.

[0190] Neutron Flux Density Measurement: Install an ex-core nuclear measurement system RPN power range detector outside the core to monitor the distribution of neutrons in the core and the power output.

[0191] Coolant Flow Sensor: Install 2 vortex flow meters at the inlet and outlet of the coolant to monitor the change of coolant flow rate.

[0192] Data Acquisition and Processing

[0193] Sampling Frequency: Set to 10 times per second (10 Hz) to ensure high-frequency monitoring.

[0194] Denoising: Apply a moving average filter with a window size of N = 5 seconds to smooth the sensor data.

[0195]

[0196] Normalization: Standardize each sensor data and calculate the mean and standard deviation .

[0197] Extract the mean, variance, kurtosis, skewness, maximum, minimum, spectral energy, and wavelet transform coefficients of each sensor to form a comprehensive feature vector.

[0198] Data Segmentation: Use the sliding window technique with a window length time step and a step size , generating 10,000 samples for subsequent model training.

[0199]

[0200] CNN Model Training

[0201] The convolutional neural network architecture is as follows:

[0202] Input Layer: Each sample contains data of time steps, and each time step has

[0203] Sensor data.

[0204] Convolutional Layer 1: 64 (3,3) convolutional kernels with the ReLU activation function;

[0205] Pooling Layer 1: Max pooling with a window size of (2,2);

[0206] Convolutional Layer 2: 128 (3,3) convolutional kernels with the ReLU activation function;

[0207] Pooling Layer 2: Max pooling with a window size of (2,2);

[0208] Fully Connected Layer 1: 256 nodes with the ReLU activation function; ;

[0209] Training Parameters

[0210] Loss function: Binary cross-entropy

[0211] Optimization algorithm: Adam, learning rate

[0212] Batch size: 64

[0213] Number of training epochs: 50

[0214] Target network update frequency: Update once every 1000 steps

[0215] Model optimization: Regularly update model parameters, retrain using new operating data, and predict the probability of xenon oscillation occurrence , improving the generalization ability and prediction accuracy of the model.

[0216] CNN output: The trained CNN model outputs a probability of xenon oscillation occurrence for each input sample , whose value ranges from 0 to 1. This probability is used to determine whether xenon oscillation is likely to occur and serves as the basis for subsequent control strategies.

[0217] Due to the performance of xenon oscillation: The axial power offset AOref exceeds the normal operating range of ±3% FP. Under normal conditions, the average core temperature is 309.5 degrees Celsius, the pressure is 15.48 MP, and AOref ±3 is the safety control range (coolant flow rate is not an important reference in actual applications). The local temperature rise and local pressure increase in the core, and the deviation degree of AOref exceeding ±3 (this indicator is the most important) determines the probability of xenon oscillation occurrence: Through the local temperature rise and local pressure increase in the core, when AOref deviates from the normal value by 2%, 2.5%, 3%, 3.5%, 4%, it corresponds to the CNN outputs of 0.2, 0.4, 0.6, 0.8, 1 respectively); Sample A: After the input data is processed by the CNN model, the output , indicating an 85% probability of xenon oscillation occurrence. Sample B: After the input data is processed by the CNN model, the output , indicating a 30% probability of xenon oscillation occurrence. Sample C: After the input data is processed by the CNN model, the output , indicating a 60% probability of xenon oscillation occurrence. Based on these probability values, the system can set a threshold of 50%. When exceeds this threshold, the control rods and boron concentration are automatically adjusted to suppress the occurrence of xenon oscillation.

[0218] DQN model training

[0219] Environment settings:

[0220] State space (S): Contains 20 sensor data (temperature, pressure, neutron flux, coolant flow rate) and 5 control variables (positions of 4 control rods and boron concentration).

[0221] Initial action set: Based on the 3AO method , each control rod has 3 possible actions (rise by 1%, hold, fall by 1%)

[0222] Expanded action set: The reinforcement learning algorithm adds more refined control actions , the control rod has 2 possible actions (rise by 0.5%, fall by 0.5%), and the boron concentration has 3 actions (increase by 0.1 ppm, hold, decrease by 0.1 ppm).

[0223] Comprehensive action space (A): , the boron concentration has 3 actions (increase by 0.1 ppm, hold, decrease by 0.1 ppm). The total number of actions is .

[0224] Reward function (R): Considering power fluctuations, safety parameters, and control energy consumption comprehensively, different weight coefficients are set.

[0225] Reward mechanism: The reward function comprehensively considers power fluctuations and safety parameters (the axial power offset AOref exceeds the normal operating range of ±3% FP), and sets different weight coefficients to guide the intelligent agent to optimize the control strategy.

[0226] Network structure:

[0227] Input layer: Shape (20), that is, 20 sensor data.

[0228] Hidden layer: Two fully connected layers, with 128 and 64 nodes, and the activation function ReLU.

[0229] Output layer: 6561 nodes, corresponding to the Q values of each action in the action space.

[0230] Training parameters:

[0231] Learning rate (η): 0.001

[0232] Discount factor (γ): 0.99

[0233] Size of the experience replay buffer: 100,000

[0234] Batch size: 64

[0235] Update frequency of the target network: Update once every 1000 steps

[0236] Exploration strategy: ε-greedy strategy, initial = 1.0, gradually reduced to 0.1.

[0237] Algorithm Fusion and Deployment: The 3AO method fuses the DQN algorithm, and deploys the trained model into the real-time monitoring system to receive sensor data in real time, output the optimal control actions, that is, the adjustment amounts of control rods and boron concentration, and monitor the control effect and system stability. The output control actions are as follows:

[0238] (1) Under normal operation: When the system runs stably and the xenon oscillation probability output by the CNN model is lower than the set threshold, the DQN model will choose to maintain the current control rod position and boron concentration without adjustment.

[0239] (2) Responding to xenon oscillation: Based on the prediction results output by the CNN, set the predicted xenon oscillation threshold at 50%. If the predicted xenon oscillation probability exceeds 50%, automatic adjustment will be executed.

[0240] Control rod adjustment strategy: The control rod drops by 0.5% to inhibit the reaction rate and mitigate the impact of xenon oscillation.

[0241] Boron concentration adjustment strategy: The boron concentration increases by 0.1 ppm to absorb excess neutrons and further stabilize the nuclear reaction.

[0242] (3) Routine maintenance adjustment: During routine maintenance inspections, the system detects a slight deviation between the position of the control rod and the preset optimal position.

[0243] Control rod adjustment strategy: Automatically fine-tune it back to the best position, and the adjustment amplitude is .

[0244] Boron concentration adjustment strategy: Boron concentration adjustment .

[0245] (4) Quick response to abnormal situations: The sensor suddenly detects an abnormally high neutron flux, indicating a potential safety risk.

[0246] Control rod adjustment strategy: Immediately reduce the reactor power, and the system quickly and continuously lowers the control rod by 0.5%.

[0247] Boron concentration adjustment strategy: To quickly stabilize the reactor, the system increases the boron concentration .

[0248] Adjustment strategy

[0249] a. Data acquisition: The PID controller receives data from sensors, including the current position of the control rod and the current value of boron concentration. These data are obtained in real time through high-precision position encoders installed on the control rod and boron concentration sensors to ensure the accuracy and timeliness of the data.

[0250] b. Error calculation: The PID controller compares the current system state with the preset target state and calculates the system error. The system error includes the deviations in the control rod position and boron concentration.

[0251] c. Error processing:

[0252] Proportional (P): According to the magnitude of the current error, corresponding adjustments are made immediately to reduce the error.

[0253] Integral (I): Accumulates past errors to eliminate long - term steady - state errors.

[0254] Derivative (D): Predicts the changing trend of future errors, increasing the stability and response speed of the system.

[0255] d. Adjustment amount calculation: Through the combined action of the above three links, the PID controller calculates precise control instructions, namely the position adjustment amount of the control rod and the change amount of boron concentration. These adjustment amounts ensure that the adjustments of the control rod and boron concentration can accurately respond to the changes in the system state and maintain the core operation parameters within the preset target range.

[0256] e. Execution and feedback: The PID controller sends the calculated control instructions to the servo motor and intelligent flow control valve to perform specific control rod position adjustment and boron concentration regulation. The real - time feedback data of the execution device is input into the PID controller again to form a closed - loop control, ensuring the continuous optimization and stability of the adjustment process.

[0257] Collaborative work of intelligent algorithms and PID controllers: The 3AO method integrated with the DQN algorithm is responsible for strategy optimization and decision - making, providing preliminary control instructions (control rod position adjustment amount and boron concentration change amount). The PID controller then finely adjusts these instructions based on the real - time feedback data to ensure the accuracy and stability of the adjustment process. The two work together to achieve efficient suppression of xenon oscillation and ensure the safe and stable operation of the nuclear reactor.

[0258] Dynamic optimization adjustment

[0259] Error calculation: For the measured value of the temperature sensor in the input parameters, calculate its error from the target. Since the xenon oscillation cannot be completely suppressed, the system will analyze this error and recommend further adjustment of the control rod.

[0260] Performance evaluation: , Close to 1 indicates that the system control accuracy is very high, the error is very small, almost reaching the ideal state. An evaluation value between 0.8 and 0.9 indicates that the control effect of the system is good, the error is small, and most safety and efficiency requirements are met. An evaluation value between 0.7 and 0.8, although the control effect is not the most ideal, is still within an acceptable range, and the system needs to be further optimized to improve performance. An evaluation value below 0.7 usually means that there is a large room for improvement in the system's control strategy, the error is large, and it may be necessary to reconsider the control logic or parameter adjustment.

[0261] According to the latest operation data and feedback results collected, the control algorithm is updated regularly. The parameters of the convolutional neural network and the DQN model are adjusted to improve the prediction accuracy and control effect.

[0262] Finally, it should be noted that: in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0263] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An automatic suppression and auxiliary warning prompt method for xenon oscillation based on the 3AO method, characterized in that It includes the following steps: Step 1: Monitor the operating parameters of the reactor core in real time, and use a convolutional neural network to identify the occurrence of xenon oscillations; Step 2: Based on the 3AO method algorithm, calculate the optimal control rod and boron concentration adjustment scheme to adapt to the dynamic changes under different operating conditions; Step 3: According to the instructions of the control module, automatically give a warning to prompt the operator to adjust the control rod position and boron concentration; Step 4: Provide real-time feedback on the adjustment effect, analyze the difference between the adjustment result and the expected target, and dynamically optimize the subsequent adjustment strategy; The specific process of Step 1 is as follows: When xenon oscillations occur, they will cause local temperature and power increase in the reactor core, which may further lead to the problem of fuel element melting. First, install high-precision sensors at key parts of the reactor core to comprehensively monitor the operating state of the reactor core; Temperature sensors: Uniformly arrange temperature sensors around the central axis to monitor the temperature change of the reactor core and identify the hot spot area and temperature anomalies; Pressure sensors: Install pressure sensors at the top and bottom of the reactor core, distributed in four quadrants, to monitor the coolant pressure and identify pressure fluctuations and anomalies; Neutron flux: The out-of-core nuclear measurement system RPN power range detector; Coolant flow sensors: Install flow sensors at the inlet and outlet of the coolant to monitor the coolant flow and ensure the stable operation of the cooling system; Denoise, normalize, and extract features from the collected original sensor data; Use a moving average filter to remove high-frequency noise, perform normalization to eliminate the influence of different dimensions, and then extract features. Extract the time-domain features of each sensor: mean, variance, kurtosis, skewness, maximum value, minimum value; Frequency domain features: spectral energy, main frequency, and time-frequency domain features: wavelet transform coefficients are used to form a comprehensive feature vector. The sliding window technique is adopted to divide continuous data into multiple overlapping windows, with each window X i being input into a convolutional neural network model as a sample; X i = {x i , x i+1 , …, x i+T-1} where T is the time step of the window length, and X i is the i-th sample, and x i is the data at time step i; The convolutional neural network architecture is as follows: Input layer: Shape (T, S), that is, each sample contains data of T time steps, and each time step has S sensor data; Convolutional layer 1: 64 (3, 3) convolutional kernels, activation function ReLU; Pooling layer 1: Max pooling, window size (2, 2); Convolutional layer 2: 128 (3, 3) convolutional kernels, activation function ReLU; Pooling layer 2: Max pooling, window size (2, 2); Fully connected layer 1: 256 nodes, activation function ReLU; Output layer: 1 node, activation function Sigmoid, output the probability of xenon oscillation occurrence Loss function and optimization algorithm Use binary cross-entropy as the loss function to measure the difference between the model prediction probability and the actual label: Among them, As the binary cross-entropy loss function guides the model optimization, N is the number of samples, and y i is the true label of the sample, 0 indicates no xenon oscillation, 1 indicates xenon oscillation occurs, represents the predicted probability of the model for the i-th sample; Use the Adam optimizer and combine the gradient descent method to adaptively adjust the learning rate: m t = β1m t-1 + (1 - β1)g t where m t and v t are the estimates of the first and second moments respectively, and are the bias-corrected forms of the first and second moments respectively, g t denotes the gradient, β1 and β2 are the decay rates of the first and second moments respectively, η is the learning rate, ∈ is a small constant to prevent division by zero, θ t-1 is the value of the parameter before update, θ t is the value of the parameter after update; is the decay factor of the first moment, is the decay factor of the second moment; After updating the model parameters, evaluate the model performance on the validation set, adjust the model learning rate according to the validation results, improve the model convergence speed and stability, set a threshold for the real-time monitored data, and predict the probability of xenon oscillations occurring through the convolutional neural network to predict whether xenon oscillations occur.

2. The xenon oscillation automatic suppression auxiliary warning prompt method based on the 3AO method according to claim 1, characterized in that The specific process of Step 2 is as follows: Select the fusion of the deep Q network algorithm and the 3AO method. Use the initial control scheme provided by the 3AO method, and the DQN algorithm performs autonomous learning and policy optimization on this basis to calculate the optimal control rod and boron concentration adjustment scheme to adapt to the dynamic changes under different operating conditions, effectively improve the suppression effect of xenon oscillations, and ensure the safety and stability of the reactor core operation; The fusion process is as follows: Action space initialization: The basic control actions defined by the 3AO method are used as the initial action set of the reinforcement learning agent. Based on the 3AO method, the action space of the agent is expanded, and more refined control actions are introduced to finely adjust the control rod position and boron concentration to improve the control accuracy. A 3AO = {a1, a2, a3} A RL = {a4, a5, a6, a7, a8} A = A 3AO ∪A RL = {a1, a2, a3, a4, a5, a6, a7, a8} Among them, A 3AO represents the basic control actions defined by the 3AO method, and A RL represents the actions extended by reinforcement learning. A is the comprehensive action space. a1, a2, and a3 are respectively the control rod rising, holding, and descending, and a4, a5, a6, a7, and a8 are respectively the newly added actions; Policy initialization: The initial policy of the agent preferentially selects the actions a1, a2, a3 provided by the 3AO method. By introducing the ε-greedy policy, the expanded actions a4, a5, a6, a7, a8 are gradually explored. The hybrid control policy is as follows: Among them, arg max represents the independent variable that makes the function obtain the maximum value, a t represents the specific control action to be selected and executed, represents a decision rule, a i is an action selected from the action space A, such that at the current state s t the Q - function value is maximized, that is, at state s t select the action that obtains the maximum expected return according to the evaluation of the current value function parameter θ; Policy optimization process: Action selection, using the ε-greedy policy, randomly selects an action a t ∈ A sometimes to promote exploration; The Q-network is updated, and the loss function is used to measure the difference between the target network and the main Q-network: Among them, is the loss function of the DQN algorithm, and E is used to represent the expected value of a variable. is the experience replay buffer. represents a set of transition samples (s randomly sampled from the experience replay buffer t , a t , r t , s t+1 ). is the expectation on the experience replay buffer, r is the immediate reward, γ is the discount factor, θ is the current network parameter, and θ - is the parameter of the target Q-network. represents the maximum Q-value under the next state s t+1 , and Q(s, a; θ) represents the Q-value of taking the action a t under the current state s t . Target network synchronization: Every certain number of steps, copy the main Q-network parameters θ to the target network: θ - ← θ, Experience replay, storing the state transition (s t , a t , r t , s t+1 ) into the experience replay buffer Randomly sampling a small batch for training; Through multiple rounds of training, the agent gradually optimizes the control strategy; The system receives sensor data in real time, including the temperature, pressure, neutron power, and coolant flow rate of the reactor core. Through the interaction with the environment, the agent learns to select the optimal action in different states to maximize the cumulative reward. The system adjusts the control rod position and boron concentration according to the optimal action. The actions include raising, lowering, or maintaining the current position of the control rod and adjusting the specific value of the boron concentration. The adjustment effect is monitored through the feedback module to dynamically optimize the control strategy.

3. A method for automatically suppressing xenon oscillation and providing auxiliary warning prompts based on the 3AO method according to claim 2, characterized in that, a4, a5, a6, a7, a8 are respectively: the control rod is raised by 0.5%, the control rod is lowered by 0.5%, the boron concentration is increased by 0.1 ppm, the boron concentration is maintained, and the boron concentration is decreased by 0.1 ppm.

4. The xenon oscillation automatic suppression auxiliary warning and prompt method based on the 3AO method according to claim 2, characterized in that, The control actions output by the agent gradually optimizing the control strategy are as follows: 1) Under normal operation: When the system is operating stably and the xenon oscillation probability output by the convolutional neural network (CNN) model is lower than the set threshold, the DQN model will choose to maintain the current control rod position and boron concentration without adjustment. 2) Responding to xenon oscillation: Based on the prediction result output by the CNN, a predicted xenon oscillation threshold of 50% is set. If the predicted xenon oscillation probability exceeds 50%, the adjustment will be automatically executed. Control rod adjustment strategy: The control rod is lowered by 0.5% to inhibit the reaction rate and mitigate the impact of xenon oscillation. Boron concentration adjustment strategy: The boron concentration is increased by 0.1 ppm to absorb excess neutrons and further stabilize the nuclear reaction.

5. A method for automatically suppressing xenon oscillation and assisting in early warning prompt based on the 3AO method according to claim 1, characterized in that, The specific process of step 3 is as follows: The position adjustment amount of the control rod and the change amount of the boron concentration calculated by the control module are converted into specific execution instructions through the intelligent control algorithm. After parsing the instructions, the specific control rod position adjustment and boron concentration change requirements are determined. The intelligent algorithm is responsible for policy optimization and decision-making, while the PID controller is responsible for specific execution and fine adjustment to ensure the efficient combination of the intelligent algorithm and the execution device and work together to achieve efficient xenon oscillation suppression. Execution instruction generation: After the algorithm receives the optimal execution plan, it calculates the specific execution steps and necessary control parameters, and continuously adjusts K p ,K i ,K d parameters, the algorithm can ensure the smoothness and continuity of the entire adjustment process; calculate the real-time control instruction u t : e(τ) is a variable in the integration process; τ is the integration variable, representing the time interval from the initial time 0 to the current time τ. When τ is equal to t, e(τ) = e(t); where, u t is the real-time control instruction, e(t) is the system error, that is, the difference between the target position or boron concentration and the current reading; K p , K i , K d are the proportional, integral, and differential coefficients respectively; the PID controller determines and adjusts the position of the control rod and the boron concentration by inputting the current system error e(t), the integral of the error, and the differential of the error, that is, the PID controller determines and adjusts the position of the control rod and the boron concentration by inputting the current position of the control rod and the current value of the boron concentration; Control rod drive: A high-precision servo motor is used to drive the control rod to ensure high-precision and fast response for the control rod position adjustment. The servo motor can achieve fine adjustment according to the instructions of the control module to reduce the position error. In addition, the servo motor is equipped with a position encoder to real-time feedback the position data of the control rod for the feedback module to monitor and optimize. Boron concentration regulation: Use an intelligent flow control valve to control the mixing and distribution of the boron solution to ensure precise regulation of the boron concentration. The intelligent flow control valve can achieve fast and precise flow control according to algorithm instructions; the actuator valve is equipped with a flow sensor to monitor the flow rate and concentration of the boron solution in real time and feed back to the control module for dynamic adjustment.

6. The automatic suppression and auxiliary warning prompt method for xenon oscillation based on the 3AO method according to claim 5, characterized in that, The PID controller converts the input data into a regulating quantity through the following workflow: Data acquisition: The PID controller receives data from sensors, including the current position of the reactor control rod and the current value of the boron concentration; this data is obtained in real time through a high-precision position encoder installed on the control rod and a boron concentration sensor to ensure the accuracy and timeliness of the data; Error calculation: The PID controller compares the current system state with the preset target state and calculates the system error; the system error includes the deviation of the control rod position and the boron concentration; Error processing: The PID controller processes the current value, cumulative value, and rate of change of the system error according to the three control links of proportional, integral, and derivative: Proportional P makes corresponding adjustments immediately according to the magnitude of the current error to reduce the error; Integral I accumulates past errors to eliminate long-term steady-state errors; Derivative D predicts the change trend of future errors to increase the stability and response speed of the system; Adjustment quantity calculation: Through the combined effect of three links: data acquisition, error calculation, and error processing, the PID controller calculates the precise control command u t , that is, the position adjustment quantity of the control rod and the change quantity of the boron concentration; These regulating quantities ensure that the adjustment of the control rod and boron concentration can accurately respond to changes in the system state and maintain the core operating parameters within the preset target range; Execution and feedback: The PID controller sends the calculated control instructions to the servo motor and the intelligent flow control valve to perform specific control rod position adjustment and boron concentration regulation; the real-time feedback data of the actuator is input into the PID controller again to form a closed-loop control to ensure the continuous optimization and stability of the adjustment process.

7. A method for automatically suppressing xenon oscillations and providing auxiliary warning prompts based on the 3AO method according to claim 1, characterized in that, The specific process of step 4 is as follows: After the control adjustment is completed, the system will continue to collect the operating parameters of the reactor core, temperature, pressure, neutron flux, and coolant flow rate, as well as the current position of the control rod and the current value of the boron concentration; this data provides a basis for the system to judge the adjustment effect; The xenon oscillation monitoring interface provides clear xenon oscillation information feedback to the operator by displaying the operating parameters and control status of the core in real time; the operator monitors the real-time status of the core through this interface, promptly identifies the occurrence of xenon oscillation, and manually intervenes or adjusts the control strategy according to the feedback information to ensure the stable and safe operation of the reactor; Dynamic optimization adjustment: Data acquisition: After the control rod position and boron concentration are adjusted, the system continuously monitors and records the operating parameters of the core, namely temperature, pressure, neutron flux, coolant flow rate, and the current position of the control rod and the current value of the boron concentration; Data comparison: Compare the real-time collected feedback data with the preset target state and calculate the deviation of each parameter; Error identification: Identify the source of the deviation through comparative analysis and determine whether the xenon oscillation suppression is sufficient; Adjustment suggestion: If it is found that the xenon oscillation suppression is insufficient, the system will suggest further lowering the control rod or increasing the boron injection amount to achieve more stable control of the nuclear chain reaction; System performance evaluation and optimization: The system calculates the overall performance index of the system by evaluating the multiple measurement errors within the evaluation period to measure the quality of the adjustment effect; Optimization process: According to the latest operation data and feedback results, the system regularly performs the following steps: Data update: Input the latest collected operation parameters and control status data into the system; Process adjustment: According to the performance evaluation results, adjust the control strategy and operation steps, including the convolutional neural network model and the DQN model, to improve the adaptability and stability of the system; Continuous improvement: Through multiple rounds of cycles, continuously optimize the control process to ensure the continuous improvement of the xenon oscillation suppression effect.

Citation Information

Patent Citations

  • Method For Estimating A Future Value Of The Axial Power Imbalance In A Nuclear Reactor

    US20240145106A1