Water injection decision method for primary circuit system of nuclear power plant

CN122552207APending Publication Date: 2026-08-11CHINA NUCLEAR POWER TECH RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,在紧张的事故处理过程中,操作员需在短时间内综合海量信息做出决策,负担重、压力大

Benefits of technology

[0031] The aforementioned water injection decision-making method for the primary loop system of a nuclear power plant achieves real-time perception of the risk status of the primary loop system by acquiring operational status parameters and calculating safety risk coefficients. Secondly, when the risk is controllable, a predictive model is used to analyze historical state sequences, obtaining predictive results including pressure change rate, water level change rate, and water injection efficiency score. This provides a basis for decision-making based on historical evolution, overcoming the limitations of relying solely on the current instantaneous state. Furthermore, by integrating the current operating status, remaining available makeup water resources, and the aforementioned predictive results, a target water injection adjustment amount is generated through data processing using an adjustment model. Finally, water injection operations are executed based on this optimized adjustment amount, ensuring that each water injection is based on multi-dimensional information integration and forward-looking analysis. Through this method, the approach achieves a shift from static threshold response to dynamic multi-factor optimization decision-making, thereby improving the accuracy of primary loop water injection decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122552207A_ABST
    Figure CN122552207A_ABST
Patent Text Reader

Abstract

This application relates to a water injection decision-making method for the primary loop system of a nuclear power plant. The method includes: acquiring operating status parameters of the primary loop system and determining a safety risk coefficient for the primary loop system based on the operating status parameters; if the safety risk coefficient does not reach a preset risk threshold, inputting a sequence of status parameters into a pre-trained prediction model for prediction processing to obtain a prediction result; wherein the sequence of status parameters includes multiple operating status parameters arranged in chronological order within a preset historical period; inputting pre-acquired remaining available makeup water resources, the operating status parameters, and the prediction result into a pre-trained adjustment model for data processing to obtain a target water injection adjustment amount; and performing water injection operations on the primary loop system based on the target water injection adjustment amount. This method can improve the accuracy of water injection decisions for the primary loop system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of nuclear power plant technology, and in particular to a water injection decision method for the primary loop system of a nuclear power plant. Background Technology

[0002] The primary coolant loop of a nuclear power plant is the core of the nuclear reactor coolant system. Under accident conditions (such as coolant loss accidents or main steam pipe ruptures), the primary coolant loop may face risks such as a sudden drop in pressure and a rapid drop in water level. Timely and appropriate water injection is a key measure to prevent core overheating and ensure reactor safety.

[0003] Currently, the decision-making process for accidental water injection in the primary circuit of nuclear power plants mainly relies on comprehensive judgment and manual intervention based on operator experience.

[0004] However, during the tense process of handling an accident, operators need to make decisions based on massive amounts of information in a short period of time, which is a heavy burden and a lot of pressure. Moreover, manual decision-making makes it difficult to accurately predict the long-term and cascading effects of water injection (such as the subsequent effects on system temperature distribution and pressure recovery rate), and the accuracy of the decision cannot be guaranteed. Summary of the Invention

[0005] Therefore, it is necessary to provide a water injection decision method for the primary loop system of a nuclear power plant that can improve the accuracy of water injection decision-making in the primary loop system, in order to address the above-mentioned technical problems.

[0006] In a first aspect, this application provides a water injection decision-making method for the primary loop system of a nuclear power plant, including:

[0007] Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters;

[0008] If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into the pre-trained prediction model for prediction processing to obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0009] The remaining available water replenishment resources, operating status parameters, and prediction results obtained in advance are input into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0010] The primary loop system is injected with water based on the target water injection adjustment amount.

[0011] Secondly, this application also provides a water injection decision device for the primary loop system of a nuclear power plant, comprising:

[0012] The acquisition module is used to acquire the operating status parameters of the primary loop system and determine the safety risk coefficient of the primary loop system based on the operating status parameters.

[0013] The prediction processing module is used to input the state parameter sequence into a pre-trained prediction model for prediction processing when the safety risk coefficient does not reach the preset risk threshold, and obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0014] The data processing module is used to input the pre-acquired remaining available water replenishment resources, operating status parameters and prediction results into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0015] The execution module is used to perform water injection operations on the primary loop system based on the target water injection adjustment amount.

[0016] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0017] Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters;

[0018] If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into the pre-trained prediction model for prediction processing to obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0019] The remaining available water replenishment resources, operating status parameters, and prediction results obtained in advance are input into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0020] The primary loop system is injected with water based on the target water injection adjustment amount.

[0021] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0022] Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters;

[0023] If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into the pre-trained prediction model for prediction processing to obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0024] The remaining available water replenishment resources, operating status parameters, and prediction results obtained in advance are input into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0025] The primary loop system is injected with water based on the target water injection adjustment amount.

[0026] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0027] Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters;

[0028] If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into the pre-trained prediction model for prediction processing to obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0029] The remaining available water replenishment resources, operating status parameters, and prediction results obtained in advance are input into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0030] The primary loop system is injected with water based on the target water injection adjustment amount.

[0031] The aforementioned water injection decision-making method for the primary loop system of a nuclear power plant achieves real-time perception of the risk status of the primary loop system by acquiring operational status parameters and calculating safety risk coefficients. Secondly, when the risk is controllable, a predictive model is used to analyze historical state sequences, obtaining predictive results including pressure change rate, water level change rate, and water injection efficiency score. This provides a basis for decision-making based on historical evolution, overcoming the limitations of relying solely on the current instantaneous state. Furthermore, by integrating the current operating status, remaining available makeup water resources, and the aforementioned predictive results, a target water injection adjustment amount is generated through data processing using an adjustment model. Finally, water injection operations are executed based on this optimized adjustment amount, ensuring that each water injection is based on multi-dimensional information integration and forward-looking analysis. Through this method, the approach achieves a shift from static threshold response to dynamic multi-factor optimization decision-making, thereby improving the accuracy of primary loop water injection decisions. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is an internal structural diagram of a computer device in one embodiment;

[0034] Figure 2 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in one embodiment;

[0035] Figure 3 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0036] Figure 4 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0037] Figure 5 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0038] Figure 6 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0039] Figure 7 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0040] Figure 8 This is a flowchart illustrating the water injection decision-making method for the primary loop system of a nuclear power plant in another embodiment;

[0041] Figure 9 This is a structural block diagram of the water injection decision device for the primary loop system of a nuclear power plant in one embodiment. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0043] The primary coolant circuit (i.e., the reactor coolant pressure boundary) of a nuclear power plant is the first line of defense for nuclear safety, bearing the core safety functions of maintaining core cooling, dissipating decay heat, and containing radioactive materials. Its system integrity and operational stability are fundamental prerequisites for ensuring safe reactor operation. Under design-basis accidents or beyond-design-basis accidents, such as loss-of-coolant incidents or steam generator heat transfer tube ruptures, the primary coolant circuit may face extreme risks including rapid coolant loss, sudden pressure drops, core exposure, and temperature increases. In such cases, timely and precise emergency water injection is one of the most critical proactive safety measures to prevent core meltdown, mitigate accident consequences, and ensure the integrity of the containment structure.

[0044] Currently, the mainstream water injection decision-making and control methods in the industry mainly follow a deterministic safety analysis and procedure-driven model. However, their core logic has the following limitations: Existing technologies typically rely on pre-set, single or limited threshold parameters (such as low pressure thresholds or low water level thresholds) as the criteria for triggering water injection. This threshold triggering mechanism is essentially reactive, meaning that a response is only initiated after the parameter actually exceeds its limit. However, severe accidents are complex and variable, with strong dynamic coupling of parameters. By the time the threshold is triggered, the system state may have already deteriorated rapidly, leading to a compressed critical intervention window. Water injection decisions inherently lag and cannot adapt to the rapid dynamic evolution of accidents.

[0045] It is evident that existing water injection decision-making technologies for the primary loop of nuclear power plants have significant shortcomings in accuracy when facing complex and dynamic accident scenarios. Therefore, this application proposes a water injection decision-making method for the primary loop system of nuclear power plants to address the aforementioned problems.

[0046] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 1 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores relevant data in the water injection decision-making process of the nuclear power plant's primary loop system. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a water injection decision-making method for the nuclear power plant's primary loop system.

[0047] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0048] In one exemplary embodiment, such as Figure 2 As shown, a water injection decision method for the primary loop system of a nuclear power plant is provided, which is then applied to... Figure 1 Taking the server in the example, the explanation includes the following steps 201 to 204. Wherein:

[0049] Step 201: Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters.

[0050] Among them, the operating status parameters refer to the physical quantities used to characterize the real-time operating conditions of the primary loop system, which may include pressure, water level, temperature and current water injection flow rate.

[0051] A safety risk coefficient refers to one or more indicators used to quantify the degree to which a current system deviates from a safe state. Its purpose is to synthesize continuous, multi-dimensional operational parameters into a scalar or vector that can be used for rapid risk assessment, providing a basis for subsequent decision-making regarding whether to initiate advanced intelligent systems. Specifically, one implementation of the safety risk coefficient may include a pressure risk coefficient calculated based on the current pressure and a preset pressure safety threshold, and a water level risk coefficient calculated based on the current water level and a preset water level safety threshold.

[0052] In this embodiment, the server receives a real-time data stream from a data acquisition system. This data stream includes pressure, water level, temperature, and water injection flow rate data of key measuring points in the first loop, collected and preprocessed at a fixed frequency. Subsequently, the server calculates the current pressure risk coefficient and water level risk coefficient based on preset pressure risk thresholds and water level risk thresholds, which together constitute the safety risk coefficient at the current moment.

[0053] In another embodiment, the server receives not only basic operating status parameters but also derived parameters such as the subcooling degree of the primary loop and the secondary water level of the steam generator, to more comprehensively construct a state profile of the primary loop system. When determining the safety risk coefficient, the server can use a weighted fusion algorithm to synthesize the pressure risk coefficient and the water level risk coefficient according to preset weights (e.g., the water level risk has a higher weight) to obtain a comprehensive risk score, which serves as a single safety risk coefficient for subsequent judgment.

[0054] Step 202: If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into the pre-trained prediction model for prediction processing to obtain the prediction result.

[0055] The status parameter sequence includes multiple operating status parameters arranged in chronological order within a preset historical time period.

[0056] The preset risk threshold refers to one or more threshold values ​​used to distinguish between routine decision-making modes and emergency intervention modes. When the safety risk coefficient is lower than this threshold, it indicates that the state of the primary loop system is still within the range that can be optimized through refined and forward-looking strategies, thus initiating the subsequent intelligent prediction and decision-making process.

[0057] The state parameter sequence refers to the input data format constructed for the predictive model, which includes the operating state parameters at the current moment and over a previous period (e.g., one data point per minute over the past 10 minutes) to capture the dynamic evolution trend of the system state.

[0058] A predictive model is a machine learning model that has been trained to learn the dynamic characteristics of a one-loop system and predict its short-term future behavior. Its input is a sequence of state parameters, and its output is a prediction of the changing trends of key system parameters over a future period.

[0059] In this embodiment, the server compares the calculated safety risk coefficient (e.g., a comprehensive risk score of 0.25) with a preset conventional risk threshold (e.g., 0.3). Since 0.25 < 0.3, the server determines to enter the conventional decision-making mode. Subsequently, the server extracts pressure, water level, temperature, and water injection flow rate data for the most recent 10 minutes from the historical database, one point per minute, forming a 10×4 state parameter sequence matrix. This matrix is ​​input into a pre-loaded, trained LSTM prediction model in memory. After the model runs, it outputs the pressure change rate (ΔP), water level change rate (ΔL), and water injection efficiency score (E_inj) for the next 5 minutes as prediction results.

[0060] In another embodiment, the preset risk threshold can be set as a two-dimensional vector, corresponding to the pressure risk coefficient threshold and the water level risk coefficient threshold, respectively. The server needs to determine that neither of the two risk coefficients has reached its respective threshold before initiating the prediction process. The construction of the state parameter sequence can be more targeted. For example, the server can identify the current accident type as a small breach water loss accident, and then extract historical time period data that is more relevant to the characteristics of similar accidents from the historical database to construct the sequence. In addition to outputting the rate of change, the prediction model can also output the predicted value curves of future key parameters, providing a more intuitive reference for decision-making.

[0061] Step 203: Input the pre-acquired remaining available water replenishment resources, operating status parameters and prediction results into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount.

[0062] The remaining available water resources refer to the total amount of cooling capacity that can be injected into the primary circuit from the emergency water supply tank or the water tank of the safety injection system at the current moment.

[0063] An adjustment model refers to an intelligent decision-making model that has been specially trained and can output the optimal operation instructions based on the current system state, resource constraints, and future predictions.

[0064] In this embodiment, the server reads the remaining water tank capacity of the current high-pressure injection system from the power plant database as the remaining available water replenishment resource. Simultaneously, the server packages the current operating status parameters (pressure, water level, temperature, current injection flow rate), the aforementioned remaining water resource data, and the prediction results (ΔP, ΔL, E_inj) obtained from the prediction model into a single state vector. This state vector is input into another pre-loaded reinforcement learning adjustment model trained using the Proximal Policy Optimization (PPO) algorithm. The model internally evaluates and calculates based on the learned policy, ultimately outputting a specific flow adjustment amount, such as increasing the injection flow rate by 50 kg / s; this value is then determined as the target injection adjustment amount.

[0065] Step 204: Perform water injection operation on the primary loop system based on the target water injection adjustment amount.

[0066] Among them, the water injection operation refers to converting the instructions generated by intelligent decision-making into actual control signals, driving the valves, pumps and other actuators of the water injection system to change the flow rate of coolant injected into the primary circuit.

[0067] In this embodiment, the server sends the calculated target water injection adjustment (e.g., "+50 kg / s") to the flow controller of the high-pressure injection system via a control network. After receiving the instruction, the flow controller compares it with the current actual flow rate, calculates the required valve opening change or pump speed adjustment, and drives the corresponding actuator to gradually adjust the actual water injection flow rate to the target value, thereby completing a water injection operation based on intelligent decision-making.

[0068] In another embodiment, the execution process may include safety checks and a smooth transition. For example, before sending control commands, the server compares the target adjustment amount with the system's physical limits (such as the maximum allowable flow rate) and performs limiting processing. After the command is issued, the server does not adjust the flow rate all at once, but may break it down into several smaller, gradual adjustment command sequences and send them to the actuators in a time-sharing manner to achieve a smooth change in flow rate and avoid excessive thermal-hydraulic shock to the primary loop system. Simultaneously, the server continuously monitors the actual response of the system after execution and compares it with the expected response, providing data for subsequent online model learning.

[0069] In the aforementioned water injection decision-making method for the primary loop system of a nuclear power plant, real-time perception of the risk status of the primary loop system is achieved by acquiring operational status parameters in real time and calculating safety risk coefficients. Secondly, when the risk is controllable, a predictive model is used to analyze historical state sequences, obtaining predictive results including pressure change rate, water level change rate, and water injection efficiency score. This provides a basis for decision-making based on historical evolution, overcoming the limitations of relying solely on the current instantaneous state. Furthermore, by integrating the current operating status, remaining available makeup water resources, and the aforementioned predictive results, a target water injection adjustment amount is generated through data processing using an adjustment model. Finally, water injection operations are performed based on this optimized adjustment amount, ensuring that each water injection is based on multi-dimensional information integration and forward-looking analysis. Through this method, the approach achieves a shift from static threshold response to dynamic multi-factor optimization decision-making, thereby improving the accuracy of primary loop water injection decisions.

[0070] In one exemplary embodiment, such as Figure 3 As shown, the above-mentioned "inputting the state parameter sequence into a pre-trained prediction model for prediction processing to obtain the prediction result" includes steps 301 to 303. Wherein:

[0071] Step 301: Determine the weight values ​​of the operating state parameters at different historical moments in the state parameter sequence based on the water level values ​​in the state parameter sequence.

[0072] The weight values ​​refer to the coefficients assigned to the data at each historical moment in the state parameter sequence. The determination of these coefficients aims to enable the predictive model to dynamically focus on historical information most relevant to the current security state, thereby improving the accuracy of predictions and its sensitivity to risks.

[0073] In this embodiment, the server extracts a sequence of state parameters of length 10 (i.e., including the current time t and the previous 9 historical times). This sequence includes the pressure P(t), water level L(t), temperature T(t), and injection flow rate Q_inj(t) for each time point. The server reads the water level value L(t) at the current time t and, based on a pre-defined attention calculation submodule, calculates a weight value α_i between 0 and 1 for each historical time point i (i=t-9,...,t) in the sequence. The operation logic of this submodule is as follows: it fuses the current water level L(t) with the hidden state that encodes the sequence information up to time i, and outputs the weight through an activation function, so that when the water level L(t) is low, those times in history where the water level was also low or rapidly decreasing will receive higher weights.

[0074] In another embodiment, the determination of weight values ​​can be more directly correlated with the pattern of water level changes. For example, the server first analyzes the water level change trend in the state parameter sequence and calculates the rate of change of water level at each historical moment. When the current water level value L(t) is below a certain attention threshold, the server assigns higher weights to historical moments with negative rates of change (i.e., water level decline) and large absolute values, because the dynamics of these moments may indicate the cause or evolution of the current risk. The specific values ​​of the weights can be generated using a monotonically increasing function related to the absolute value of the rate of change of water level.

[0075] Step 302: Weight the weight values ​​and the corresponding running state parameters to obtain weighted time series features.

[0076] Weighted time-series features refer to the new time-series data representation obtained after weighted processing. It emphasizes historical information that is more relevant to the current focus of attention (the risk status indicated by the current water level value) while weakening information with lower relevance, thus forming an enhanced time-series context representation that focuses on specific risk patterns.

[0077] In this embodiment, the server performs scalar multiplication on the calculated 10 weight values ​​α_i (i=t-9,...,t) with the corresponding 10 state parameter vectors x_i=[P(i),L(i),T(i),Q_inj(i)] to obtain 10 weighted feature vectors α_i*x_i. These 10 weighted vectors are arranged in chronological order to form a new matrix with the same dimension as the original sequence. This matrix is ​​the weighted temporal feature output in this step. In this feature matrix, the parameter values ​​of historical moments with high weights are relatively amplified, while the parameter values ​​of moments with low weights are relatively reduced.

[0078] Step 303: Determine the prediction result based on the weighted time series features.

[0079] In this embodiment, the server can input the weighted temporal feature matrix (10×4-dimensional) obtained in step 302 into the subsequent fully connected neural network layer of the prediction model. This network layer performs nonlinear transformation and compression on the weighted temporal information, ultimately outputting a prediction result vector containing three scalar values: [ΔP, ΔL, E_inj]. Here, ΔP represents the predicted average pressure change rate (Pa / min) over the next 5 minutes, ΔL represents the predicted average water level change rate (m / min) over the next 5 minutes, and E_inj represents the predicted water injection efficiency score, calculated based on the ratio of the predicted ΔL to the average water injection flow rate over the next 5 minutes.

[0080] In an exemplary embodiment, the above prediction results include at least the pressure change rate, water level change rate, and water injection efficiency score, based on which, such as Figure 4 As shown, the above "determining the prediction result based on weighted time series features" includes steps 401 to 402. Wherein:

[0081] Step 401: Based on the weighted time series characteristics, predict the pressure change rate and water level change rate of the primary loop system within a future set time period, and determine the water level change amount based on the water level change rate.

[0082] Among them, the weighted time series feature refers to the historical state parameter sequence after dynamic weight adjustment, which highlights historical information that is highly correlated with the current risk state.

[0083] The future timeframe refers to the length of time covered by the forecast, such as the next 5 minutes. The selection of this timeframe requires a trade-off between the timeliness and accuracy of the forecast; it should be long enough to assess the short- to medium-term effects of the water injection measures, but short enough to ensure the reliability of the forecast under rapidly evolving accident conditions.

[0084] The pressure change rate refers to the average rate of change of the primary loop system pressure (such as pressurizer pressure) over a predicted future set period, measured in Pa / min. A positive value indicates a pressure increase, and a negative value indicates a pressure decrease. The water level change rate refers to the average rate of change of the primary loop system water level (such as pressurizer water level or core water level) over a predicted future set period, measured in m / min. A positive value indicates a water level increase, and a negative value indicates a water level decrease. The water level change amount refers to the absolute value of the cumulative water level change over a predicted future set period, calculated based on the predicted water level change rate. It is typically calculated as the product of the water level change rate and the period length, measured in meters.

[0085] In this embodiment, the server inputs the weighted temporal feature matrix into the final fully connected output layer of the prediction model. This output layer is configured to directly output two scalar values: ΔP (pressure change rate) and ΔL (water level change rate). Based on the set prediction period length T (e.g., 5 minutes), the server calculates the predicted total water level change using the formula ΔL_total = ΔL * T.

[0086] Step 402: Determine the water injection efficiency score based on the proportional relationship between the water level change and the corresponding average water injection flow rate within a future set time period.

[0087] The average injection flow rate refers to the predicted or planned average mass of coolant injected into one circuit per unit time during a pre-defined future injection period, assuming the implementation of the injection strategy. This value may differ from the current injection flow rate and reflects the anticipated injection operation.

[0088] Water injection efficiency score is a performance indicator used to quantify the contribution of a unit injection flow rate to water level recovery. It is the ratio of water level rise to injection flow rate. The higher the score, the more efficient the water injection behavior is in raising the water level under given operating conditions and injection strategies. It provides an important economic consideration for subsequent decision optimization and helps to achieve efficient resource utilization while ensuring safety.

[0089] In this embodiment, the server obtains the planned water injection flow rate Q_plan for a future set time period (e.g., maintaining the current flow rate, or a flow rate based on a preliminary decision). The server calculates the average water injection flow rate Q_avg during this time period (if Q_plan is constant, then Q_avg = Q_plan). Subsequently, the server uses the water level change ΔL_total calculated in step 401 to calculate the water injection efficiency score E_inj using the formula E_inj = ΔL_total / Q_avg, with units of m / (kg / s) or equivalent units.

[0090] In an exemplary embodiment, the training process of the above prediction model includes:

[0091] Multiple sets of training samples are obtained; training samples with water levels below the risk threshold are oversampled, and a hybrid loss function is used to backpropagate and optimize the model parameters to update the model parameters until the iterative calculation is completed and the prediction model is obtained.

[0092] Among them, the hybrid loss function is used to constrain the prediction error of pressure change rate, water level change rate, and water injection efficiency score.

[0093] Training samples refer to the input-output data pairs used to train the prediction model. The input to each training sample is typically a combination of system parameters (including primary loop geometric parameters, accident release source terms, and water injection operation parameters) generated by Latin hypercube sampling under a specific accident condition. This is followed by transient simulation using a high-fidelity thermal-hydraulic program (such as RELAP5), resulting in a time-varying sequence of primary loop key state parameters (pressure P, water level L, temperature T, and water injection flow rate Q_inj). The output of the sample is the corresponding prediction target value, such as the true values ​​of pressure change rate, water level change rate, and water injection efficiency score over future periods.

[0094] A hybrid loss function is a composite loss function composed of multiple individual loss functions combined with specific weights. Its purpose is to simultaneously optimize multiple prediction objectives of the model, using weight allocation to reflect the differences in importance between different objectives. In this method, the hybrid loss function contains at least three components, used to measure and constrain the deviations between the model's predicted values ​​and actual values ​​for pressure change rate, water level change rate, and water injection efficiency score. The core of designing this loss function is to ensure highly accurate predictions of core safety parameters (pressure and water level change rate) while effectively controlling the prediction deviations of the water injection efficiency score, which reflects economic efficiency and strategy effectiveness, thereby achieving multi-objective synergistic optimization.

[0095] Backpropagation optimization refers to the process in neural network training where the error signal, calculated using the loss function, is propagated backward along the network structure from the output layer to the input layer, and the gradient of the loss function with respect to each model parameter is calculated according to the chain rule. Subsequently, optimization algorithms are used to update the model parameters based on these gradients, gradually reducing the value of the loss function so that the model's predictive ability continuously approximates the patterns contained in the training data.

[0096] In this embodiment, the server loads N pre-prepared training samples from the training database. Before each training iteration, the server scans the water level label values ​​of all training samples and identifies all samples with water levels below a set risk threshold (e.g., L_critical = 1.5 meters). The server then copies these low-water-level samples K times (e.g., K = 2), increasing their frequency in the current training batch to (K+1) times the original frequency, while keeping the frequency of normal-water-level samples unchanged. Next, the server constructs a mixed loss function L_total = α*MSE(ΔP) + β*MSE(ΔL) + γ*MAE(E_inj), where MSE is the mean squared error, MAE is the mean absolute error, and α, β, and γ are preset weight coefficients (e.g., α = 1.0, β = 1.0, γ = 0.5). After each forward propagation to obtain a predicted value, the server calculates L_total and uses the backpropagation algorithm to calculate the gradient, updating the weights and bias parameters of the LSTM prediction model using the Adam optimizer. This process is repeated over multiple epochs until the model's performance on the validation set stabilizes. The final model saved is the trained prediction model.

[0097] In one exemplary embodiment, such as Figure 5 As shown, the above-mentioned "inputting the pre-acquired remaining available water replenishment resources, operating status parameters, and prediction results into a pre-trained adjustment model for data processing to obtain the target water injection adjustment amount" includes steps 501 to 504. Wherein:

[0098] Step 501: Determine the pressure risk coefficient based on the current pressure value and candidate water injection adjustment amount in the operating status parameters; and determine the water level risk coefficient based on the current water level value and candidate water injection adjustment amount in the operating status parameters.

[0099] Specifically, when determining the pressure risk coefficient and the water level risk coefficient, the pressure risk coefficient is verified using the pressure change rate in the prediction results, and the water level risk coefficient is corrected using the water level change rate in the prediction results.

[0100] The candidate water injection adjustment amount refers to the specific change in water injection flow rate that the reinforcement learning agent considers at the current decision moment, such as increasing by 50 kg / s, decreasing by 20 kg / s, or keeping it unchanged. This adjustment amount is the action option that the agent internally evaluates and selects.

[0101] Pressure risk coefficient and water level risk coefficient are indicators that quantify the degree to which the current pressure state and water level state deviate from the safe range, respectively.

[0102] In this embodiment, the server first sets a candidate water injection adjustment amount ΔQ_candidate. To evaluate this action, the server assumes that the adjustment is performed immediately and calculates the adjusted instantaneous pressure P_instant = P_current + f_P(ΔQ_candidate).

[0103] Next, the server calculates the instantaneous water level L_instant = L_current + f_L(ΔQ_candidate), where f_P and f_L are simplified instantaneous impact estimation functions. Based on P_instant, the design pressure P_design, and the pressure critical value P_critical, the base pressure risk coefficient R_p_base is calculated. Simultaneously, based on L_instant and the water level critical value L_critical, the base water level risk coefficient R_L_base is calculated. Subsequently, the server reads the future pressure change rate ΔP_pred and water level change rate ΔL_pred given by the prediction model under the current state. ΔP_pred is used to correct R_p_base, for example, R_p = R_p_base * (1 + k_p * |ΔP_pred|), where k_p is the correction coefficient. If ΔP_pred is negative (pressure decrease), it may further increase the perceived risk. Similarly, ΔL_pred is used to correct R_L_base to obtain R_L.

[0104] Step 502: Determine the safety assessment value based on the pressure risk coefficient and the water level risk coefficient.

[0105] The safety assessment value is a comprehensive scalar used to quantify the expected level of overall system safety after implementing candidate water injection adjustments. It is the result of combining the risk coefficients of both pressure and water level dimensions according to their importance to safety. Typically, water level risk is given a higher weight because it is directly related to core cooling.

[0106] In this embodiment, the server calculates the security assessment value S_safety using a weighted summation method: S_safety = -(w_p*R_p + w_L*R_L), where w_p and w_L are preset weight coefficients, and w_L is usually greater than w_p. The negative sign indicates that the higher the risk coefficient (the less secure), the lower the security assessment value (the greater the penalty). This assessment value will be used for subsequent comprehensive decision-making.

[0107] Step 503: Determine the economic evaluation value based on the remaining available water replenishment resources and the current water injection flow rate in the operating status parameters.

[0108] The economic assessment value is a scalar used to quantify resource utilization efficiency. It encourages the economical use of limited emergency cooling water resources (remaining available makeup water resources) while discouraging the use of excessive and potentially unnecessary injection flows (current injection flow). Its goal is to achieve long-term, efficient utilization of makeup water resources while meeting safety requirements.

[0109] In this embodiment, the server calculates the economic evaluation value S_econ as: S_econ = c * W_remaining - d * (Q_current)^2. Where W_remaining is the remaining available replenishment water resources (unit: kg), Q_current is the current injection flow rate (unit: kg / s), and c and d are positive coefficients. The first term encourages retaining more water resources, while the second term (negative term) penalizes high injection flow rates, and the use of a square term makes the penalty increase sharply with the flow rate, thereby suppressing unnecessary excessive flow.

[0110] Step 504: The candidate water injection adjustment amount that maximizes the comprehensive evaluation value composed of the safety evaluation value and the economic evaluation value is determined as the target water injection adjustment amount.

[0111] The comprehensive evaluation value is the final evaluation index that combines the safety evaluation value and the economic evaluation value according to the overall system optimization objective. Finding the candidate water injection adjustment amount that maximizes this comprehensive evaluation value is essentially finding the optimal balance between the potentially conflicting objectives of safety and economy. This is fundamentally a multi-objective optimization problem, which is solved in this method by constructing a unified evaluation function and optimizing it.

[0112] In this embodiment, the server defines a comprehensive evaluation value: Total = Ssafety + λ * Second, where λ is an adjustment coefficient (0 < λ < 1, typically small to ensure safety priority) used to balance the importance of safety and economy. For the current decision moment, the server enumerates or samples multiple candidate water injection adjustment amounts ΔQcandidate within its action space (e.g., from -100 kg / s to +100 kg / s, with a step size of 10 kg / s). For each candidate amount, the server sequentially executes steps 501 to 503 to calculate the corresponding Ssafety and Second, thereby obtaining Total. Finally, the server selects the candidate water injection adjustment amount that maximizes the Total value as the final determined target water injection adjustment amount ΔQtarget and outputs it for execution.

[0113] In an exemplary embodiment, the training process of the above-described adjusted model includes:

[0114] A primary-loop water injection simulation environment based on a prediction model and a primary-loop parameterized model is established. The state space, action space, and reward function of the primary-loop water injection simulation environment are obtained, and an untrained initial model is constructed. The state space is used to characterize the state of the simulation environment at any given time, and includes at least the pressure change rate, water level change rate, and water injection efficiency score output by the prediction model, as well as the remaining available water replenishment resources. The reward function is used to determine the instantaneous reward values ​​for safety and economic rewards based on the changes in state values ​​after the execution of actions. The model is trained in the primary-loop water injection simulation environment, and an adjusted model is obtained after training.

[0115] In this embodiment, the server integrates a pre-trained LSTM prediction model and a primary loop physical characteristic model to construct a simulation platform for reinforcement learning training. This environment can simulate the dynamic response of the primary loop under accident conditions. The server defines a state space, action space, and reward function for the simulation environment. The state space includes the pressure change rate, water level change rate, water injection efficiency score, and remaining available water replenishment resources output by the prediction model; the action space is the continuous adjustment of the water injection flow rate; the reward function is designed as a weighted sum of safety rewards and economic rewards. The server uses a proximal policy optimization algorithm to allow the initialized agent model to interact and test with the simulation environment in multiple rounds, iteratively updating the model parameters by collecting state, action, and reward data until the policy performance is stable. When the model's performance on the validation set reaches a preset standard (e.g., the average reward value stabilizes above a threshold), the server stops training and saves the final network parameters, obtaining an adjusted model that can be used for real-time decision-making. In an exemplary embodiment, such as... Figure 6 As shown, the above-mentioned "training the model in a primary loop water injection simulation environment and obtaining an adjusted model after training" includes steps 601 to 603. Wherein:

[0116] Step 601: Obtain the current state observation value output by the primary loop water injection simulation environment, input the current state observation value into the initial model, and use the initial model to output the candidate water injection adjustment amount based on the current state observation value.

[0117] Here, the current state observation refers to all the information vectors output by the simulation environment at a certain moment that are available for the model to perceive. The initial model refers to the neural network policy model that has not yet completed training.

[0118] In this embodiment, the server reads the current state vector s_t from the primary loop water injection simulation environment. This vector contains the predicted pressure change rate, water level change rate, water injection efficiency score, and remaining replenishment volume. The server inputs this state vector into the PPO policy network, which outputs an action probability distribution through forward propagation. After sampling, the candidate water injection adjustment amount a_t is obtained.

[0119] Step 602: Input the candidate water injection adjustment amount into the primary loop water injection simulation environment for execution, and obtain the next state observation value fed back by the simulation environment and the instantaneous reward calculated by the reward function.

[0120] The immediate reward refers to the scalar evaluation value calculated by a preset reward function based on the change in the environmental state after the action is performed.

[0121] In this embodiment, the server sends the candidate water injection adjustment amount a_t to the simulation environment. The environment updates its primary loop state based on this action and calls the reward function to calculate the immediate reward r_t. Simultaneously, the environment generates the state observation value s_{t+1} for the next time step and feeds it back to the server.

[0122] Step 603: Based on the current state observation, candidate water injection adjustment amount, next state observation and immediate reward, iteratively update the model parameters to learn the target water injection decision strategy that maximizes the cumulative reward.

[0123] Here, cumulative reward refers to the sum of all immediate rewards obtained by the agent in a complete training round. Model parameter updates aim to enable the agent to learn to select action sequences that yield higher cumulative rewards.

[0124] In this embodiment, the server collects multiple sets of data (s_t, a_t, r_t, s_{t+1}) to form an experience replay buffer. Based on these data, the advantage function is calculated, and the parameters of the policy network and value network are updated through gradient backpropagation, enabling the model to gradually improve its decision-making strategy.

[0125] In one exemplary embodiment, the method further includes:

[0126] If neither the pressure risk coefficient nor the water level risk coefficient reaches the risk threshold, water injection operation is performed on the primary loop system based on the target water injection adjustment amount.

[0127] In this embodiment, the server monitors the calculated pressure risk coefficient R_p and water level risk coefficient R_L in real time. When R_p < 0.3 and R_L < 0.3 (where 0.3 is an example of a preset risk threshold), the server determines that the current situation is low-risk and can adopt refined intelligent decision-making. Subsequently, the server sends the target water injection adjustment amount (e.g., +25 kg / s) previously calculated through the prediction model and adjustment model to the flow regulation unit of the high-pressure safety injection system through the control interface to perform precise flow adjustment.

[0128] In another embodiment, a more conservative strategy can be adopted for determining the risk threshold. For example, a risk buffer can be set up so that even if both R_p and R_L are below the threshold, as long as either one shows a rapid upward trend in the recent period (e.g., in the past 30 seconds) (with a derivative greater than zero), the server may moderately increase the execution priority or slightly amplify the adjustment amount to respond to potential risk accumulation in advance.

[0129] In one exemplary embodiment, the method further includes:

[0130] When the pressure risk coefficient and water level risk coefficient reach or exceed the risk threshold, switch to emergency intervention mode; emergency intervention mode is used to execute the preset maximum flow water injection command.

[0131] Among them, the emergency intervention mode refers to the special control logic activated when the system risk reaches a critical level. This mode will bypass the conventional intelligent decision-making process and directly execute the pre-set control instructions with the primary goal of ensuring safety to the greatest extent.

[0132] The maximum flow rate water injection command refers to the control command that controls the water injection system to inject water at its maximum design capacity or the safe limit flow rate under current operating conditions.

[0133] In this embodiment, the server monitors the pressure risk coefficient R_p and the water level risk coefficient R_L in real time. Once R_p ≥ 0.8 or R_L ≥ 0.8 (where 0.8 is a preset emergency risk threshold example), the server immediately triggers an interruption mechanism, stops the ongoing routine intelligent decision-making process, and switches system control to emergency intervention mode.

[0134] Upon entering emergency intervention mode, the server ignores the target water injection adjustment from the adjustment model and instead sends a setpoint command to the high-pressure safety injection system. This command instructs the system to immediately increase the water injection flow rate and maintain it at the equipment's maximum permissible safe flow rate Q_max (e.g., 500 kg / s). This command is sent via a high-priority control channel to ensure immediate execution.

[0135] In one exemplary embodiment, such as Figure 7 As shown, the above method also includes:

[0136] Step 701: After performing the water injection operation, obtain the actual state response results of the primary loop system.

[0137] The actual state response result refers to the data obtained by sensors after the water injection operation, which reflects the state evolution of the primary loop system over a period of time (such as one or several decision cycles). It usually includes changes in key parameters such as pressure, water level, and temperature.

[0138] In this embodiment, after sending a water injection control command (whether it is the target adjustment amount in normal mode or the maximum flow command in emergency mode), the server starts a timer. Over the next 5 minutes, the server continuously receives time-series data of parameters such as actual pressure P_real(t), water level L_real(t), and temperature T_real(t) from the data acquisition system, and uses this as the actual status response result of this water injection operation.

[0139] In another embodiment, in addition to directly measuring the parameters, the server also calculates the actual pressure change rate ΔP_real and water level change rate ΔL_real as more intuitive response indicators. Simultaneously, if the injection flow rate fails to fully reach the target value during execution due to actuator limitations, the server also records the actual injection flow rate sequence Q_inj_real(t) for subsequent, more accurate performance evaluation.

[0140] Step 702: Determine the error between the actual state response result and the predicted result.

[0141] Among them, prediction error refers to the quantitative difference between the predictions made by the model (such as pressure change rate ΔP_pred, water level change rate ΔL_pred, water injection efficiency score E_inj_pred) and the corresponding results actually observed (ΔP_real, ΔL_real, or E_inj_real calculated based on actual data).

[0142] In this embodiment, the server compares the actual water level change rate ΔL_real obtained in step 701 with the predicted water level change rate ΔL_pred previously output by the prediction model, and calculates its absolute error: e_L = |ΔL_pred - ΔL_real|. Similarly, the absolute error e_P of the pressure change rate is calculated. Then, the server calculates a comprehensive error index, such as a weighted error sum: e_total = α * e_P + β * e_L, where α and β are weighting coefficients.

[0143] In another embodiment, the calculation of prediction error can focus more on "directional" errors. For example, if the model predicts that the water level will rise (ΔL_pred>0), but the actual water level falls (ΔL_real<0), this fundamental prediction error will be penalized with a high value, rather than just the numerical difference. Errors can also be dynamically weighted according to the current risk level, with prediction biases occurring during high-risk periods being considered more severe.

[0144] Step 703: If the error exceeds the preset error threshold, the sample containing the actual state response result and the corresponding water injection operation is determined as a new training sample; the new training sample is used to update the prediction model and / or adjust the model parameters in an incremental learning manner.

[0145] Newly added training samples refer to data tuples consisting of a complete interaction of state, action, and result, which can be used for further training of the model. Incremental learning refers to fine-tuning the parameters of an existing model using a small amount of newly generated data without retraining on a large scale, enabling it to adapt to slow changes in system characteristics or compensate for deficiencies in existing knowledge.

[0146] In this embodiment, the server compares the comprehensive error e_total calculated in step 702 with a preset error threshold e_threshold (e.g., 0.1). If e_total > e_threshold, the server determines that the current prediction deviates significantly from the actual situation and needs to update the model using new knowledge. The server packages the complete data of this decision cycle (including the state sequence before the decision, the target water injection adjustment amount, and the obtained actual state response results) into a new training sample and stores it in an online learning buffer. Subsequently, the server samples a small batch (containing new samples and other recent samples) from this buffer with a small learning rate and performs a parameter update iteration on the prediction model and the adjustment model respectively.

[0147] In another embodiment, the triggering of incremental learning can be more intelligent. The server can establish a sliding window statistics of error, and only trigger a model update when the error exceeds a threshold in multiple consecutive decisions, or when the error shows a systematic upward trend, in order to avoid model instability caused by a single random fluctuation. For updates to predictive models, the focus may be on correcting the accuracy of their state inferences; for updates to adjustment models, the focus may be on correcting their reward expectations or policy choices.

[0148] In one exemplary embodiment, such as Figure 8 As shown, the above method also includes:

[0149] Step 1: Parametric Modeling of Primary Loop Injection

[0150] The model is built upon the detailed geometry and thermo-hydraulic characteristics of the primary loop system. By precisely dividing the computational domain geometry, the model can adapt to complex piping layouts, ensuring precise capture of fluid flow and heat transfer paths. Specifically, geometric modeling requires adherence to the actual layout of the primary loop system, including the precise dimensions, elevations, and spatial relationships of pressure vessels, steam generators, main pumps, pressurizers, and connecting pipes. This information is derived from nuclear power plant design drawings, equipment specifications, and actual installation data, ensuring the accuracy of the geometric model.

[0151] In terms of physical modeling, the governing equations include the mass conservation equation, momentum conservation equation, and energy conservation equation. The mass conservation equation describes the relationship between the mass change of the coolant and the injection water flow rate in the primary loop, ensuring mass balance. The momentum conservation equation characterizes the flow characteristics of the coolant in the loop, considering the effects of pressure gradient, frictional resistance, and gravity. The energy conservation equation simulates the heat transfer process in the primary loop, including heat exchange between the coolant and the wall, power release from heating elements, and energy input from the injection water. These equations are solved discretizedly using the finite volume method or the finite element method to form a dynamic simulation model of the primary loop system.

[0152] During the modeling process, the impact mechanism of the water injection system on the primary loop state must be carefully considered. Water injection not only directly increases the coolant content in the loop, affecting water level and pressure, but also introduces fluids at different temperatures, altering the local temperature distribution and thus affecting the thermal-hydraulic response. Therefore, the model needs to couple water injection parameters (such as injection temperature, pressure, and flow rate) with primary loop state parameters (pressure, temperature, and water level) to establish a multiphysics coupling relationship, ensuring the ability to simulate the complex interactions during the water injection process.

[0153] Through the comprehensive modeling method described above, the model can reflect the dynamic changes of the primary loop under accident conditions in real time, providing high-fidelity data support for subsequent deep learning prediction and reinforcement learning optimization. This refined modeling method not only improves prediction accuracy but also lays a solid foundation for the intelligent and automated water injection decision-making.

[0154] Step 2: Multi-condition sampling and analysis of the primary loop water injection

[0155] In the decision-making process for primary loop water injection in nuclear power plants, multi-condition sampling calculation is a crucial step in achieving accurate accident management and optimizing water injection strategies. This process involves coupled calculations of multiple parameters and requires comprehensive consideration of the characteristics of the water injection system, the geometry of the primary loop, and the dynamic changes in accident scenarios.

[0156] The input parameter set X for multi-parameter coupled computation contains four types of key parameters, defined as follows:

[0157]

[0158] Here, Loop refers to parameters related to the characteristics of the primary loop system itself. V_loop represents the volume information (unit: m³) and elevation (unit: m) of each region or segment of the primary loop. J_loop represents the characteristics of the connection channels between compartments or segments, including flow area (unit: m²), upstream and downstream connection relationships, local drag coefficients, etc., used to describe the flow path and resistance characteristics of the coolant. S_loop represents the thermal component information of the primary loop system, such as core fuel assemblies, steam generator heat transfer tubes, pipe walls, etc., including their heat exchange area (unit: m²), wall thickness (unit: m), material thermal properties (such as thermal conductivity, specific heat capacity), spatial orientation (angle, °), and location, used to calculate the heat exchange between the system and the outside world. Release refers to the mass energy items released into the primary loop under accident conditions; different source items correspond to different accident types (such as small breach loss-of-coolant accidents, main steam pipe ruptures). Q_steam represents the steam flow rate (kg / s) released into the primary loop, indicating the amount of steam leaking from the secondary side to the primary side or for emergency depressurization. h_steam represents the specific enthalpy of the released steam (J / kg), determining the energy carried by the released steam. Q_water represents the liquid flow rate (kg / s) released into the primary loop, such as the injection flow rate of the safety injection tank drain or the Emergency Core Cooling System (ECCS). h_water represents the specific enthalpy (J / kg) of the released liquid, indicating the energy possessed by the released liquid. Injection refers to the relevant operating parameters of the water injection system (such as the high-pressure safety injection system and the low-pressure safety injection system). T_inj represents the initiation time of the water injection action (unit: seconds), indicating the time interval from the initial moment of the accident to the start of water injection. Q_inj represents the water injection flow rate (unit: kg / s), indicating the mass of coolant injected into the primary circuit per unit time. P_inj represents the water injection pressure (unit: Pa), indicating the pressure of the coolant at the injection point, affecting the injection capacity and flow direction. Temp_inj represents the water injection temperature (unit: K), indicating the temperature of the injected coolant, affecting the overall thermal balance of the primary circuit and potential thermal shock. Model refers to the model-related parameters required during the modeling process. These parameters may not be actual equipment parameters, but are used to correct or calibrate the model to approximate physical reality. D_nozzle is the equivalent diameter of the injection nozzle (in meters), which affects the jet pattern and mixing efficiency of the injected fluid. V_nozzle is the initial velocity of the injection fluid at the nozzle outlet (in meters per second), which affects the penetration depth and momentum exchange of the jet.

[0159] To comprehensively cover the parameter space and ensure sampling efficiency, Latin hypercube sampling (LHS) is employed to sample all continuous parameters in the aforementioned input parameter set X. LHS is a stratified sampling technique that divides the range of each parameter into N equally probable intervals, randomly selects a value from each interval, and then randomly combines these values ​​into N sets of working conditions. This method ensures uniform coverage across each parameter dimension and minimizes the correlation between samples, thus efficiently representing the uncertainty of the entire parameter space with a smaller sample size.

[0160] After randomly generating N sets of operating condition combinations, each set of operating conditions is subjected to transient simulation calculations using a validated high-precision thermal-hydraulic system program. The simulation calculations accurately solve the mass, momentum, and energy equations of the primary loop system, obtaining the dynamic impact history of water injection operations on key state parameters (pressure, water level, and temperature) of the primary loop.

[0161] After each set of operating conditions is simulated, its time series data is output, including:

[0162]

[0163] Wherein, P(t) is the pressure change over time at key points in the primary loop system (such as the pressurizer) (unit: Pa); L(t) is the water level change over time at the primary loop system (such as the pressurizer water level or the core water level) (unit: m); T(t) is the temperature change over time at key points in the primary loop system (such as the core outlet or cold tube section) (unit: K); and Q_inj(t) is the water injection flow rate change over time (unit: kg / s), reflecting the execution process of the water injection strategy. These input-output data pairs constitute the basic dataset for subsequent machine learning model training and validation.

[0164] Step 3: LSTM Deep Learning Prediction for One-Loop Water Injection

[0165] Under nuclear power plant accident conditions, there is a strong nonlinear coupling relationship between pressure transients, water level drops, and temperature changes in the primary loop system. Traditional simplified models or empirical formulas struggle to accurately capture this dynamic characteristic, leading to biases in the prediction of water injection effectiveness. To address this issue, this application proposes a Long Short-Term Memory (LSTM) deep learning model specifically optimized for primary loop water injection scenarios. This model accurately predicts the changing trends of the primary loop state and the effectiveness of water injection behavior within a short future period.

[0166] (1) Input feature design and scene adaptation

[0167] The model's input features are closely built around the real-time interaction between the primary loop state and the water injection operation. The input is a fixed-length historical time-series window containing time-series data of multiple key parameters of the primary loop over a past period, designed to provide the model with sufficient contextual information to learn system dynamics. The input time-series window structure is defined as follows:

[0168]

[0169] Where: x_t represents the feature vector at time step t, and its specific composition is as follows:

[0170]

[0171] Here, P(t) is the pressure (Pa) at time t, L(t) is the water level (m) at time t, T(t) is the temperature (K) at time t, and Q_inj(t) is the water injection flow rate (kg / s) at time t. These parameters together characterize the overall state of the primary loop and the water injection operation being performed at time t.

[0172] k is the length of the time window, i.e., the number of historical time steps the model reviews. In this embodiment, after analyzing the typical dynamic response time of a single loop, k=10 is set. This means that when the model makes predictions, it will consider the system state and operational history of the past 10 time steps, which is long enough to cover the critical delay period from the occurrence of water injection to the initial response of the system.

[0173] Δt is the interval between adjacent time steps, i.e., the sampling step size. In this embodiment, Δt is set to 1 minute. The choice of this step size balances two aspects: first, it requires sufficiently dense sampling to capture rapid dynamics (such as changes in the initial stage of a pressure drop); second, it needs to consider the model's computational efficiency and real-time requirements. A step size of 1 minute is considered to be able to effectively capture the main dynamic characteristics under primary loop fault conditions, while also meeting the timeliness requirements of online prediction.

[0174] (2) LSTM network architecture and specific optimization

[0175] Although LSTM itself can learn time dependencies, in primary circuit accident management, water level is the most critical parameter characterizing the core cooling state and directly related to nuclear safety. The model needs to have higher sensitivity and attention to changes in water level, especially low or rapidly decreasing trends. Therefore, this application's embodiment introduces a WaterLevelAttentionSub-module into the network. This module dynamically calculates the attention weights:

[0176]

[0177] Here, h_t represents the LSTM hidden state at the current time step t, encoding the sequence information learned by the model up to time t. L(t) is the real-time water level value (m) at time step t. [h_t;L(t)] represents concatenating the hidden state vector with the water level scalar value to form a new vector. W_h and b_h are the trainable weight matrix and bias vector of this attention submodule. σ is the sigmoid activation function, which compresses the output to the (0,1) interval. The closer its value is to 1, the more important the historical information at that moment is for dealing with the current water level state, thus amplifying the importance of features highly correlated with the critical water level state, making the model prediction more focused on risk avoidance.

[0178] The model's output objective (i.e., the prediction task) is defined as three key metrics for the next 5 minutes (i.e., the next 5 time steps, since Δt = 1 min):

[0179] ΔP: The rate of change of pressure (Pa / min) over the next 5 minutes, indicating whether the pressure tends to stabilize or deteriorate.

[0180] ΔL: The rate of change of water level (m / min) in the next 5 minutes, which directly indicates whether the water level is rising or continuing to fall, and is a core indicator of safety.

[0181] E_inj: Water Injection Efficiency Score ((m / kg)*s), is a comprehensive performance index defined as E_inj = ΔL / Q_inj (average). This index quantifies the water level rise effect brought about by a unit water injection flow rate. The higher the E_inj value, the better the water injection efficiency of the current water injection strategy (or operating condition). This index provides an economic basis for subsequent optimization decisions.

[0182] (3) Specialized training strategies based on water level risk

[0183] To improve the prediction accuracy and reliability of the model in high-risk scenarios (i.e., critical water level scenarios), two specific strategies were adopted during the model training phase:

[0184] Oversampling: When preparing the training dataset, duplicate and oversample the samples with water level readings L≤1.5m (this threshold can be set according to the specific pile type). This increases the proportion of low water level samples in the training data, forcing the model to pay more attention to these high-risk conditions and optimizing its prediction performance in such scenarios.

[0185] Hybrid Loss Function: The design of this loss function considers the accuracy of multiple prediction targets simultaneously.

[0186]

[0187] Wherein: MSE(ΔP,ΔL) is the Mean Squared Error, calculated as the average of the sum of squared errors between the predicted and actual values ​​of the pressure change rate (ΔP) and water level change rate (ΔL), respectively. MSE penalizes larger errors more severely, helping the model prioritize the accuracy of the predictions of the two key safety parameters, ΔP and ΔL. MAE(E_inj) is the Mean Absolute Error, calculated as the average of the absolute values ​​of the differences between the predicted and actual values ​​of the injection efficiency score (E_inj) and E_inj_true. MAE penalizes errors relatively linearly. A coefficient of 0.5 is used to balance the magnitude and importance of the MSE and MAE terms, indicating that the training process slightly emphasizes the accurate prediction of ΔP and ΔL, while also constraining the prediction bias of E_inj.

[0188] Through the network structure, attention mechanism, and training strategy designed above, this LSTM model can accurately predict the short-term evolution trend of the first-loop state and the water injection effect, providing a reliable basis for real-time optimization decisions.

[0189] Step 4: PPO reinforcement learning water injection optimization decision

[0190] In handling primary loop accidents at nuclear power plants, water injection decisions face a multi-objective optimization challenge: rapidly alleviating pressure and water level crises to ensure safety, while cautiously utilizing limited makeup water resources to avoid premature depletion. Traditional rule-based control or simple optimization algorithms struggle to dynamically balance these often conflicting objectives. This application proposes a reinforcement learning agent based on a proximal policy optimization algorithm, enabling it to learn through interaction with the primary loop simulation environment, ultimately acquiring a safe, efficient, and robust water injection decision strategy.

[0191] (1) State space and risk quantification

[0192] The state s_t observed by the reinforcement learning agent at each decision step needs to contain sufficient information to characterize the current state of the system and available resources. This application defines the state space as follows:

[0193]

[0194] Where: P(t): Current primary loop pressure (Pa), reflecting the system's pressure status. L(t): Current primary loop water level (m), reflecting the coolant charge and a core safety parameter. T(t): Current primary loop average temperature or critical point temperature (K), affecting system pressure and thermal stress. Q_inj(t): Current water injection flow rate (kg / s), representing the current action taken. W(t): Current remaining available emergency water replenishment resources (kg), representing constraints on future actions and directly reflecting economic considerations.

[0195] In order to transform continuous system state parameters (P, L) into quantitative indicators characterizing the degree of risk, the embodiments of this application define the following two risk coefficients:

[0196] Pressure risk coefficient R_p:

[0197]

[0198] Wherein, P(t): the current real-time pressure measurement (Pa). P_design: the design pressure of the primary loop system (Pa), which is the upper limit of the pressure that the system can safely withstand over a long period of time. P_critical: the critical pressure risk value of the primary loop (Pa), usually set to a value lower than P_design (e.g., 80% of P_design). When P(t) is lower than P_critical, the pressure risk is considered very low (R_p=0); when P(t) is higher than P_critical, R_p begins to increase linearly from 0; when P(t) reaches or exceeds P_design, R_p reaches its maximum value of 1.0. This coefficient quantifies the level at which the current pressure approaches a dangerous level.

[0199] Water level risk coefficient R_L:

[0200] Where L(t): the current real-time water level measurement (m). L_critical: the critical water level risk value (m), typically a low absolute water level value (e.g., 4.0m) or a relative value (e.g., the water level corresponding to exposed reactor core). When L(t) is higher than L_critical, the water level risk is considered very low (R_L=0); when L(t) is lower than L_critical, R_L begins to increase linearly from 0; when L(t) drops to 0, R_L reaches its maximum value of 1.0. This coefficient quantifies the risk level caused by the current low water level.

[0201] (2) Design of multi-objective reward function

[0202] The reward function is the core element guiding the learning direction of the agent. The reward function designed in this embodiment comprehensively considers both safety and economy objectives:

[0203] Safety reward R_safe: Safety is the highest priority. The reward function incentivizes the agent to avoid risks by penalizing high-risk agents.

[0204]

[0205] This reward is always negative or zero; the higher the risk, the greater the penalty (negative reward).

[0206] Using the squared terms (R_p², R_L²) means that when the risk coefficient is close to 1 (high risk), the penalty will increase sharply, which forces the agent to try its best to avoid entering the high-risk area, which is in line with the nuclear safety principle.

[0207] The weight of the water level risk coefficient R_L (-15) is higher than the weight of the pressure risk coefficient R_p (-10). This is because, in typical accident scenarios, the consequences of core exposure and melting caused by excessively low water levels are generally considered to be more severe and irreversible than short-term overpressure (as long as it does not exceed the pressure limit).

[0208] Economic incentive R_econ: Resource utilization should be optimized while ensuring security.

[0209]

[0210] W: Remaining water resources (kg). This term is positive, encouraging the agent to maintain a high water reserve to prepare for long-term accidents. -0.2*(Q_inj)²: This term is negative, representing a penalty for large-volume water injection. Using a squared term means that the larger the flow rate, the faster the penalty increases. This encourages the agent to use the smallest possible, smoothest water injection flow rate while meeting safety requirements, avoiding drastic flow fluctuations and excessive consumption. The coefficient 0.1 is used to adjust the magnitude of the overall economic reward term, ensuring it does not overwhelm the safety reward and that safety is always the top priority. Total reward R_total: The total reward obtained by the agent at each step is the sum of the safety reward and the economic reward.

[0211]

[0212] The design of this reward function gives the agent dynamic priorities when making decisions: when the risk coefficient is high (R_p or R_L is large), R_safe (a large negative value) dominates the total reward, and the agent's primary goal is to take action (increase the injection) to reduce the risk; when the risk coefficient is low, R_safe is close to zero, R_econ begins to play a greater role, and the agent will tend to adopt a more resource-efficient strategy.

[0213] (3) Implementation of PPO algorithm

[0214] This application employs the Proximal Policy Optimization (PPO) algorithm as the training framework for reinforcement learning agents, which has been widely validated for its advantages in training stability, sampling efficiency, and handling of continuous action spaces. This algorithm belongs to existing general technologies, and its core idea is to ensure the stability of the training process and avoid drastic fluctuations in policy performance by limiting the magnitude of policy updates.

[0215] In implementation, the agent's policy network π_θ(a_t|s_t) takes the state s_t as input, parameterizes a stochastic policy, and outputs the probability distribution of action a_t (i.e., the adjustment amount of the water injection flow). The algorithm updates the parameters θ of the policy network by maximizing a clipped objective function:

[0216]

[0217] Where: E_t is the expectation. r_t(θ) = π_θ(a_t|s_t) / π_{θ_old}(a_t|s_t), representing the probability ratio of choosing action a_t in state s_t between the old and new policies. A_t is the advantage function, used to evaluate the advantage of taking action a_t in state s_t compared to the average policy. ε is a hyperparameter used to define the pruning range, limiting the fluctuation range of the probability ratio r_t(θ), thereby preventing excessively large single policy updates. The advantage function A_t is usually calculated using the generalized advantage estimation (GAE) method to balance the bias and variance of the estimation. The value function network V_φ(s_t) is trained simultaneously to fit the state values ​​and provide a benchmark for the calculation of the advantage function.

[0218] Step 5: Real-time Water Injection Decision-Making Plan

[0219] The ultimate goal of this application is to develop an intelligent decision-making system that can be used to guide the handling of primary circuit accidents in nuclear power plants in real time. This system integrates three core technologies: high-speed synchronous acquisition of multi-source data, a hierarchical response decision-making mechanism, and online adaptive learning, to achieve accurate, reliable, and efficient water injection control under accident conditions.

[0220] (1) Multi-source data synchronization and processing

[0221] Sensor network deployment and data acquisition: High-frequency, high-precision sensors are deployed at key locations in the primary loop system to form a real-time data acquisition network, mainly including:

[0222] Pressure sensor: The range covers the expected operating pressure range of the primary circuit (e.g., 0~20MPa), accurately measuring regulator pressure, hot and cold section pressure, etc.

[0223] Water level gauges (e.g., differential pressure transmitters) accurately measure the water level of a pressure regulator, with a range covering the entire variation range.

[0224] Thermocouples: Their range covers the possible temperature range under accident conditions (e.g., 0~400°C), measuring the temperature at critical locations such as the core outlet and hot / cold sections.

[0225] Data synchronization and preprocessing: All sensor data is synchronously acquired through a high-bandwidth, low-latency data acquisition system, with a sampling frequency that meets the requirements for dynamic process capture (e.g., 10Hz). The acquired raw data undergoes preprocessing steps such as filtering, outlier removal, and unit conversion to form a clean, consistent state vector s_t that can be used as model input.

[0226] (2) Hierarchical response dynamic decision-making mechanism

[0227] The system switches between different operating modes based on the current real-time risk assessment level to achieve an optimal balance between security and response speed.

[0228] Normal mode (steady-state or low-risk control): When the system pressure and water level are both within safe ranges (e.g., R_p<0.3, R_L<0.3), the system operates in normal mode.

[0229] Decision frequency: A decision cycle is triggered once every Δt = 1 second.

[0230] Decision-making process:

[0231] The latest preprocessed state data s_t is input into the trained LSTM prediction model. The LSTM model predicts the pressure change trend ΔP, water level change trend ΔL, and injection efficiency E_inj for the next 15 minutes. The current state s_t and the LSTM prediction results (as additional state information) are then input into the trained PPO agent policy network. The PPO policy network outputs a suggested injection flow rate adjustment: ΔQ_inj = π(s_t) (unit: kg / s). The system sends ΔQ_inj to the control system for execution or provides it to the operator as a decision suggestion.

[0232] Emergency Mode (Rapid Intervention):

[0233] When real-time monitoring data indicates a rapid deterioration in the system's condition, and it is about to enter or has already entered a high-risk zone, the emergency mode should be triggered immediately. Triggering conditions:

[0234] P(t)>0.8*P_critical (pressure too low, seriously endangering integrity) or L(t)<1.2*L_critical (water level too low, seriously endangering core cooling).

[0235] Once any condition is met, the system immediately executes the preset maximum water injection logic to inject coolant as quickly as possible to prevent the accident from worsening.

[0236] ΔQ_inj(t)=min(k*(Q_max - Q_inj(t)), 0) (usually positive, indicating an increased flow rate)

[0237] Where: Q_max is the maximum flow rate (kg / s) that the water injection system can provide, which is the physical upper limit of the system design. k is a gain coefficient (0 < k ≤ 1), used to control the rate of flow increase and avoid excessive impact on the system.

[0238] The emergency mode has the highest priority to ensure absolute priority of safety in critical situations.

[0239] (3) On-site feedback and online model fine-tuning

[0240] To address the possible differences (model errors) between the actual power plant and the simulation model, as well as the possible aging or characteristic changes of equipment over time, the system integrates an online learning and adaptive mechanism:

[0241] Verification of execution effect:

[0242] After each water injection action (or a series of actions) is executed, the system continuously collects the state response data of the actual primary loop, for example, records {P_real(τ), L_real(τ), T_real(τ)} in the next few minutes.

[0243] Calculation of prediction error:

[0244] Compare the actually observed state changes (such as ΔP_real, ΔL_real) with the corresponding predictions (ΔP_pred, ΔL_pred) made by the previous LSTM model, and calculate the prediction error:

[0245] e = (ΔP_pred - ΔP_real)² + (ΔL_pred - ΔL_real)²

[0246] This error quantifies the deviation between the current model prediction and the actual system behavior.

[0247] Online model fine-tuning: Set an error threshold (for example, e_threshold = 10%). If the calculated error e > e_threshold, trigger the incremental learning mechanism: The system takes the latest actual operation data (state, action, result) as a new training sample. With a small learning rate, use this new sample (or a small batch containing such new samples) to fine-tune the parameters of the LSTM prediction model and the PPO policy network.

[0248] In an exemplary embodiment, the above method further includes:

[0249] Step 1: Obtain multiple sets of training samples;

[0250] Step 2 involves oversampling training samples with water levels below the risk threshold and using a hybrid loss function to backpropagate and optimize the model parameters to update them until the iterative calculation is completed and the prediction model is obtained. The hybrid loss function is used to constrain the prediction errors of pressure change rate, water level change rate, and water injection efficiency score.

[0251] Step 3: Establish a primary loop water injection simulation environment based on the prediction model and the primary loop parameterized model;

[0252] Step 4: Obtain the state space, action space, and reward function of the primary water injection simulation environment, and construct an untrained initial model. The state space is used to characterize the state of the simulation environment at any given time. The state space includes at least the pressure change rate, water level change rate, and water injection efficiency score output by the prediction model, as well as the remaining available water replenishment resources. The reward function is used to determine the immediate reward values ​​of the safety reward item and the economic reward item based on the change of the state value after the action is performed.

[0253] Step 5: Obtain the current state observation value output by the primary loop water injection simulation environment, input the current state observation value into the initial model, and use the initial model to output the candidate water injection adjustment amount based on the current state observation value;

[0254] Step 6: Input the candidate water injection adjustment amount into the primary loop water injection simulation environment for execution, and obtain the next state observation value fed back by the simulation environment and the instantaneous reward calculated by the reward function;

[0255] Step 7: Based on the current state observation, candidate water injection adjustment amount, next state observation and immediate reward, iteratively update the model parameters to learn the target water injection decision strategy that maximizes the cumulative reward.

[0256] Step 8: Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters;

[0257] Step 9: If neither the pressure risk coefficient nor the water level risk coefficient reaches the risk threshold, determine the weight values ​​of the operating state parameters at different historical moments in the state parameter sequence based on the water level values ​​in the state parameter sequence.

[0258] Step 10: Weight the weight values ​​and the corresponding running state parameters to obtain weighted time series features;

[0259] Step 11: Based on the weighted time series characteristics, predict the pressure change rate and water level change rate of the primary loop system within a future set time period, and determine the water level change amount based on the water level change rate.

[0260] Step 12: Determine the water injection efficiency score based on the proportional relationship between the water level change and the corresponding average water injection flow rate within a future set time period; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical time period;

[0261] Step 13: Determine the pressure risk coefficient based on the current pressure value and candidate water injection adjustment amount in the operating status parameters; and determine the water level risk coefficient based on the current water level value and candidate water injection adjustment amount in the operating status parameters; wherein, when determining the pressure risk coefficient and the water level risk coefficient, the pressure risk coefficient is verified by the pressure change rate in the prediction results, and the water level risk coefficient is corrected by the water level change rate in the prediction results.

[0262] Step 14: Determine the safety assessment value based on the pressure risk coefficient and the water level risk coefficient;

[0263] Step 15: Determine the economic evaluation value based on the remaining available water replenishment resources and the current water injection flow rate in the operating status parameters;

[0264] Step 16: The candidate water injection adjustment amount that maximizes the comprehensive evaluation value composed of the safety assessment value and the economic assessment value is determined as the target water injection adjustment amount.

[0265] Step 17: Perform water injection operation on the primary loop system based on the target water injection adjustment amount.

[0266] Step 18: If the pressure risk coefficient and water level risk coefficient reach or exceed the risk threshold, switch to emergency intervention mode; emergency intervention mode is used to execute the preset maximum flow water injection command.

[0267] Step 19: After performing the water injection operation, obtain the actual state response results of the primary loop system;

[0268] Step 20: Determine the error between the actual state response result and the predicted result;

[0269] Step 21: If the error exceeds the preset error threshold, the sample containing the actual state response result and the corresponding water injection operation is determined as a new training sample.

[0270] The steps involve using newly added training samples to update the prediction model and / or adjust the model parameters in an incremental learning manner.

[0271] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0272] Based on the same inventive concept, this application also provides a water injection decision device for a nuclear power plant primary loop system to implement the water injection decision method for the nuclear power plant primary loop system described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the water injection decision device for a nuclear power plant primary loop system provided below can be found in the limitations of the water injection decision method for a nuclear power plant primary loop system described above, and will not be repeated here.

[0273] In one exemplary embodiment, such as Figure 9 As shown, a water injection decision-making device for the primary loop system of a nuclear power plant is provided, comprising: an acquisition module 901, a prediction processing module 902, a data processing module 903, and an execution module 904, wherein:

[0274] The acquisition module 901 is used to acquire the operating status parameters of the primary loop system and determine the safety risk coefficient of the primary loop system based on the operating status parameters.

[0275] The prediction processing module 902 is used to input the state parameter sequence into a pre-trained prediction model for prediction processing when the safety risk coefficient does not reach the preset risk threshold, and obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period.

[0276] The data processing module 903 is used to input the pre-acquired remaining available water replenishment resources, operating status parameters and prediction results into the pre-trained adjustment model for data processing to obtain the target water injection adjustment amount;

[0277] The execution module 904 is used to perform water injection operations on the primary loop system based on the target water injection adjustment amount.

[0278] In an exemplary embodiment, the prediction processing module 902 is specifically used to determine the weight values ​​of the operating state parameters at different historical moments in the state parameter sequence based on the water level values ​​in the state parameter sequence; to perform weighted processing on the weight values ​​and the corresponding operating state parameters to obtain weighted time series features; and to determine the prediction result based on the weighted time series features.

[0279] In an exemplary embodiment, the above prediction results include at least the pressure change rate, the water level change rate, and the water injection efficiency score. The prediction processing module 902 is specifically used to predict the pressure change rate and the water level change rate of the primary loop system in a future set time period based on weighted time series characteristics, and determine the water level change amount based on the water level change rate; and determine the water injection efficiency score based on the proportional relationship between the water level change amount and the corresponding average water injection flow rate in the future set time period.

[0280] In an exemplary embodiment, the water injection decision device of the primary loop system of the nuclear power plant is specifically used to acquire multiple sets of training samples; to oversample the training samples whose water level is lower than the risk threshold, and to use a hybrid loss function to backpropagate and optimize the model parameters in order to update the model parameters until the iterative calculation is completed and the prediction model is obtained; wherein, the hybrid loss function is used to constrain the prediction error of pressure change rate, water level change rate, and water injection efficiency score.

[0281] In an exemplary embodiment, the data processing module 903 is specifically used to determine a pressure risk coefficient based on the current pressure value and candidate water injection adjustment amount in the operating status parameters; and to determine a water level risk coefficient based on the current water level value and candidate water injection adjustment amount in the operating status parameters; wherein, when determining the pressure risk coefficient and water level risk coefficient, the pressure risk coefficient is verified using the pressure change rate in the prediction results, and the water level risk coefficient is corrected using the water level change rate in the prediction results; a safety assessment value is determined based on the pressure risk coefficient and water level risk coefficient; an economic assessment value is determined based on the remaining available water replenishment resources and the current water injection flow rate in the operating status parameters; and the candidate water injection adjustment amount that maximizes the comprehensive assessment value composed of the safety assessment value and the economic assessment value is determined as the target water injection adjustment amount.

[0282] In an exemplary embodiment, the water injection decision device for the primary loop system of the nuclear power plant is specifically used to establish a primary loop water injection simulation environment based on a prediction model and a primary loop parameterized model; obtain the state space, action space, and reward function of the primary loop water injection simulation environment, and construct an untrained initial model; wherein, the state space is used to characterize the state of the simulation environment at any given time, and the state space includes at least the pressure change rate, water level change rate, and water injection efficiency score output by the prediction model, as well as the remaining available makeup water resources; the reward function is used to determine the immediate reward values ​​of the safety reward item and the economic reward item based on the change of the state value after the action is performed; the model is trained in the primary loop water injection simulation environment, and an adjusted model is obtained after the training is completed.

[0283] In an exemplary embodiment, the water injection decision device for the primary loop system of the nuclear power plant is specifically used to acquire the current state observation value output by the primary loop water injection simulation environment, input the current state observation value into the initial model, and use the initial model to output candidate water injection adjustment amounts based on the current state observation value; input the candidate water injection adjustment amounts into the primary loop water injection simulation environment for execution, and acquire the next state observation value fed back by the simulation environment and the immediate reward calculated by the reward function; based on the current state observation value, candidate water injection adjustment amount, next state observation value, and immediate reward, iteratively update the model parameters to learn the target water injection decision strategy that maximizes the cumulative reward.

[0284] In an exemplary embodiment, the aforementioned safety risk coefficients include pressure risk coefficients and water level risk coefficients. The aforementioned water injection decision device for the primary loop system of the nuclear power plant is specifically used to perform water injection operations on the primary loop system based on the target water injection adjustment amount when neither the pressure risk coefficient nor the water level risk coefficient reaches the risk threshold.

[0285] In an exemplary embodiment, the water injection decision device of the primary loop system of the nuclear power plant is specifically used to switch to emergency intervention mode when the pressure risk coefficient and water level risk coefficient reach or exceed the risk threshold; the emergency intervention mode is used to execute a preset maximum flow water injection command.

[0286] In an exemplary embodiment, the water injection decision device for the primary loop system of the nuclear power plant is specifically used to obtain the actual state response result of the primary loop system after performing the water injection operation; determine the error between the actual state response result and the prediction result; if the error exceeds a preset error threshold, determine the sample containing the actual state response result and the corresponding water injection operation as a new training sample; and use the new training sample to update the prediction model and / or adjust the model parameters of the model in an incremental learning manner.

[0287] The modules in the water injection decision-making device of the primary loop system of the aforementioned nuclear power plant can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0288] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0289] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0290] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0291] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0292] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0293] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A water injection decision method for the primary loop system of a nuclear power plant, characterized in that, The method includes: Obtain the operating status parameters of the primary loop system, and determine the safety risk coefficient of the primary loop system based on the operating status parameters; If the safety risk coefficient does not reach the preset risk threshold, the state parameter sequence is input into a pre-trained prediction model for prediction processing to obtain the prediction result; wherein, the state parameter sequence includes multiple operating state parameters arranged in chronological order within a preset historical period. The remaining available water replenishment resources, the operating status parameters, and the prediction results are input into a pre-trained adjustment model for data processing to obtain the target water injection adjustment amount. The primary loop system is injected with water based on the target water injection adjustment amount.

2. The method according to claim 1, characterized in that, The step of inputting the state parameter sequence into a pre-trained prediction model for prediction processing to obtain the prediction result includes: Based on the water level values ​​in the state parameter sequence, determine the weight values ​​of the operating state parameters at different historical moments in the state parameter sequence; The weight values ​​are weighted together with the corresponding running state parameters to obtain weighted time series features; The prediction result is determined based on the weighted time series features.

3. The method according to claim 2, characterized in that, The prediction results include at least the pressure change rate, water level change rate, and water injection efficiency score. Determining the prediction results based on the weighted time-series characteristics includes: Based on the weighted time series characteristics, the pressure change rate and water level change rate of the primary loop system are predicted within a future set time period, and the water level change amount is determined based on the water level change rate. The water injection efficiency score is determined based on the proportional relationship between the water level change and the average water injection flow rate corresponding to the future set time period.

4. The method according to claim 1, characterized in that, The training process of the prediction model includes: Obtain multiple sets of training samples; Training samples with water levels below the risk threshold are oversampled, and a hybrid loss function is used to backpropagate and optimize the model parameters to update the model parameters until the iterative calculation is completed to obtain the prediction model; wherein, the hybrid loss function is used to constrain the prediction error of pressure change rate, water level change rate, and water injection efficiency score.

5. The method according to claim 1, characterized in that, The step of inputting the pre-acquired remaining available water replenishment resources, the operating status parameters, and the prediction results into a pre-trained adjustment model for data processing to obtain the target water injection adjustment amount includes: Based on the current pressure value and candidate water injection adjustment amount in the operating status parameters, a pressure risk coefficient is determined; and based on the current water level value and candidate water injection adjustment amount in the operating status parameters, a water level risk coefficient is determined; wherein, when determining the pressure risk coefficient and the water level risk coefficient, the pressure risk coefficient is verified using the pressure change rate in the prediction results, and the water level risk coefficient is corrected using the water level change rate in the prediction results; The safety assessment value is determined based on the pressure risk coefficient and the water level risk coefficient. The economic evaluation value is determined based on the remaining available water replenishment resources and the current water injection flow rate in the operating status parameters. The candidate water injection adjustment amount that maximizes the comprehensive evaluation value formed by the safety assessment value and the economic assessment value is determined as the target water injection adjustment amount.

6. The method according to claim 1, characterized in that, The training process of the adjusted model includes: Establish a primary loop water injection simulation environment based on the aforementioned prediction model and primary loop parameterized model; The state space, action space, and reward function of the first-loop water injection simulation environment are obtained, and an untrained initial model is constructed. The state space characterizes the state of the simulation environment at any given time, and includes at least the pressure change rate, water level change rate, and water injection efficiency score output by the prediction model, as well as the remaining available water replenishment resources. The reward function determines the immediate reward values ​​for safety and economic rewards based on the changes in state values ​​after an action is performed. The model is trained in the primary loop water injection simulation environment, and the adjusted model is obtained after the training is completed.

7. The method according to claim 6, characterized in that, The process of training the model in the primary loop water injection simulation environment and obtaining the adjusted model after training includes: Obtain the current state observation value output by the primary loop water injection simulation environment, input the current state observation value into the initial model, and use the initial model to output candidate water injection adjustment amounts based on the current state observation value; The candidate water injection adjustment amount is input into the primary loop water injection simulation environment for execution, and the next state observation value fed back by the simulation environment and the instantaneous reward calculated by the reward function are obtained. Based on the current state observation, the candidate water injection adjustment amount, the next state observation, and the immediate reward, the model parameters are iteratively updated to learn the target water injection decision strategy that maximizes the cumulative reward.

8. The method according to claim 1, characterized in that, The safety risk coefficient includes a pressure risk coefficient and a water level risk coefficient, and the method further includes: If neither the pressure risk coefficient nor the water level risk coefficient reaches the risk threshold, water injection operation is performed on the primary loop system based on the target water injection adjustment amount.

9. The method according to claim 8, characterized in that, The method further includes: If the pressure risk coefficient and the water level risk coefficient reach or exceed the risk threshold, switch to emergency intervention mode; the emergency intervention mode is used to execute a preset maximum flow water injection command.

10. The method according to claim 1, characterized in that, The method further includes: After performing the water injection operation, the actual state response result of the primary loop system is obtained; Determine the error between the actual state response result and the predicted result; If the error exceeds a preset error threshold, the sample containing the actual state response result and the corresponding water injection operation will be identified as a new training sample. The model parameters of the prediction model and / or the adjustment model are updated using the newly added training samples in an incremental learning manner.