Nuclear power plant safety assessment agent model based on adaptive construction of classifier and regression prediction algorithm and application thereof

CN118332410BActive Publication Date: 2026-08-21SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410603624.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2026-08-21
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

[0010](3)准确模拟核动力装置热工水力行为的系统仿真模型更精细化、复杂化,RISMC事件序列计算耗时长,全部事件序列仿真计算量激增,RISMC工程应用难以承受

Benefits of technology

[0036]上述模型创建的训练过程中,创新性的引入了分类器模型并经过训练可以获得超平面边界区域,在后续用于训练的数据采样时,通过在超平面边界区域附近的数据采样降低数据采样的数量。此外,代理模型中还引入了回归预测模型,该模型通过回归精度的计算和判断来决定代理模型的最终训练准确度,在不满足代理模型精度情况下,通过调用分类器模型超平面边界区域的采样数据能够实现自适应过程,从而缩小后续RELAP5程序的计算量,提高计算效率。训练支持向量机算法(SVM),可获得超平面,用以区分超限与否,达到分类的目的;通过超平面与支持向量间获得新数据点,用以迭代训练K最近邻算法(KNN),减少了数据需求量,提高训练效率与模型精度。训练获得K最近邻算法,用以预测超限与否,达到预测的效果,与传统的RISMC方法相比,计算效率大大提高。获得的代理模型,可预测并拟合出关键安全参数分布(如燃料包壳峰值温度(PCT))曲线,用于安全裕度分析。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118332410B_ABST
    Figure CN118332410B_ABST
Patent Text Reader

Abstract

The application discloses a nuclear power plant safety evaluation proxy model based on a classifier and a regression prediction algorithm, and a construction method thereof, which comprises the following steps: selecting safety accident uncertainty input parameters, constructing a physical model, obtaining a risk evaluation model, performing dimension reduction on sensitive analysis uncertainty parameters, running a RELAP5 program to obtain calculation results to construct an initial data set, training an initial classifier model and determining an initial hyperplane boundary of the classifier, obtaining a regression prediction model by using a regression prediction algorithm, judging whether the model accuracy meets the requirements, if not, sampling in the hyperplane boundary area by using the classifier model to generate a data set for running the RELAP5 program as the RELAP5 program parameter calculation results, and performing iterative calculation until the regression prediction model accuracy meets the specified requirements. The model obtains a hyperplane by using the classifier model to distinguish whether the limit is exceeded or not, so that the classification purpose is achieved; the regression prediction model is used to predict whether the limit is exceeded or not, so that the prediction effect is achieved. The proxy model can not only quickly calculate a mass accident scene consequence and a risk level, but also greatly improves the calculation efficiency of an accident scene with a long simulation calculation time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical field of this application belongs to nuclear power plant safety analysis, and in particular to the use of big data to generate analysis models for nuclear power plant safety assessment methods, namely, a nuclear power plant safety assessment proxy model and its application based on adaptive construction of classifiers and regression prediction algorithms. This model uses machine learning and adaptive sampling strategies to improve the computational efficiency of nuclear power safety assessment. Background Technology

[0002] Currently, nuclear energy, as a clean and efficient energy source, has been widely used in countries around the world. Nuclear safety is the most critical and important factor for the sustainable development of nuclear power. During the operation of nuclear power plants, analyzing and assessing potential safety accidents through simulation is an important way to ensure nuclear power safety. Generally, nuclear power plant safety analysis methods are mainly divided into two categories: deterministic safety analysis methods and probabilistic safety analysis methods. For deterministic safety analysis, my country's Nuclear Safety Administration and the International Atomic Energy Agency have released four safety analysis options, as shown in Table 1.

[0003] Table 1: Options for Deterministic Security Analysis

[0004]

[0005] Traditional deterministic safety analysis methods, employing concepts like defense-in-depth and single-failure principles, struggle to comprehensively assess nuclear power plant operating conditions and system equipment status, potentially exhibiting either excessive or insufficient conservatism. Conversely, static analysis methods used in probabilistic safety analysis (such as fault tree and event tree methods) are ill-suited for simulating dynamic time sequences. Furthermore, traditional safety analyses often only consider cognitive uncertainty, neglecting stochastic uncertainty. Consequently, the safety margins assessed by a single safety analysis method are insufficiently accurate. Therefore, internationally, the best estimation and uncertainty analysis methods of deterministic approaches have been integrated with probabilistic risk quantification techniques, resulting in a more advanced safety analysis technique—Risk-informed Safety Margin Characterization (RISMC). RISMC comprehensively considers different types of uncertainty and uses high-fidelity physical models and computer simulations to analyze accident probabilities and consequences. RISMC is part of the U.S. Department of Energy's Light Water Reactor Sustainability Project. The aim of this approach is to improve the safety and economics of nuclear power plants, thereby extending their operational lifespan. The advantage of the RISMC method lies in its ability to fully utilize existing physical and computational models, combined with high-performance computing technology, to conduct a more comprehensive and in-depth analysis of the risks of nuclear power plants, consider the impact of different types of uncertainties, improve the accuracy and credibility of risk assessment, and provide a more scientific and reasonable basis and guidance for the design, operation and management of nuclear power plants.

[0006] In the RISMC methodology, safety margin is defined by the relationship between a power plant’s capacity and its load capacity. Specifically, it is the probability that the power plant’s load will be lower than its capacity under a specific anticipated accident condition, i.e., safety margin = P (C>=L). Figure 1 The intersection of the two distribution curves visually illustrates the measurement of this safety margin. The surrogate model built based on the RISMC method holds promise for future applications in deeply exploring the impact of design and operational changes on the safety margin. For example, aging of power plant materials may cause a leftward shift in the power plant capacity curve, while increased burnup and reactor power may lead to a rightward shift in the load curve. Compared to traditional safety analysis methods, the RISMC method more closely reflects actual operating conditions and is more effective in uncovering potential margin space. Therefore, using a surrogate model based on the RISMC method can more accurately predict and manage the operating status and safety margin of power plants, thereby comprehensively integrating nuclear safety and optimizing the design, operation, and maintenance decisions for nuclear power plant risk guidance. In the field of nuclear safety, using surrogate models to quickly predict and assess the safety performance of nuclear facilities under different scenarios provides strong support for decision-making. Through surrogate models, it is easier to study the impact of various factors on the safety performance of nuclear facilities, thereby identifying potential safety risks and developing corresponding countermeasures.

[0007] RISMC accident scenario generation employs the DPRA method, primarily including Monte Carlo (MC) risk simulation, Dynamic Event Tree (DET) methods, hybrid MC and DET methods, and GO-FLOW methods. Among these, MC risk simulation and Discrete Dynamic Event Tree (DDET) methods have better engineering practicality. However, all of these require significant computational resources and time. Applying RISMC analysis to actual nuclear power plants still involves complex issues, particularly in RISMC computational efficiency, quantification and propagation of uncertainties (margin analysis), and intelligent decision guidance, which are mainly due to the following shortcomings:

[0008] (1) In the process of modeling complex accident scenarios, it is easy to cause an excessive number of accident sequences and branch explosion;

[0009] (2) In order to meet the high computational accuracy and high requirements for verification of RISMC models, a massive number of samples are needed to support the RISMC analysis of a single accident sequence.

[0010] (3) The system simulation model that accurately simulates the thermal-hydraulic behavior of nuclear power plants is more refined and complex. The calculation of RISMC event sequences takes a long time, and the amount of calculation for the simulation of all event sequences has increased dramatically, making it difficult for RISMC engineering applications to bear.

[0011] To simplify computation and improve efficiency, surrogate models are commonly used. A surrogate model (also called a meta-model or approximate model) is a mathematical model constructed without sacrificing accuracy. It has low computational cost, short computation time, and produces results close to those of numerical analysis or physical experiments. It can address the complexities of computation and information exchange in multidisciplinary optimization. The basic principles of surrogate models include two aspects: firstly, the selection of sample points for model construction, which falls under experimental design; and secondly, data fitting and prediction model building, which constitutes the core of the surrogate model. Representative surrogate models include RSM, Kriging, and RBF surrogate models. Based on the core idea of ​​global function approximation, surrogate models use mathematical approximation methods such as interpolation, fitting, and regression, based on a finite amount of sample data, to construct a mathematical approximation of a computationally expensive complex model. In computer simulations, surrogate models are used to simplify complex models, improving computational efficiency and reducing costs. It can be a simple formula, polynomial, interpolation function, neural network, etc.

[0012] This application proposes improvements to the design of the proxy model using big data technology, thereby enhancing the efficiency of nuclear power safety analysis and calculation. Summary of the Invention

[0013] The purpose of this application is to introduce classifiers and regression prediction algorithms into the design of a proxy model to construct an adaptive nuclear power plant safety assessment proxy model. The technical solution to achieve the above-mentioned objective is a nuclear power plant safety assessment proxy model adaptively constructed based on classifiers and regression prediction algorithms. This model is trained and constructed using the following method.

[0014] S1: Identify the nuclear safety issues of the nuclear power plant to be analyzed, clarify the risk-sensitive accident conditions, and select important uncertainty input parameters related to the safety accidents. The nuclear safety issues and risk-sensitive accident conditions of the nuclear power plant are obtained by reviewing, but not limited to, the power plant design report or power plant accident literature. Uncertainty input parameters include, but are not limited to, reactor power, breach size, core heat flux density parameters, their distribution characteristics, and sampling range parameters.

[0015] S2: Based on the selection results of step S1 and the accident conditions, construct a system-level physical model of the nuclear power equipment and its system risk assessment model. The construction of the system-level physical model of the nuclear power equipment is achieved through a system simulation program (this patent uses RELAP5 as an example). This model includes, but is not limited to, various transient analyses of light water reactor systems, loss-of-coolant accidents (LOCA), and anticipated transients (ATWS) models of failure to shut down in an emergency. It is not limited to component models of pipes, valves, pumps, thermal components, electric heaters, jet pumps, turbines, separators, safety injection tanks, and control system components, reactor neutron physics models, and special process models including, but not limited to, cross-sectional abrupt flow, crossflow models, choke flow, boron tracking, and non-condensable gas transport. The system risk assessment model is constructed by combining probabilistic safety assessment (PSA) to establish fault trees and event trees for quantitative analysis, and calculating the failure probability and risk level of the system.

[0016] S3: Based on the evaluation results of the system risk assessment model in step S2, sensitivity analysis is used to reduce the dimensionality of the uncertainty parameters in step S1, that is, the parameters with the greatest impact on safety indicators are selected as the feature values ​​of the dataset; the sensitivity analysis here is the Spearman correlation coefficient method, and the Spearman coefficient is calculated by the following formula: Spearman's correlation is a statistic that measures the correlation between two variables. It is based on the rank of the ordinal variable rather than its actual value; in the above formula, ρ is the Spearman correlation coefficient, which takes values ​​of [-1, +1], and R0 is the index of the variable. Xi and R Yi ρ and Y are the ranks of X and Y, respectively. Since X and Y are both ordinal variables, their ranks can be regarded as the relative positions of the observations. n is the total number of observation samples. The magnitude of the absolute value of ρ indicates the strength of the correlation between the input and output parameters. The larger the absolute value, the stronger the correlation between the parameters. The sign of the coefficient indicates whether the parameters are positively or negatively correlated. That is, if ρ>0, they are positively correlated, and vice versa. When ρ=0, the parameters are not correlated.

[0017] S4: Use the dimensionality-reduced data obtained in S3 to run the RELAP5 program to obtain high-fidelity calculation results; combine the dimensionality-reduced data, high-fidelity calculation results, and labels indicating whether the calculation results exceed the limits to form the initial dataset;

[0018] S5: Using the initial data from step S4, train the initial classifier model and determine the initial hyperplane boundary region of the classifier.

[0019] S6: Using the initial data and regression prediction algorithm from step S4, train and obtain the regression prediction model;

[0020] S7: Determine whether the accuracy of the regression prediction model obtained in step S6 meets the specified requirements. If it does, the entire surrogate model is completed; if it does not, proceed to the next step.

[0021] S8: Using the classifier model from step S5, sample the hyperplane boundary region to generate a dataset for running the RELAP5 program. This dataset is the result of calculations using the input parameters of the RELAP5 program. Repeat the S5-S8 process until the accuracy of the regression prediction model reaches the specified requirements, and obtain the final nuclear power plant safety assessment proxy model containing the classifier model and the regression prediction model.

[0022] The classifier model in S5 above is a Support Vector Machine (SVM) model. The specific training process for obtaining the hyperplane boundary region is as follows:

[0023] S51: Set the results to be divided into those exceeding the limit and those not exceeding the limit. Use "1" as the label for data exceeding the limit and "0" as the label for data not exceeding the limit.

[0024] S52: Given a binary classification training set D = [(x1,y1),(x2,y2),…,(x... m ,y m )], where X i ∈Rn is the eigenvector, Y i ∈{-1,+1} is the class label, and a separating hyperplane ω is found in the feature space. T The equation x+b=0 is given, which makes the two classes of samples located on opposite sides of the hyperplane, with the greatest distance to the hyperplane. These samples can be determined by solving the following system of equations:

[0025] Where: ω is the normal vector, which determines the direction of the hyperplane; b is the displacement term, which determines the distance between the hyperplane and the origin;

[0026] S53: To maximize the interval, we need to minimize ||ω|| 2 Simultaneously, the constraint condition y must be satisfied. i (ω T x i +b)≥1, meaning each sample is correctly classified;

[0027] S54: Using the Lagrange multiplier method, the above problem is transformed into a dual problem, namely, maximizing the Lagrange function: Where λ i ≥0 represents a Lagrange multiplier;

[0028] S55: For nonlinearly separable problems, the algorithm uses a Gaussian kernel function to map the nonlinearly separable problem from the original feature space to a higher-dimensional Hilbert space, thus transforming it into a linearly separable problem. The expression for the Gaussian kernel function is: Setting the partial derivatives of the Lagrange function with respect to ω and b to zero and introducing a Gaussian kernel function, the Lagrange function then transforms into: Where, λ i For Lagrange multipliers, the KKT conditions must be satisfied for inequality constraints, i.e., the following must be met.

[0029] S56: By solving the dual problem, we can obtain the optimal ω and b, as well as the optimal λ. i Thus, support vectors are used to construct the optimal hyperplane, i.e. the boundary region of the hyperplane.

[0030] The regression prediction algorithm in S6 is used to train and obtain the regression prediction model. The KNN model is constructed using the k-nearest neighbor (KNN) algorithm. The specific process is as follows:

[0031] S61: Set the results to be divided into those exceeding the limit and those not exceeding the limit. Use "1" as the label for data exceeding the limit and "0" as the label for data not exceeding the limit.

[0032] S62: Use the Minkowski distance to calculate the distance between each test sample and each training sample. The formula for calculating the Minkowski distance is as follows:

[0033] S63: Sort the distances from smallest to largest based on the calculated results, select the k nearest samples as the neighboring samples in the test samples, calculate the proportion of each type of sample in the k neighboring samples, and select the S type with the largest proportion as the predicted value output.

[0034] S7: Determine whether the accuracy of the KNN model meets the specified requirements. The judgment criteria can be: cross-validation average accuracy ≥ ε or new sample prediction accuracy ≥ δ for three consecutive times; here ε and δ are preset parameter values ​​less than 1, which are given by the security analyst.

[0035] This proxy model, by inputting nuclear power plant accident operating condition parameters, can predict whether the accident operating condition exceeds limits, the consequences of the accident, and determine whether the consequences are failure or success, thereby guiding the accident path evolution and providing early warning of accident consequences. This proxy model is also applied to risk-informed safety margin characterization (RISMC) in the implementation of nuclear power plant risk assessment.

[0036] In the training process of the aforementioned model, a classifier model is innovatively introduced. After training, the hyperplane boundary region can be obtained. During subsequent data sampling for training, the number of data samples is reduced by sampling data near the hyperplane boundary region. Furthermore, a regression prediction model is introduced into the surrogate model. This model determines the final training accuracy of the surrogate model through the calculation and judgment of regression accuracy. If the accuracy of the surrogate model is not met, an adaptive process can be achieved by calling the sampled data from the hyperplane boundary region of the classifier model, thereby reducing the computational load of the subsequent RELAP5 program and improving computational efficiency. Training the Support Vector Machine (SVM) algorithm yields a hyperplane used to distinguish between exceeding the limit and achieving classification. New data points are obtained between the hyperplane and support vectors to iteratively train the K-Nearest Neighbor (KNN) algorithm, reducing data requirements and improving training efficiency and model accuracy. The K-Nearest Neighbor algorithm, trained to predict whether the limit is exceeded, achieves predictive results, and its computational efficiency is significantly improved compared to the traditional RISMC method. The obtained surrogate model can predict and fit the distribution curves of key safety parameters (such as fuel cladding peak temperature (PCT)) for safety margin analysis. Attached Figure Description

[0037] Figure 1 A schematic diagram illustrating the concept of safety margin characteristics for risk guidance;

[0038] Figure 2 A schematic diagram illustrating the process of training a proxy model;

[0039] Figure 3 A schematic diagram for finding the hyperplane for Support Vector Machine (SVM);

[0040] Figure 4 A schematic diagram illustrating the KNN classification performance for different k values;

[0041] Figure 5 This is a schematic diagram of the limit surface obtained through the proxy model of this application. Detailed Implementation

[0042] The present invention will be further described below with reference to embodiments, but is not limited to the contents of the specification. It should be noted that the terms "upper," "lower," "front," "rear," "left," "right," "top," "bottom," "inner," "outer," "center," "vertical," and "horizontal," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the purpose of simplifying the description of the present invention, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0043] To clearly illustrate the proxy model construction process of the technical solution in this application, the following is a detailed explanation in conjunction with the appendix. Figure 2 Please explain. For example... Figure 2 As shown, this application employs a support vector machine and the k-nearest neighbor algorithm, combined with an adaptive sampling strategy to construct a surrogate model. Specific technical steps include:

[0044] S1: Identify the nuclear safety issues to be analyzed, clarify the risk-sensitive accident conditions, review the power plant design report or power plant accident literature, select important uncertainty input parameters related to the accident, such as reactor power, breach size or core heat flux density, and clarify the distribution characteristics and sampling range of these parameters.

[0045] S2: A system-level physical model is pre-constructed, research objectives are clarified, physical variables are selected, a mathematical model is determined, and an appropriate method is used to solve the mathematical model. Finally, the model is verified and optimized. Taking the RELAP5 program used in this application as an example, its modeling model mainly covers various transient analyses of light water reactor systems, including loss-of-coolant accidents (LOCA), anticipated transients without emergency shutdown (ATWS), and other transients such as feedwater loss, loss of external power supply, plant-wide power outage, and turbine tripping. RELAP5 component models include pipes, valves, pumps, thermal components, heaters, jet pumps, turbines, separators, safety injection tanks, and control system components. Furthermore, to more accurately simulate specific phenomena, RELAP5 also includes some special process models, such as abrupt cross-sectional flow, crossflow models, choke flow, boron tracking, and non-condensable gas transport.

[0046] S3: Construct a system risk assessment model. First, clarify the application goals and scope of the PSA model, such as assessing the safety risks of a nuclear power plant. Define the systems and processes to be analyzed, and the potential risk events of interest. For identifying potential risk events, analyze the operating principles of the system or process, historical data, and expert opinions to identify events or failure modes that may lead to risk. Classify and prioritize these events based on their impact on system safety. Next, build a fault tree, starting from the top event and analyzing all possible causes of the event layer by layer, constructing an event tree to describe a series of possible subsequent events after a specific event occurs, to assess the system's response and consequences under different scenarios. Use probability theory and mathematical statistics methods to quantitatively analyze the fault tree and event tree, calculating the system's failure probability and risk level. Finally, validate the model using actual operating data or new research findings.

[0047] S4: In nuclear power plant system simulation models, the sheer number of input parameters and their high coupling significantly increases the number of calculations required for system simulation, leading to a sharp rise in computational load. Furthermore, the excessively large dimensionality of the sample space in the uncertainty sampling process during surrogate model construction further exacerbates the explosive growth in training computation. In reality, some parameters have a relatively weak impact on key safety indicators (such as peak fuel cladding temperature). To effectively address these challenges, this paper employs sensitivity analysis to reduce the dimensionality of uncertainty parameters, such as the Spearman correlation coefficient method, a single sensitivity analysis approach, to accurately evaluate the sampled parameters and select the parameters with the greatest impact on safety indicators as the feature values ​​of the dataset. The Spearman coefficient is calculated using the following formula: Spearman's correlation is a statistic that measures the correlation between two variables. It is based on the rank of the ordinal variable rather than its actual value; in the above formula, ρ is the Spearman correlation coefficient, which takes values ​​of [-1, +1], and R0 is the index of the variable. Xi and R Yi ρ and Y are the ranks of X and Y, respectively. Since X and Y are both ordinal variables, their ranks can be regarded as the relative positions of the observations. n is the total number of observation samples. The magnitude of the absolute value of ρ indicates the strength of the correlation between the input and output parameters. The larger the absolute value, the stronger the correlation between the parameters. The sign of the coefficient indicates whether the parameters are positively or negatively correlated. That is, if ρ>0, they are positively correlated, and vice versa. When ρ=0, the parameters are not correlated.

[0048] S5: Modify the input card based on the obtained data, run the RELAP5 program to obtain high-fidelity calculation results, and label the data that exceeds the limit and those that do not (e.g., "1" indicates exceeding the limit and failure, "0" indicates safety). Use the key parameter data, high-fidelity calculation results, and data labels together as the initial dataset. Furthermore, the dataset is divided into a training set and a validation set according to a certain ratio. The training set should contain data with both exceeding and not exceeding the limit.

[0049] S6: In this application, the classifier is SVM. First, the initial training set is used to train the SVM model using the SVM training method described above, and the hyperparameters are optimized. The analyzed problem may not be linearly separable, therefore a suitable kernel function needs to be selected. The SVM model is constructed to classify the initial data, and simultaneously obtains the initial hyperplane and support vectors. A hyperplane is a linear subspace in n-dimensional Euclidean space with a co-dimensional dimension of one; that is, a hyperplane is an (n-1)-dimensional subspace in n-dimensional space. Specifically, a hyperplane in two-dimensional space is a line, while a hyperplane in three-dimensional space is a plane. This is a purely mathematical concept, not a real physical concept. A hyperplane can divide a linear space into two disjoint parts. In mathematics, the hyperplane H is a mapping subspace from n-dimensional space to (n-1)-dimensional space, defined by an n-dimensional vector and a real number. Let d be a non-zero vector in n-dimensional Euclidean space R, and a be a real number. Then, in R, the condition d... X The set of points X equal to a is called a hyperplane in R. The concept of a hyperplane has wide applications in machine learning, computer vision, and other fields. For example, in Support Vector Machines (SVMs), hyperplanes are used to separate data into different categories. Figure 3 As shown, in the two-dimensional space, in the support vector machine model, the hyperplane is the equation ω T The line corresponding to x+b=0. The specific process of obtaining the hyperplane boundary region through training is as follows:

[0050] S61: Set the results to be divided into those exceeding the limit and those not exceeding the limit. Use "1" as the label for data exceeding the limit and "0" as the label for data not exceeding the limit.

[0051] S62: Given a binary classification training set D = [(x1,y1),(x2,y2),…,(x... m ,y m )], where X i ∈Rn is the eigenvector, Y i ∈{-1,+1} is the class label, and a separating hyperplane ω is found in the feature space. T The equation x+b=0 is given, which makes the two classes of samples located on opposite sides of the hyperplane, with the greatest distance to the hyperplane. These samples can be determined by solving the following system of equations:

[0052] Where: ω is the normal vector, which determines the direction of the hyperplane; b is the displacement term, which determines the distance between the hyperplane and the origin;

[0053] S63: To maximize the interval, we need to minimize ||ω|| 2 Simultaneously, the constraint condition y must be satisfied. i (ω T x i+b)≥1, meaning each sample is correctly classified;

[0054] S64: Using the Lagrange multiplier method, the above problem is transformed into a dual problem, namely, maximizing the Lagrange function: Where λ i ≥0 represents a Lagrange multiplier;

[0055] S65: For nonlinearly separable problems, the algorithm uses a Gaussian kernel function to map the nonlinearly separable problem from the original feature space to a higher-dimensional Hilbert space, thus transforming it into a linearly separable problem. The expression for the Gaussian kernel function is: Setting the partial derivatives of the Lagrange function with respect to ω and b to zero and introducing a Gaussian kernel function, the Lagrange function then transforms into: Where, λ i For Lagrange multipliers, the KKT conditions must be satisfied for inequality constraints, i.e., the following conditions must be met.

[0056] S66: By solving the dual problem, we can obtain the optimal ω and b, as well as the optimal λ. i Therefore, support vectors are used to construct the optimal hyperplane, i.e., the boundary region of the hyperplane, such as... Figure 5 As shown.

[0057] S7: Train the KNN model using the initial training set, and use cross-validation to select the optimal k value and distance metric to achieve maximum accuracy. The accuracy measurement process focuses on the validation set. The trained KNN model is used to predict on the validation set, and the prediction results are compared with the actual high-fidelity simulation results on the validation set to obtain the accuracy of the KNN model. Figure 4 The diagrams illustrating the KNN classification performance for different k values ​​are provided. The specific process of constructing the KNN model is as follows:

[0058] S71: The results are divided into those exceeding the limit and those not exceeding the limit. "1" is used as the label for data exceeding the limit, and "0" is used as the label for data not exceeding the limit.

[0059] S72: Use the Minkowski distance to calculate the distance between each test sample and each training sample. The formula for calculating the Minkowski distance is as follows:

[0060] S73: Sort the distances from smallest to largest based on the calculated results, select the k nearest samples as the neighboring samples in the test samples, calculate the proportion of each type of sample in the k neighboring samples, and select the S type with the largest proportion as the predicted value output.

[0061] S8: Determine if the accuracy of the KNN model meets the specified requirements. The criteria can be: average cross-validation accuracy ≥ ε or prediction accuracy for new samples ≥ δ three consecutive times. For demonstration purposes, ε = 0.8 and δ = 0.95. If the accuracy requirement is met, the surrogate model training is considered complete. If the accuracy requirement is not met, proceed to the next step.

[0062] S9: Step 4 has obtained the hyperplane and support vectors through the SVM model. The coordinate indices of the support vector machine and the hyperplane are obtained, and new sample points are inserted between them to ensure that the sampling points are concentrated near the hyperplane. Simultaneously, another batch of sample points is randomly sampled from the entire space to prevent overfitting, achieving adaptive sampling. The input card is modified based on the newly generated data points, and the RELAP5 program is run to obtain high-fidelity calculation results for the new data points. Following the dataset processing method in Step 3, a new dataset is obtained for iterative training of the KNN model, improving its accuracy. Adaptive sampling further reduces the training data range and the number of training data points, improving model training efficiency.

[0063] S10: Repeat the above steps until the hyperplane, support vectors are stable, or the surrogate model prediction accuracy meets the requirements, at which point the surrogate model training is considered complete.

[0064] After the surrogate model is trained, it can achieve the following functions: Using the surrogate model, it can predict whether an accident condition exceeds limits. By inputting key parameters of the accident condition, the surrogate model can quickly predict the consequences of the accident condition, determining whether the consequence is failure or success, thus realizing accident path evolution guidance and accident consequence early warning functions. Based on the RISMC method, by inputting key parameters of the accident condition, the surrogate model can efficiently predict deterministic simulation results (such as directly providing the peak fuel cladding temperature in a small breach loss-of-water accident), and can efficiently fit the load distribution curve, combining it with the capacity distribution curve to perform dynamic probabilistic risk calculation. Safety margins are divided into deterministic safety margins (DSM) and probabilistic safety margins (PSM). A probabilistic safety margin is the probability that the load of a safety parameter is greater than the capacity, i.e., PSM = P(L>C). A deterministic safety margin is defined by the ratio or difference between capacity and load, i.e., DSM = LC or DSM = L / C. For the failure space and success space in the whole space, the hyperplane between them can be obtained by the surrogate model. Therefore, the surrogate model can realize the adaptive sampling function. Compared with the whole space sampling of Monte Carlo sampling and Latin hypercube sampling, the sample points of adaptive sampling are concentrated in the more sensitive region, that is, the region between the support vector and the hyperplane. This can reduce the number of samplings. At the same time, since the sample is concentrated in the more sensitive region, it has a greater impact on the model accuracy.

[0065] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is impossible to exhaustively list all embodiments here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms, characterized in that, The proxy model is constructed as follows; S1: Identify the nuclear safety issues of the nuclear power plant to be analyzed, clarify the risk-sensitive accident conditions, and select important uncertainty input parameters related to the safety accidents; S2: Based on the selection results of step S1 and the accident conditions, construct a system-level physical model of the nuclear power equipment and obtain its system risk assessment model; S3: Based on the evaluation results of the physical model and system risk assessment model in step S2, use sensitivity analysis to perform dimensionality reduction on the uncertainty parameters in step S1, that is, select the parameters that have the greatest impact on safety indicators as the feature values ​​of the dataset. S4: Use the dimensionality-reduced data obtained in S3 to run the RELAP5 program to obtain high-fidelity calculation results; combine the dimensionality-reduced data, high-fidelity calculation results, and labels indicating whether the calculation results exceed the limits to form the initial dataset; S5: Using the initial data from step S4, train the initial classifier model and determine the initial hyperplane boundary region of the classifier; the classifier model here is a Support Vector Machine (SVM) model, and the specific training process to obtain the hyperplane boundary region is as follows: S51: Set the results to be divided into those exceeding the limit and those not exceeding the limit. Use "1" as the label for data exceeding the limit and "0" as the label for data not exceeding the limit. S52: Given a binary classification training set D = [(x1,y1),(x2,y2),…,(x... m ,y m )] , where X i ∈R n It is the eigenvector, Y i Let ∈{-1,+1} be the class label, and find a dividing hyperplane in the feature space. T The equation x + b = 0 is given, such that the two classes of samples are located on opposite sides of the hyperplane, with the greatest distance to the hyperplane. These can be determined by solving the following system of equations: ,in: is the normal vector, which determines the direction of the hyperplane; b is the displacement term, which determines the distance between the hyperplane and the origin. S53: To maximize the interval, we need to minimize || || 2 Simultaneously, the constraint condition y must be satisfied. i ( T x i + b)≥ 1, meaning each sample is correctly classified; S54: Using the Lagrange multiplier method, the above problem is transformed into a dual problem, namely, maximizing the Lagrange function: , where λ i ≥0 represents a Lagrange multiplier; S55: For nonlinearly separable problems, the algorithm uses a Gaussian kernel function to map the nonlinearly separable problem from the original feature space to a higher-dimensional Hilbert space, thus transforming it into a nonlinearly separable problem. The expression for the Gaussian kernel function is: Let the Lagrange function be paired with... With the partial derivatives of b being zero and a Gaussian kernel function introduced, the Lagrange function is transformed into: ;in, i For Lagrange multipliers, the KKT conditions must be satisfied for inequality constraints, i.e., the following conditions must be met. ; S56: The optimal solution can be obtained by solving the dual problem. and b, and the optimal i Thus, the support vectors are used to construct the optimal hyperplane, that is, the boundary region of the hyperplane; S6: Use the initial data from step S4 to train a regression prediction model using a regression prediction algorithm; S7: Determine whether the accuracy of the regression prediction model obtained in step S6 meets the specified requirements. If it does, the entire surrogate model is completed; if it does not, proceed to the next step. S8: Using the classifier model from step S5, sample its hyperplane boundary region to generate a dataset for running the RELAP5 program. This dataset is used as input parameters for the RELAP5 program to calculate the results. Repeat the S5-S8 process until the accuracy of the regression prediction model reaches the specified requirements, and obtain the final nuclear power plant safety assessment proxy model containing the classifier model and the regression prediction model.

2. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 1, characterized in that, Nuclear safety issues and risk-sensitive accident conditions in S1 are obtained by reviewing, but not limited to, power plant design reports or power plant accident literature. Uncertainty input parameters include, but are not limited to, reactor power, breach size or core heat flux density parameters, their distribution characteristics and sampling range parameters.

3. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 1, characterized in that, In S2, the system-level physical model of nuclear power equipment is constructed using the RELAP5 program. This model includes, but is not limited to, various transient analyses of light water reactor systems, predicted transient models of loss-of-coolant accidents and failure to shut down in an emergency, and component models of piping, valves, pumps, thermal build-ups, point reactor dynamics, electric heaters, jet pumps, turbines, separators, safety injection tanks, and control system components, as well as special process models including, but not limited to, cross-sectional abrupt flow, crossflow models, choke flow, boron tracking, and non-condensable gas transport. The system risk assessment model in S2 is constructed by combining probabilistic safety assessment to establish a fault tree for quantitative analysis, calculating the system's failure probability and risk level.

4. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 1, characterized in that, The sensitivity analysis in step S3 uses the Spearman correlation coefficient method. The Spearman coefficient is calculated using the following formula: Spearman's correlation coefficient is a statistic that measures the correlation between two variables. It is based on the rank of the ordinal variable rather than its actual value. In the above formula, ρ is the Spearman correlation coefficient, which takes values ​​of [-1, +1], and R0 is the index of the variable. Xi and R Yi ρ represents the ranks of X and Y, respectively. Since X and Y are both ordinal variables, their ranks can be seen as the relative positions of the observations. n is the total number of observed samples. The absolute value of ρ indicates the strength of the correlation between the input and output parameters. The larger the absolute value, the stronger the correlation between the parameters. The sign of the coefficient indicates whether the parameters are positively or negatively correlated. That is, if ρ>0, they are positively correlated, and vice versa. When ρ=0, the parameters are uncorrelated.

5. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 1, characterized in that, The regression prediction algorithm in S6 is used to train a regression prediction model, which is a KNN model constructed using the k-nearest neighbor (KNN) algorithm. The specific process is as follows: S61: Set the results to be divided into those exceeding the limit and those not exceeding the limit. Use "1" as the label for data exceeding the limit and "0" as the label for data not exceeding the limit. S62: Use the Minkowski distance to calculate the distance between each test sample and each training sample. The formula for calculating the Minkowski distance is as follows: ; S63: Sort the distances from smallest to largest based on the calculated results, select the k nearest samples as the neighboring samples in the test samples, calculate the proportion of each type of sample in the k neighboring samples, and select the S type with the largest proportion as the predicted value output.

6. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 5, characterized in that, S7 determines whether the accuracy of the KNN model meets the specified requirements. The judgment criteria can be: the average accuracy of cross-validation ≥ ε or the prediction accuracy of new samples ≥ δ for three consecutive times; here ε and δ are preset parameter values ​​less than 1.

7. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to claim 1, characterized in that, During the S8 process of repeating S5-S8, the classifier model and regression prediction model are iteratively updated and trained to ensure the adaptive updating of the dataset running the RELAP5 program and the accuracy of the model.

8. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to any one of claims 1-7, characterized in that, The nuclear power plant safety assessment proxy model constructed using this method can predict whether the accident conditions exceed the limits, the consequences of the accident conditions, and judge whether the consequences are failure or success by inputting the nuclear power plant accident operating condition parameters, thereby realizing accident path evolution guidance and accident consequence early warning.

9. The method for adaptively constructing a proxy model for nuclear power plant safety assessment based on classifiers and regression prediction algorithms according to any one of claims 1-7, characterized in that, The nuclear power plant safety assessment proxy model constructed using this method is used to implement nuclear power plant risk assessment in the safety margin characteristic analysis of risk guidance.