Nuclear power plant equipment fault diagnosis method
Through the combination of KPCA and CNN-Bi-LSTM, the problem of data imbalance and scarcity of fault samples in complex operating conditions of nuclear power plant equipment is solved, and efficient fault identification and diagnosis is achieved.
Patent Information
- Application Number
- CN202510504745.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-29
AI Technical Summary
Nuclear power plant equipment faces the problems of data imbalance and scarce failure samples under complex working conditions, resulting in insufficient generalization ability and diagnostic accuracy of traditional fault diagnosis methods.
KPCA is used for nonlinear mapping and feature decomposition of data, combined with CNN-Bi-LSTM model for fault diagnosis, and transfer learning strategies are used to optimize the adaptability of the model under different operating conditions.
It improves the accuracy and robustness of fault identification of nuclear power plant equipment under complex operating conditions, especially in the diagnosis accuracy of data imbalance and sample scarcity scenarios.
Smart Images

Figure CN120386327A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nuclear power plant equipment safety, and more particularly, to a method for diagnosing faults of nuclear power plant equipment. Background Art
[0002] The state monitoring and fault diagnosis technology of static equipment in nuclear power plants is a key area in nuclear power safety management, and its importance is directly related to the safe and stable operation of nuclear power plants. With the increasing complexity of nuclear power plant equipment and the variability of operating environments, traditional fault diagnosis methods face challenges such as unbalanced data volumes and scarce fault types. Especially under different operating conditions, the fault samples of some equipment are few, resulting in poor generalization ability of diagnostic models. Therefore, how to accurately identify abnormal states and fault types of equipment under incomplete and unbalanced data has become an important issue for ensuring the safe operation of nuclear power plants.
[0003] At present, there is a relatively extensive research on the state monitoring method for nuclear power plants. Binsen et al. proposed a hybrid state monitoring method based on sparse autoencoder and isolation forest to address the problem that the complex operating conditions and numerous sensors in nuclear power plants result in high-dimensional data collection, making it difficult to conduct state monitoring. The experimental results show that the proposed method has achieved high monitoring accuracy on different operating condition datasets of nuclear power plants. Xueying et al. proposed a state monitoring method based on denoising autoencoder and one-class support vector machine to solve the problem that the high-dimensional feature parameter data obtained from nuclear power plants affects the subsequent state monitoring effect, greatly improving the accuracy of nuclear power plant monitoring. Wei et al. combined a statistical-based method with principal component analysis (PCA) to solve the problem of high false alarm rate in traditional state monitoring methods for nuclear power plants. The experimental results show that the proposed method greatly reduces the false alarm probability during the monitoring process. Shiqiao et al. proposed a monitoring algorithm based on denoising diffusion probability model to address the problem that it is difficult to conduct state monitoring using supervised learning algorithms due to the limited availability of labeled data in nuclear power plants. The experimental results show that compared with traditional state monitoring algorithms, this model has better robustness and higher monitoring accuracy. In summary, the existing state monitoring methods for nuclear power plants have played an important role in improving the reliability of equipment operation, but still face problems such as high data dimension, complex feature extraction, and limited abnormal detection accuracy. To address these challenges, this paper uses the KPCA method for state monitoring. Compared with traditional PCA, KPCA can map the original data to a high-dimensional feature space through a kernel function, extract more discriminative non-linear features in this space, thereby improving the accuracy and robustness of abnormal state detection. This characteristic makes KPCA particularly suitable for state monitoring under complex operating conditions of nuclear power plants, and can more effectively capture the non-linear relationships in the equipment operation data, thus enhancing the reliability of static equipment state monitoring in nuclear power plants.
[0004] In terms of fault diagnosis and transfer learning, Gui et al. proposed a hybrid transformer model to improve the accuracy of fault diagnosis in nuclear power plants. Compared with other deep learning fault diagnosis models, this model has a higher diagnostic accuracy and better local and global interpretability. Jie et al. proposed an interpretable artificial intelligence method based on game theory to address the poor interpretability of current deep learning models and applied it to fault diagnosis in nuclear power plants. Wenzhe et al. proposed a fault diagnosis method based on symmetric point graphs and residual neural networks for the problems of complex equipment, numerous compound faults, and difficult fault identification in nuclear power plants. The model was evaluated using multiple metrics, and the results showed that this method has a high diagnostic accuracy for compound fault diagnosis. Jiangkuan et al. proposed a transfer learning method based on maximum mean discrepancy and convolutional neural network (CNN) for the problems of numerous operating conditions, significant differences in data under different operating conditions, and poor generalization performance of existing fault diagnosis models in nuclear power plants. The results showed that this method exhibited high diagnostic accuracy in most transfer tasks. Zhichao et al. proposed a fault diagnosis method that combines deep learning and transfer learning for the problem of poor generalization performance of traditional diagnostic algorithms due to inconsistent fault data distributions under different operating conditions in nuclear power plants, and verified the effectiveness of this method using simulation data. In summary, existing fault diagnosis algorithms for nuclear power plants effectively improve the safety and reliability of nuclear power plants. However, in the case of data imbalance, scarce fault samples, and significant changes in operating conditions, the generalization ability and diagnostic accuracy of traditional methods still have certain limitations. Summary of the Invention
[0005] This specification provides a method for diagnosing faults in nuclear power plant equipment to overcome at least one technical problem existing in the related art.
[0006] The embodiments of the specification provide a method for diagnosing faults in nuclear power plant equipment, including:
[0007] S1. Data acquisition, including: obtaining normal state data and fault data of static equipment in a nuclear power plant under 100% and 80% operating conditions. The fault data includes steam generator heat transfer tube rupture, steam pipeline rupture, cold leg rupture, and hot leg rupture fault data. Moreover, the fault data at 80% operating condition is divided into fault data with the same and different fault degrees as those at 100% operating condition according to the fault degree.
[0008] S2. KPCA state monitoring, including:
[0009] S21. Nonlinear mapping of data, including: mapping the multi-dimensional data set X = {x1, x2, …, x nProject the original data into a high-dimensional feature space H through the non-linear mapping function φ(x) shown in formula (1), and establish a state monitoring model in this feature space;
[0010] φ:R n →H,x i →φ(x i ) (1)
[0011] S22. Kernel matrix calculation, including: calculating the inner product in the high-dimensional space using the kernel function to construct the kernel matrix K, as shown in formula (2):
[0012] K=[k(x i ,x j )] N×N (2)
[0013] The Gaussian kernel function is as shown in formula (3), and the polynomial kernel function is as shown in formula (4):
[0014]
[0015] k(x i ,x j )=(<x i ,x j >+c) d (4)
[0016] Centrally process the kernel matrix as shown in formula (5):
[0017] K′=K-1 N K-K1 N +1 N K1 N (5)
[0018] Among them, the symbol 1 N represents an N×N all-ones matrix;
[0019] S23. Eigenvalue decomposition and dimensionality reduction projection, including: eigenvalue decomposition of the centered kernel matrix as shown in formula (6), and projected data as shown in formula (7):
[0020] K′v=λv (6)
[0021]
[0022] In formula (7), z i is the value projected onto the i-th principal component, and α j is the eigenvector normalization coefficient;
[0023] S24. Abnormal state discrimination: Through the T shown in formula (8) 2Judge abnormality by the SPE statistic shown in the statistic formula (9):
[0024]
[0025] SPE = ||x - x proj || 2 (9)
[0026] In formulas (8) and (9), m represents the number of retained principal components, x represents the original data point, and x proj represents the reconstructed value of the data in the principal component space;
[0027] S3. Fault diagnosis based on CNN-Bi-LSTM deep transfer learning:
[0028] If the fault data of the target working condition is greater than the first threshold, fault diagnosis is performed through the CNN-Bi-LSTM model; if the fault data of the target working condition is less than the second threshold or in the case of data loss, the model is optimized through transfer learning, specifically including:
[0029] S31. CNN feature extraction, including: extracting local spatial features through the convolutional layer shown in formula (10), the pooling layer shown in formula (12), and the fully connected layer shown in formula (13):
[0030]
[0031] P l+1 = max(y l+1 [(x:y);(L:W)]) (12)
[0032] P(x) = f(ωx + b i ) (13)
[0033] In formula (10), the symbol l represents the l-th layer of the network, the symbol represents the output of the j-th neuron in the l-th layer, the symbol M j represents the input feature map, the symbol represents the input of the i-th neuron in the (l - 1)-th layer, the symbol ω represents the weight matrix, the symbol represents the network bias of the i-th layer and j-th neuron, the symbol σ represents the activation function; in formula (12), the symbol P l+1 represents the maximum value obtained within the pooling window range, the symbol y l+1 represents the feature map output by the activation layer, the symbols x and y represent the starting coordinates of the pooling window, and the symbols L and W represent the length and width of the pooling window; in formula (13), the symbol P(x) represents the output of the fully connected layer, the symbol ω represents the weight matrix, x is the unfolded one-dimensional feature vector, b i is the bias, and f(·) is the activation function;
[0034] S32. Bi-LSTM time series modeling, including: capturing time series dependencies through bidirectional LSTM. The forward LSTM calculates the hidden layer information in the forward order of the time series, and the backward LSTM calculates the hidden layer information in the reverse order of the time series. The hidden layer information of both is fused to achieve bidirectional time series modeling, which specifically includes the following steps:
[0035] S321. Forward LSTM calculation, including: passing forward in the time series and performing forget gate calculation through formula (14) as follows:
[0036] f t = σ(W f ·[h t-1 , x t + b f ) (14)
[0037] Perform input gate calculation through formulas (15), (16), and (17) as follows:
[0038] i t = σ(W i ·[h t-1 , x t + b i ) (15)
[0039]
[0040] Perform output gate calculation through formulas (18) and (19) to update the forward hidden state as follows:
[0041] o t = σ(W o ·[h t-1 , x t + b o ) (18)
[0042] h i = o t ⊙ tanh(c t ) (19)
[0043] Among them, in formula (14), the forget gate output f t ∈ [0, 1], representing the proportion of information to be retained; σ is the sigmoid activation function; h t-1 is the hidden state at the previous moment; x t is the input unit at the current moment; W f is the weight matrix of the forget gate controller; b f is the bias of the forget gate controller; in formulas (15), (16), and (17), is the temporary memory unit; Wi is the weight matrix of the input gate controller; b i is the bias of the input gate controller; W c is the weight matrix for updating the cell state; b c is the bias for updating the cell state; In equations (18) and (19), W o is the weight matrix of the output gate controller, b o is the bias of the output gate controller, h t is the hidden state at the current time;
[0044] S322. Backward LSTM calculation, including: backward propagation in time series, updating the backward hidden state through formula (20)
[0045]
[0046] S323. Concatenate the forward hidden state and the backward hidden state, and generate the final output o through a fully connected layer t , as shown in formulas (21) and (22):
[0047]
[0048] In equations (20), (21), and (22), W f and W b represent the weight matrices for forward propagation and backward propagation respectively; and represent the hidden layer memory unit information output by the forward propagation and backward propagation of the model respectively; o t represents the output information of the Bi-LSTM model; W o is the weight matrix of the output layer of the model; b o is the bias unit of the output layer of the model;
[0049] S33. Transfer learning, including: training the CNN-Bi-LSTM model under the source working conditions, extracting the parameters of the CNN convolutional layer, retaining the initial weight matrix of the Bi-LSTM layer, and adapting to the target working conditions through feature transfer and model fine-tuning, so as to realize the knowledge transfer of the model between different working conditions.
[0050] In some alternative embodiments, the activation function of the CNN in step S3 adopts the ReLU function, as shown in formula (11):
[0051]
[0052] Among them, the symbol x l+1 represents the upper-layer input feature, the symbol y l+1 represents the output feature after activation, and the symbol f(·) represents the ReLU activation operation.
[0053] In some alternative embodiments, in step S3, the source operating condition data is preprocessed by normalization, and among the parameters of the trained CNN-Bi-LSTM model, the weights of the CNN convolutional layer are frozen, and the fully connected layer is fine-tuned to adapt to the target operating condition.
[0054] In some alternative embodiments, the anomaly threshold for KPCA state monitoring in step S2 is determined by the 95% confidence intervals of the T 2 statistic and the SPE statistic of the normal operating condition data.
[0055] In some alternative embodiments, the number of hidden layer neurons in the Bi-LSTM layer is optimized by the accuracy of the validation set, and the optimal value is determined by the trial-and-error method or Bayesian optimization.
[0056] The beneficial effects of the embodiments of this specification are as follows: The technical solution of this application uses CNN-Bi-LSTM for fault diagnosis and combines a transfer learning strategy to enhance the adaptability of the model under different operating conditions. CNN can automatically extract deep features and improve the recognition ability of fault patterns, while the Bidirectional Long Short-Term Memory (Bi-LSTM) can effectively capture time series dependence information and enhance the dynamic perception ability of the equipment operating state. In addition, by combining the transfer learning method, the existing source operating condition data can be used to optimize the fault recognition ability of the model under the target operating condition, thereby improving the diagnostic accuracy under the conditions of small samples and unbalanced data distribution. This method can better adapt to the complex operating environment of nuclear power plants and provide a more efficient and robust solution for the intelligent fault diagnosis of static equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of this specification or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is a flowchart of a method for diagnosing faults in nuclear power plant equipment provided by the embodiments of this specification;
[0059] Figure 2 It is a mapping relationship diagram of the original space mapped to the high-dimensional feature space;
[0060] Figure 3 It is a schematic diagram of the CNN structure;
[0061] Figure 4 It is the structural diagram of the LSTM network;
[0062] Figure 5 It is the network structural diagram of Bi-LSTM;
[0063] Figure 6 It is the structural diagram of the CNN-Bi-LSTM model. Specific implementation manners
[0064] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] It should be noted that the terms "include" and "have" in the embodiments of this specification and their any deformations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0066] Figure 1 It is a schematic flowchart of a method for diagnosing equipment faults in a nuclear power plant provided by an embodiment of this specification. This process may include the following steps.
[0067] Step 1: First, perform data collection work. The collected data and characteristic parameters are shown in Table 1 and Table 2.
[0068] Table 1 Data collection
[0069]
[0070] Table 2 Names of characteristic parameters
[0071]
[0072]
[0073]
[0074] The operating conditions of a nuclear power plant can be set to two representative states, namely 100% and 80%. These two operating conditions are chosen because they can cover the operating conditions at different power levels of the nuclear power plant, thus providing multi-dimensional data support for subsequent condition monitoring and fault diagnosis. The data acquisition work can be carried out in two steps. First, data is acquired for the normal states of the two operating conditions. These data serve as the basis for subsequent analysis and are used to construct a normal operating condition model. The normal state data reflects the operating characteristics of the nuclear power plant equipment under stable and fault-free conditions and is a reference for judging whether the equipment is abnormal. Then, fault data is acquired. The fault data acquired under the 100% operating condition can include four types of fault data: steam generator heat transfer tube rupture fault data, steam pipeline rupture fault data, cold leg rupture fault data, and hot leg rupture fault data. These fault types may cause serious consequences under high-power operating conditions, and the acquisition of their data helps to deeply understand the fault modes of the equipment during high-load operation. The fault data acquisition under the 80% operating condition is more detailed and can be divided into two categories according to the degree of fault, namely the acquisition of fault data with the same degree of fault as the 100% operating condition and the acquisition of fault data with a different degree of fault from the 100% operating condition. Among them, the fault data with the same degree of fault as the 100% operating condition can include steam generator rupture fault data and cold leg rupture fault data, which can be used to compare the performance differences of the same fault under different operating conditions. The fault data packet with a different degree of fault from the 100% operating condition can include steam generator heat transfer tube rupture fault data, steam pipeline rupture fault data, cold leg rupture fault data, and hot leg rupture fault data, which can further enrich the data samples, reflect the diversity of faults under low-power operating conditions, provide a comprehensive data basis for studying the occurrence and development laws of faults under different operating conditions, and thus improve the accuracy and adaptability of the fault diagnosis model
[0075] Step 2: Develop the KPCA condition monitoring module. The principle of KPCA for condition monitoring is as follows:
[0076] (1) Nonlinear mapping of data
[0077] Assume that the multi-dimensional data set collected by the sensors of the nuclear power plant is X = {x1, x2, …, x n},where x i ∈R n represents the state characteristics of the equipment at time i. KPCA uses the nonlinear mapping function φ(x) to project the original data into a high-dimensional feature space and establish a condition monitoring model in this feature space, as shown in Formulas (1) and Figure 2 .
[0078] φ:R n →H, x i →φ(x i ) (1)
[0079] (2) Kernel matrix calculation
[0080] In the high-dimensional space, the inner product between data points is calculated through the kernel function k(x i , x j ), avoiding the computational burden of directly performing high-dimensional mapping. The calculation of the kernel matrix is shown in Equation (2). Commonly used kernel functions include the Gaussian kernel function and the polynomial kernel function, and the calculations of these two kernel functions are shown in Equations (3) and (4) respectively.
[0081] K = [k(x i , x j )] N×N (2)
[0082]
[0083] k(x i , x j ) = (<x i , x j > + c) d (4)
[0084] In Equations (3) and (4), σ represents the kernel bandwidth parameter, c represents the bias parameter, and d represents the polynomial order.
[0085] Since there may be offsets in the data, it is necessary to centralize the kernel matrix, and the processing process is shown in Equation (5).
[0086] K′ = K - 1 N K - K1 N + 1 N K1 N (5)
[0087] In Equation (5), K is the original kernel matrix, K′ is the centralized kernel matrix, and 1 N is an N×N matrix of all 1s used to remove the offset.
[0088] (3) Eigenvalue decomposition and dimensionality reduction projection
[0089] In the high-dimensional space, perform eigenvalue decomposition on the centralized kernel matrix K′, as shown in Equation (6).
[0090] K′v = λv (6)
[0091] In Equation (6), λ is the eigenvalue and v is the eigenvector.
[0092] The eigenvector v is used to construct the KPCA model and project the input data into the feature space, and this process is shown in Equation (7).
[0093]
[0094] In formula (7), z i is the value after projection onto the \(i\)-th principal component, and \(\alpha\) j is the eigenvector normalization coefficient.
[0095] (4) Condition monitoring
[0096] The system is condition-monitored using the features extracted by KPCA. The monitored indicators include: \(T\) 2 statistic and SPE statistic. The calculation processes are shown in formulas (8) and (9).
[0097]
[0098] SPE = \(\|x - \hat{x}\|\) proj \| 2 (9)
[0099] In formulas (8) and (9), \(m\) is the number of principal components retained, \(x\) is the original data point, and \(\hat{x}\) proj is the reconstructed value of the data in the principal component space.
[0100] Step 3: Development of a deep transfer learning fault diagnosis algorithm for static equipment in nuclear power plants based on the CNN-Bi-LSTM algorithm.
[0101] In this step, if the fault data of the target working condition is greater than the first threshold, fault diagnosis is performed using the CNN-Bi-LSTM model; if the fault data of the target working condition is less than the second threshold or in the case of missing data, the model is optimized through transfer learning, which is elaborated in detail below.
[0102] The first threshold and the second threshold in the above content are parameters used to determine whether the amount of fault data under the target working condition is sufficient, and thus to decide the fault diagnosis strategy. Specifically, if the fault data of the target working condition is greater than the first threshold, it indicates that the amount of fault data under this working condition is relatively sufficient. At this time, the CNN-Bi-LSTM model can be directly used for fault diagnosis because sufficient data can provide rich information for model training, enabling the model to fully learn fault features and exert its fault diagnosis ability. Under 100% working condition, the fault data such as the rupture of the steam generator heat transfer tube and the rupture of the steam pipeline are collected relatively comprehensively, and the data volume reaches or exceeds the first threshold, so the model can be directly used for diagnosis. The CNN is used to automatically extract deep local spatial features, and the Bi-LSTM is used to capture temporal dependence information to accurately identify the fault mode. The first threshold is a standard for measuring whether the fault data of the target working condition can meet the requirements of directly performing model diagnosis, and its setting depends on various factors such as model complexity, data feature dimension, and fault type diversity. Reasonably setting the first threshold can ensure that the model performance is fully exerted when the data is sufficient, improving the diagnosis efficiency and accuracy. When the fault data of the target working condition is less than the second threshold or there is a data missing working condition, it means that the data volume is insufficient. At this time, directly using the model is likely to result in overfitting of the CNN-Bi-LSTM or inaccurate diagnosis, so the model needs to be optimized through transfer learning. Under 80% working condition, some fault data are small samples or there are missing data. For example, the amount of fault data for the rupture of the steam generator heat transfer tube is small and less than the second threshold, so transfer learning is required. The second threshold is used to judge whether the amount of fault data of the target working condition is too small. If it is less than it, transfer learning needs to be used to transfer the model parameters trained in the source working condition (with sufficient data) to the target working condition and fine-tune them, and use the fault features and model learning experience in the source working condition data to optimize the fault recognition ability under the target working condition. The determination of the second threshold is also affected by various factors such as the degree of data distribution imbalance and the similarity of data between different working conditions. A suitable second threshold can timely initiate the transfer learning strategy when the data is insufficient, improving the diagnosis accuracy of the model in small sample and data imbalance scenarios.
[0103] The data generated during the operation of a nuclear power plant integrates information in both spatial and temporal dimensions. For example, the data collected by sensors at different locations at the same moment has spatial correlation, while the data collected by the same sensor at different times has temporal continuity. Traditional fault diagnosis methods cannot effectively handle such complex characteristics, resulting in limited diagnostic accuracy. To solve this technical problem, the technical solution of this application adopts a method combining CNN and Bi-LSTM to achieve efficient fault diagnosis and transfer learning. Among them, the convolutional layer of CNN performs convolution operations by sliding a convolutional kernel over the input data, performing weighted summation on the pixels or data points in the local area, and can automatically learn the local features of the data. For example, when processing the sensor data of a nuclear power plant, the convolutional layer can capture the characteristic patterns of specific sensor combinations or local areas, such as the combined features of temperature and pressure sensor data at certain key parts. The pooling layer then downsamples the feature map output by the convolutional layer, reducing the data dimension, reducing the computational complexity while retaining important features. Bi-LSTM (Bidirectional Long Short-Term Memory Network) is used to process the temporal dependence relationship of the data. It is extended on the basis of LSTM and processes the time series data from the forward and backward directions through two LSTM networks, the forward and backward ones respectively. In the fault diagnosis of a nuclear power plant, the operating state of the equipment changes over time, and Bi-LSTM can consider the influence of data at past and future moments on the current moment at the same time. For example, when monitoring the operating state of the heat transfer tubes of a steam generator, Bi-LSTM can use the temperature and pressure change trends at previous moments and the data fluctuations at subsequent moments to more accurately capture the abnormal change patterns before a fault occurs.
[0104] In summary, the technical solution of this application uses the local spatial features extracted by CNN to provide more representative inputs for Bi-LSTM, and then uses Bi-LSTM to perform temporal modeling on these features to further enhance the ability to identify fault patterns. In this way, it comprehensively and accurately analyzes the operating data of nuclear power plant equipment, improving the efficiency and accuracy of fault diagnosis.
[0105] Next, the principle of the convolutional neural network CNN mentioned above will be explained in detail. CNN is a deep learning model widely used in pattern recognition and feature extraction, and is good at automatically learning local spatial features from data. Compared with traditional feature extraction methods, CNN has the ability of end-to-end learning, can extract discriminative fault patterns from raw sensor data, thereby improving the accuracy and generalization ability of fault diagnosis. The structure of CNN is as Figure 3 shown.
[0106] As Figure 3 known, CNN is mainly composed of a convolutional layer, a pooling layer and a fully connected layer, which will be explained separately below.
[0107] (1) Convolutional layer
[0108] The convolutional layer extracts features from the input data through convolutional operations, thereby learning the local features and high-level representations of the data. The core idea of the convolutional operation is that the convolutional kernel performs a weighted sum operation on the local area of the input data. The stacking of convolutional layers and the use of multi-channel convolutional layers enable the CNN to learn more abstract and complex features. The mathematical model of the convolutional layer is shown in Equation (10).
[0109]
[0110] In Equation (10), l represents the l-th layer of the network, is the output of the j-th neuron in the l-th layer, M j is the input feature map, is the input of the i-th neuron in the (l - 1)-th layer, ω is the weight matrix, is the network bias of the j-th neuron in the l-th layer, and σ is the activation function.
[0111] The important role of the activation function is to enhance the non-linear expression ability of the network. Compared with the Sigmoid and tanh functions, ReLU has a faster calculation speed and can accelerate the forward propagation process of the neural network. Its mathematical model is shown in Equation (11).
[0112]
[0113] In Equation (11), x l+1 is the input feature from the upper layer, y l+1 is the output feature after activation, and f(·) is the ReLU activation operation.
[0114] (2) Pooling layer
[0115] The pooling layer is usually used alternately with the convolutional layer to form the basic structure of the deep CNN. The pooling operation reduces the size of the feature map and the computational complexity by aggregating the local area. Max pooling is the most commonly used pooling method in the CNN network model. It realizes the downsampling of features by selecting the maximum value in each pooling window as the output. The expression of its model is shown in Equation (12).
[0116] P l+1 =max(y l+1 [(x:y);(L:W)]) (12)
[0117] In Equation (12), P l+1 is the maximum value obtained within the pooling window, y l+1 is the feature map output by the activation layer, x and y are the starting coordinates of the pooling window, and L and W are the length and width of the pooling window.
[0118] (3) Fully Connected Layer
[0119] The fully connected layer is responsible for linearly combining the features passed from the previous layer and introducing non-linearity through the activation function to generate the final output. The output features of the convolution are mapped into a one-dimensional array through the Flatten function, and finally a one-dimensional vector for multi-classification is output. The mathematical model of the fully connected layer is shown in Equation (13).
[0120] P(x) = f(ωx + b i ) (13)
[0121] In Equation (13), P(x) is the output of the fully connected layer, ω is the weight matrix, x is the expanded one-dimensional feature vector, b i is the bias, and f(·) is the activation function
[0122] Next, the principle of LSTM will be explained in detail. The Long Short-Term Memory neural network (LSTM) is a special type of Recurrent Neural Network (RNN) designed specifically to handle temporal dependencies in long sequence data. LSTM effectively alleviates the problems of vanishing gradients and exploding gradients that occur in traditional RNNs during long sequence training by introducing forget gates, input gates, and output gates. In fault diagnosis, LSTM can capture key fault features from the temporal evolution of sensor data, enhance the model's learning ability for time-related patterns, and improve the accuracy of fault identification under complex working conditions.
[0123] LSTM is mainly composed of three gates, namely the forget gate f t , the input gate i t , and the output gate o t . These gate units continuously update the cell state c t at time step f t using recursive equations, and its structure is as Figure 4 shown.
[0124] (1) Forget Gate
[0125] The forget gate is mainly used to control when and to what extent to forget the past state. First, the forget gate receives the hidden layer state h t-1 from the previous time step and the current input x t . Then, f t determines the degree of retention of the input information from the previous time step to reduce the amount of data during model training, and the calculation process is shown in Equation (14).
[0126] f t = σ(W f · [h t-1 [[ID=5t +b f ) (14)
[0127] In formula (14), the forget gate output f t ∈ [0, 1], representing the proportion of information to be retained; σ is the Sigmoid activation function; h t-1 is the hidden state at the previous moment; x t is the input unit at the current moment; W f is the weight matrix of the forget gate controller; b f is the bias of the forget gate controller.
[0128] (2) Input gate
[0129] The main function of the input gate i t is to determine when to update the cell state and the degree of input information update. The calculation process of this module consists of two parts. First, the input gate i t decides which part of the information needs to be updated. Then, the hidden layer state and the current input are processed by the Tanh function to output a value ranging from -1 to 1. At the same time, the output value is multiplied by the Sigmoid function to obtain the latest cell state c t . The calculation process of the input gate is shown in formulas (15), (16), and (17)
[0130] i t = σ(W i · [h t-1 , x t +b i ) (15)
[0131]
[0132] In formulas (15), (16), and (17), is the temporary memory unit; W i is the weight matrix of the input gate controller; b i is the bias of the input gate controller; W c is the weight matrix for updating the cell state; b c is the bias for updating the cell state.
[0133] (3) Output gate
[0134] The output gate determines the output of the LSTM network at the current time step. The output gate generates an output vector between 0 and 1 through a Sigmoid activation function based on the state of the memory unit at the current time step and the input at the current time step, indicating which information of the memory units should be output to the output at the current time step. This output vector will be multiplied by the state of the memory unit to generate the final output.
[0135] o t = σ(W o ·[h t-1 ,x t +b o ) (18)
[0136] h t = o t ⊙tanh(c t ) (19)
[0137] In equations (18) and (19), W o is the weight matrix of the output gate controller, b o is the bias of the output gate controller, and h t is the hidden state at the current time step.
[0138] The principle of Bi-LSTM is explained below. Bi-LSTM is extended based on LSTM by introducing two independent network layers, a forward layer and a backward layer, into the traditional LSTM structure to achieve bidirectional modeling of time series data. Bi-LSTM not only considers the influence of the current time step on future states but also takes into account the impact of future time steps on the current state, thereby capturing more complete temporal features. In fault diagnosis, Bi-LSTM can more comprehensively analyze the temporal dependencies of sensor data, improving the ability to identify complex fault patterns, especially suitable for multivariate and dynamically changing industrial systems. The principle of Bi-LSTM is as Figure 5 shown.
[0139] Figure 5 In f , LSTM b represents the forward direction, and LSTM
[0140]
[0141] In equations (20), (21), and (22), W f and W b represent the weight matrices for forward propagation and backward propagation, respectively; and represent the hidden layer memory unit information output by the forward propagation and backward propagation of the model, respectively; o t represents the output information of the Bi-LSTM model; W ois the weight matrix of the model output layer; b o is the bias unit of the model output layer.
[0142] The principle of CNN-Bi-LSTM will be explained below. CNN-Bi-LSTM is a deep learning model that combines CNN and Bi-LSTM, with both the powerful feature extraction ability of CNN and the time series dependence modeling ability of Bi-LSTM. The structure of the CNN-Bi-LSTM model is as Figure 6 shown.
[0143] In the fault diagnosis of nuclear power plant equipment, when facing the situation of insufficient target working condition data (such as small samples or missing fault data), a transfer learning strategy can be adopted to train the CNN-Bi-LSTM model in the source working condition and adapt it to the target working condition. The specific process can be as follows:
[0144] First, train the CNN-Bi-LSTM model under the source working condition with relatively sufficient data. In this process, the model will automatically learn various feature patterns in the source working condition data. The convolutional layer of CNN performs convolutional operations by sliding the convolutional kernel on the input data, continuously adjusting the parameters of the convolutional kernel (i.e., the parameters of the convolutional layer) to extract local spatial features in the data, such as the key patterns presented by different combinations of sensor data in the nuclear power plant in space. Bi-LSTM adjusts its weight matrix to learn the dependence relationship of the data in the time dimension, captures the change law of the equipment operation state over time. Through the training of a large amount of source working condition data, the model can master rich and accurate feature representations.
[0145] After completing the training of the source operating condition model, extract the parameters of the CNN convolutional layer and freeze these parameters. At the same time, retain the initial weight matrix of the Bi-LSTM layer to keep it trainable during the target operating condition training. These CNN convolutional layer parameters contain important local spatial feature information learned from the source operating condition data. Transferring these CNN convolutional layer features learned from the source operating condition to the target operating condition is based on the principle that although there are differences between the source operating condition and the target operating condition, there are commonalities in some basic physical laws and fault feature patterns of nuclear power plant equipment operation. For example, in the case of a steam generator heat transfer tube rupture fault under different operating conditions, the change trends of related parameters such as temperature and pressure may be similar. These local spatial features extracted by the CNN convolutional layer (such as the spatial correlation pattern of sensor data) have cross-operating condition generality (such as equipment physical structure, sensor layout related features). After freezing, it can avoid repeated learning and improve the transfer efficiency. The Bi-LSTM layer is responsible for capturing temporal dependencies (such as the time series features of fault development). The temporal patterns under different operating conditions may vary (such as the fault development rate, signal fluctuation period). It is necessary to keep its trainability to adapt to the dynamic characteristics of the target operating condition, that is, retain its dynamic learning ability for temporal features. Thus, under the target operating condition, since the Bi-LSTM layer is not frozen, it can autonomously learn and adjust its ability to capture temporal dependencies according to the time series characteristics of the target operating condition data. For the fully connected layer, because it directly affects the classification decision of the model, only the fully connected layer is fine-tuned. By training on a small amount of data in the target operating condition, let the fully connected layer learn the data distribution and fault features unique to the target operating condition. On the basis of retaining the general spatial features extracted by the CNN and the temporal features autonomously learned by the Bi-LSTM layer, optimize the adaptability of the model to the target operating condition and improve the accuracy of fault diagnosis under the target operating condition. This method not only avoids learning from scratch in the case of data scarcity but also makes full use of the data characteristics of the target operating condition, enabling the model to more accurately identify the fault patterns under the target operating condition.
[0146] In this model, the CNN is responsible for extracting local spatial features from the input data and can effectively capture the key pattern information in the sensor data. Specifically, there are numerous sensors distributed in the nuclear power plant, which collect a large amount of data in real time. This data contains rich information about the equipment operating state. The CNN slides the convolutional kernel in the convolutional layer over the input data to perform weighted sum operations on local regions, and can automatically learn and extract the key local feature patterns. For example, when monitoring the steam generator, the CNN can accurately capture the combined features of sensor data at specific positions, such as the correlation pattern of parameters such as steam pressure, temperature, and flow rate in the local region. These patterns are often closely related to the normal or faulty state of the equipment. This local feature extraction ability enables the CNN to focus on the key parts of the data, effectively filter out redundant information, and provide a concise and valuable feature representation for subsequent analysis.
[0147]
[0147] Based on the features extracted by CNN, Bi-LSTM (Bidirectional Long Short-Term Memory Network) further models the temporal relationships of these features. The operation of nuclear power plant equipment is a dynamic process, and its state changes continuously over time. There is a strong dependence between the data at previous and current times. Bi-LSTM processes information along the forward and backward directions of the time series through two independent LSTM networks. During forward processing, it can learn how the states at previous times affect the current time; while during backward processing, it can capture the potential impact of information at future times on the current state. For example, when analyzing the changing trend of coolant flow over time, Bi-LSTM can not only use the flow data at previous times to predict the normal flow range at the current time, but also more accurately determine whether the current flow is in an abnormal state based on the flow fluctuations at subsequent times, thereby capturing more comprehensive time-dependent features and having a deeper understanding of the dynamic changes in the equipment operation state.
[0148] In an alternative embodiment technical solution, in step S3, the source condition data is preprocessed by normalization. Among the parameters of the trained CNN-Bi-LSTM model, the weights of the CNN convolutional layer are frozen, and the fully connected layer is fine-tuned to adapt to the target condition.
[0149]
[0149] The source condition data can refer to the nuclear power plant equipment operation data collected under specific and known condition conditions. These data can come from various sensors, such as temperature sensors, pressure sensors, flow sensors, etc., and reflect the operation characteristics of the equipment in normal or abnormal states. Normalization is a data preprocessing technique aimed at unifying the source condition data into a specific range, usually [0,1] or [-1,1]. In nuclear power plant equipment fault diagnosis, sensor data of different types may have different dimensions and value ranges. For example, temperature data may be between dozens of degrees Celsius and hundreds of degrees Celsius, while pressure data may be between several megapascals and dozens of megapascals. If the original data is directly used for model training, features with larger numerical ranges may have a greater impact on model training, thus masking the role of other features. Through normalization processing, the influence of dimensions can be eliminated, making all features equally important in model training, which helps to improve the convergence speed and stability of the model.
[0150] In the entire model, the CNN is mainly responsible for extracting local spatial features from the input data. When training on the source operating condition data, the convolutional layer performs convolution operations by sliding the convolution kernel over the input data, learning local feature patterns related to the equipment operating state. These features can reflect some basic characteristics of the nuclear power plant equipment under the source operating conditions, such as the combined features of sensor data at specific positions, features related to the physical structure of the equipment, etc. When the model needs to be applied to the target operating condition, although there may be differences between the target operating condition and the source operating condition, some basic physical structures of the equipment and the basic change patterns of sensor data are often similar. Therefore, freezing the weights of the CNN convolutional layer can preserve these learned effective features and prevent them from being destroyed during the fine-tuning process. Moreover, the CNN convolutional layer is responsible for extracting local spatial features (such as the spatial correlation patterns of sensor data), and these features are highly generalizable under different operating conditions (such as features related to the physical structure of the equipment and the sensor layout). After freezing, it can avoid repeated learning and improve the transfer efficiency. At the same time, freezing the weights of the convolutional layer can also reduce the number of parameters to be adjusted, reduce the risk of model overfitting, improve the generalization ability of the model, and enable the model to maintain the ability to identify key local features under different operating conditions.
[0151] In the CNN-Bi-LSTM model, the fully connected layer is responsible for linearly combining the features extracted by the previous layers and introducing non-linearity through the activation function to finally generate the classification result, that is, to judge whether there is a fault in the equipment and the type of the fault. Since there may be differences in operating parameters, environmental conditions, etc. between the target operating condition and the source operating condition, the parameters of the fully connected layer obtained by directly training with the source operating condition may not accurately meet the classification requirements of the target operating condition. Therefore, it is necessary to fine-tune the fully connected layer. During the fine-tuning process, a small amount of data under the target operating condition is used to update the parameters of the fully connected layer, enabling the model to learn the data distribution and fault characteristics unique to the target operating condition. This can optimize the adaptability of the model to the target operating condition and improve the accuracy of fault diagnosis under the target operating condition on the basis of retaining the general local features extracted by the CNN and the temporal features captured by the Bi-LSTM. By fine-tuning the fully connected layer, the model can adjust the classification boundary according to the characteristics of the target operating condition and more accurately identify the fault patterns under the target operating condition.
[0152] In the above solution, the reason for not freezing the Bi-LSTM layer is that the Bi-LSTM layer is responsible for capturing temporal dependencies (such as the time series features of fault development), and the temporal patterns may vary under different operating conditions (such as the fault development rate, signal fluctuation period). It is necessary to keep its trainability to adapt to the dynamic characteristics of the target operating condition. The role of fine-tuning the fully connected layer is that as the classification decision layer, the fully connected layer directly determines the output of the fault type. By fine-tuning the parameters of the fully connected layer, on the basis of retaining the spatial features of the CNN and the temporal features of the Bi-LSTM, it can adapt the classification boundary of the target operating condition and improve the diagnostic accuracy.
[0153] In an alternative embodiment technical solution, the abnormal threshold of KPCA state monitoring in step S2 is determined by the 95% confidence intervals of the T 2 statistic and the SPE statistic of the normal operating condition data.
[0154] In an alternative embodiment technical solution, the number of hidden layer neurons in the Bi-LSTM layer can be optimized by the accuracy of the validation set, and the optimal value is determined by the trial-and-error method or Bayesian optimization.
[0155] The confidence interval is a statistical estimation method used to measure the uncertainty between a sample statistic and a population parameter. When determining the abnormal threshold of KPCA state monitoring, choosing a 95% confidence interval means that under normal operating conditions, 95% of the data points corresponding to the T 2 statistic and the SPE statistic will fall within this interval. From a probability perspective, when the monitored statistic exceeds this interval, there is a relatively high confidence level that the device is in an abnormal state.
[0156] In actual operation, the process of determining the abnormal threshold may include: First, collect a large amount of data of the nuclear power plant equipment under normal operating conditions. Based on these data, calculate the T 2 statistic and the SPE statistic. Then, determine the 95% confidence intervals of these two statistics through statistical methods. The upper and lower limits of this interval constitute the abnormal threshold of KPCA state monitoring. During the subsequent operation of the equipment, calculate the T 2 statistic and the SPE statistic of the current data in real time, and compare them with the determined abnormal threshold. If the current statistic exceeds the threshold range, the system will determine that the equipment may be abnormal and issue a warning signal in a timely manner so that the staff can take corresponding measures. This solution determines the abnormal threshold by the 95% confidence intervals of the T 2 statistic and the SPE statistic of the normal operating condition data, which is objective and scientific. It is based on the statistical characteristics of a large amount of normal operating condition data and can avoid the subjectivity and arbitrariness of artificially setting the threshold. Moreover, the 95% confidence interval takes into account the natural fluctuations of the data while ensuring a certain degree of accuracy, and can effectively reduce the false alarm rate and improve the reliability of state monitoring.
[0157] The following takes "optimizing based on the validation set accuracy" as an example to illustrate how to determine the number of hidden layer neurons in the Bi-LSTM layer. In the technical solution of this embodiment, the validation set is a part of the data divided from the training data and is used to evaluate the performance of the model on unseen data. In the nuclear power plant equipment fault diagnosis model, in the technical solution of this embodiment, the validation set accuracy is used as an indicator to optimize the number of hidden layer neurons in the Bi-LSTM layer because the accuracy directly reflects the ability of the model to correctly identify fault types. During the training process, different numbers of hidden layer neurons will cause differences in the ability of Bi-LSTM to capture the temporal characteristics of equipment operation data. When the number of neurons is too small, the model may not be able to fully learn the complex temporal dependencies in the data, resulting in a decline in the ability to identify fault patterns and thus a lower accuracy on the validation set; while when the number of neurons is too large, the model may overfit the training data, and the generalization ability to the data in the validation set becomes poor, also reducing the accuracy. By evaluating the accuracy of the model under different numbers of hidden layer neurons on the validation set, an optimal number of neurons that balances the model's ability to capture temporal characteristics and avoid overfitting can be found, thereby improving the accuracy of the nuclear power plant equipment fault diagnosis.
[0158] In summary, the technical solution of this application combines CNN and Bi-LSTM, enabling the model to have powerful capabilities in fault diagnosis. Under complex working conditions, the operation state of the equipment is complex and changeable, and it may be affected by multiple factors simultaneously. Traditional diagnostic methods often struggle to cope. However, the model of the technical solution of this application can more comprehensively depict the operation state of the equipment and accurately identify various complex fault patterns by virtue of the local spatial features extracted by CNN and the temporal features captured by Bi-LSTM. In the case of data imbalance, that is, when the number of data samples of some fault types is far less than that of other types, the model can still avoid diagnostic biases caused by data volume differences through the extraction of key features by CNN and the effective utilization of temporal features by Bi-LSTM, demonstrating strong robustness and generalization ability, and improving the accuracy and reliability of fault diagnosis.
[0159] The technical solution of this application uses CNN-Bi-LSTM for fault diagnosis and combines a transfer learning strategy to enhance the adaptability of the model under different working conditions. CNN can automatically extract deep features and improve the recognition ability of fault patterns, while the Bidirectional Long Short-Term Memory (Bi-LSTM) can effectively capture temporal dependence information and enhance the dynamic perception ability of the equipment operation state. In addition, by combining the transfer learning method, the existing source working condition data can be utilized to optimize the fault recognition ability of the model under the target working condition, thereby improving the diagnostic accuracy under the conditions of small samples and unbalanced data distribution. This method can better adapt to the complex operation environment of nuclear power plants and provide a more efficient and robust solution for the intelligent fault diagnosis of static equipment.
[0160] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A diagnostic method for equipment failures in a nuclear power plant, characterized in that, It includes the following steps: S1. Data acquisition, including: obtaining the normal state data and fault data of the static equipment in a nuclear power plant under 100% and 80% operating conditions. The fault data includes the fault data of steam generator heat transfer tube rupture, steam pipeline rupture, cold leg rupture, and hot leg rupture. And the fault data under 80% condition is divided into fault data with the same and different fault degrees as those under 100% condition; S2. KPCA state monitoring, including: S21. Data non - linear mapping, including: projecting the multi - dimensional data set X = {x1, x2, ···, x n} collected by the sensor to the high - dimensional feature space H through the non - linear mapping function φ(x) shown in formula (1), and establishing a state monitoring model in this feature space; φ:R n →H,x i →φ(x i ) (1) S22. Kernel matrix calculation, including: calculating the inner product in the high-dimensional space by using the kernel function to construct the kernel matrix K, as shown in formula (2): K = [k(x i , x j )]N ×N (2) The Gaussian kernel function is as shown in formula (3), and the polynomial kernel function is as shown in formula (4): k(x i ,x j )=(<x i ,x j >+c)d(4) Centering the kernel matrix as shown in formula (5): K′ = K - 1 N K - K1 N +1 N K1 N (5) Among them, symbol 1 N represents an N×N all-ones matrix; S23. Eigenvalue decomposition and dimensionality reduction projection, including: performing eigenvalue decomposition on the centered kernel matrix as shown in formula (6), and projecting the data as shown in formula (7): K′v=θv (6) In Equation (7), z i is the value after projection onto the i-th principal component, and α j is the eigenvector normalization coefficient; S24. Abnormal state discrimination: Determine abnormality by the T statistic shown in formula (8) 2 and the SPE statistic shown in the statistic formula (9) of formula (8): SPE = ||x - x proj || 2 (9) In formulas (8) and (9), m represents the number of retained principal components, x represents the original data point, and x proj represents the reconstructed value of the data in the principal component space; S3. Fault diagnosis based on CNN-Bi-LSTM deep transfer learning: If the fault data of the target condition is greater than the first threshold, fault diagnosis is performed through the CNN-Bi-LSTM model; if the fault data of the target condition is less than the second threshold or in the case of data missing, the model is optimized through transfer learning, specifically including: S31. CNN feature extraction, including: extracting local spatial features through the convolutional layer shown in formula (10), the pooling layer shown in formula (12), and the fully connected layer shown in formula (13): P l+1 = max(y l+1 [(x:y);(L:W)]) (12) P(x) = f(ωx + b i ) (13) In formula (10), the symbol l represents the l-th layer of the network, and the symbol represents the output of the j-th neuron in the l-th layer. The symbol M j represents the input feature map, and the symbol represents the input of the i-th neuron in the (l - 1)-th layer. The symbol ω represents the weight matrix, and the symbol represents the network bias of the j-th neuron in the i-th layer. The symbol σ represents the activation function; in formula (12), the symbol P l+1 represents the maximum value obtained within the pooling window. The symbol y l+1 represents the feature map output by the activation layer. The symbols x and y represent the starting coordinates of the pooling window, and the symbols L and W represent the length and width of the pooling window; in formula (13), the symbol P(x) represents the output of the fully connected layer, the symbol ω represents the weight matrix, x is the unfolded one-dimensional feature vector, and b i is the bias, and f(·) is the activation function; S३२. Bi-LSTM time series modeling, including: capturing time series dependencies through bidirectional LSTM. The forward LSTM calculates the hidden layer information in the forward order of the time series, and the backward LSTM calculates the hidden layer information in the reverse order of the time series. The hidden layer information of both is fused to achieve bidirectional time series modeling, specifically including the following steps: S321. Forward LSTM calculation, including: passing forward in the time series, performing forget gate calculation through formula (14), specifically as follows: f t = σ(W f · [h t-1 , x t + b f ) (14) Performing input gate calculation through formulas (15), (16), and (17), specifically as follows: i t = σ(W i · [h t-1 , x t + b i ) (15) Performing output gate calculation through formulas (18) and (19) to update the forward hidden state, specifically as follows: o t = σ(W o ·[h t-1 ,x t +b o ) (18) h t = o t ⊙tanh(c t ) (19) Among them, in formula (14), the forget gate output f t ∈[0,1], representing the proportion of information to be retained; σ is the sigmoid activation function; h t-1 is the hidden state at the previous moment; x t is the input unit at the current moment; W f is the weight matrix of the forget gate controller; b f is the bias of the forget gate controller; in formulas (15), (16) and (17), is the temporary memory unit; W i is the weight matrix of the input gate controller; b i is the bias of the input gate controller; W c is the weight matrix for updating the cell state; b c is the bias for updating the cell state; in formulas (18) and (19), W o is the weight matrix of the output gate controller, b o is the bias of the output gate controller, h t is the hidden state at the current moment; S322. Backward LSTM calculation, including: passing backward in the time series, updating the backward hidden state through formula (20) S323. Concatenate the forward hidden state and the backward hidden state, and generate the final output o through a fully connected layer t , as shown in Formulas (21) and (22): In equations (20), (21), and (22), W f and W b represent the weight matrices for forward propagation and backward propagation, respectively; and represent the hidden layer memory unit information output by the forward propagation and backward propagation of the model, respectively; o t represents the output information of the Bi-LSTM model; W o is the weight matrix of the output layer of the model; b o is the bias unit of the output layer of the model; S33. Transfer learning, including: training the CNN-Bi-LSTM model in the source condition, extracting the parameters of the CNN convolutional layer, retaining the initial weight matrix of the Bi-LSTM layer, and adapting to the target condition through feature transfer and model fine-tuning, so as to achieve knowledge transfer of the model between different conditions.
2. The method according to claim 1, wherein The activation function of CNN in step S3 adopts the ReLU function, as shown in formula (11): Among them, the symbol x l+1 represents the upper-layer input feature, the symbol y l+1 represents the output feature after activation, and the symbol f(·) represents the ReLU activation operation.
3. The diagnostic method for equipment failures in a nuclear power plant according to claim 1, characterized in that, In step S3, the source condition data is pre-normalized. Among the parameters of the trained CNN-Bi-LSTM model, the weights of the CNN convolutional layer are frozen, and the fully connected layer is fine-tuned to adapt to the target condition.
4. The diagnostic method for equipment failures in a nuclear power plant according to claim 1, characterized in that In the step S2, the abnormal threshold of KPCA state monitoring is determined by the 95% confidence intervals of the T 2 statistic and the SPE statistic of the normal operating condition data.
5. The diagnostic method for equipment failures in a nuclear power plant according to claim 1, characterized in that The number of hidden layer neurons in the Bi-LSTM layer is optimized through the accuracy of the validation set, and the optimal value is determined by the trial-and-error method or Bayesian optimization.
Citation Information
Cited By
Power equipment simulation analysis system based on machine learning
CN120579470A
Intelligent fault diagnosis method, system and equipment for chemical production equipment and medium
CN121544574A