Nuclear power plant rotating part fault diagnosis method based on MOCO twin neural network
By using a twin neural network method based on MOCO in the fault diagnosis of rotating components of nuclear power plants, the encoder update process is optimized, and the problem of low accuracy of fault diagnosis under small sample conditions is solved, and a high accuracy of fault diagnosis is achieved.
Patent Information
- Application Number
- CN202510216043.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art is difficult to effectively diagnose rotating components of nuclear power plants under small sample conditions, and there are mutations in the feature extraction of twin neural networks after the encoder is updated, affecting the diagnostic accuracy.
Using a twin neural network method based on MOCO, the encoder update is optimized through momentum transformation and dictionary query, reducing feature extraction mutations and improving the model resistance robustness.
It significantly improves the resistance of the model, can effectively diagnose faults under small sample conditions, achieve 75% accuracy, and improves the ability to identify difficult samples.
Smart Images

Figure CN120145003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis of rotating components in nuclear power plants, and particularly to a method for fault diagnosis of rotating components in nuclear power plants based on a MOCO twin neural network. Background Art
[0002] Faults are likely to occur in the rotating components of various rotating mechanical equipment in nuclear power units. Monitoring their operating states and timely identifying and diagnosing faults are of great significance for ensuring the safe and stable operation of nuclear power units.
[0003] Currently, the data-driven method that combines sensor monitoring data with machine learning algorithms is a research hotspot in the field of bearing fault diagnosis. However, it is highly dependent on signal processing technology and expert experience, and it is difficult to efficiently extract feature information from massive data.
[0004] Deep learning theory can adaptively extract deep features of input signals, effectively solving the problem of efficiently mining a large number of signals. However, these methods require a sufficient number of samples to train the model. During the operation of nuclear power plants, complex environments often result in the inability to obtain sufficient samples of various fault states. Under the condition of small samples, the accuracy of these models will be affected.
[0005] For the problem of fault diagnosis with limited data samples, a twin neural network is usually used for fault diagnosis. However, after the encoder is updated, there is a problem of sudden change in feature extraction compared with the previous batch, which will affect the diagnostic accuracy. Summary of the Invention
[0006] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a method for fault diagnosis of rotating components in nuclear power plants based on a MOCO twin neural network. By means of momentum transformation and dictionary query, the problem of sudden change in feature extraction caused by encoder update is reduced, and the anti-robustness of the model is significantly improved, effectively solving the problem of fault diagnosis under small samples.
[0007] To achieve the above purpose, the present invention provides the following solutions:
[0008] A method for fault diagnosis of rotating components in nuclear power plants based on a MOCO twin neural network, comprising:
[0009] Obtain the sensor signals of the rotating components, preprocess the sensor signals, and obtain time-frequency features containing rich spatial position information;
[0010] Input the time-frequency features into the MoCo-SINN model to obtain input features; the MoCo-SINN model is obtained by training a Siamese neural network model using a small-sample training set and updating the model parameters in combination with the FocalLoss loss function; the small-sample training set includes: original time-frequency diagrams.
[0011] Training the Siamese neural network model using a small-sample training set includes: optimizing the process of training the Siamese neural network model using the small-sample training set through the form of a dictionary query task and momentum encoding.
[0012] Perform similarity metric calculation on the input features to realize the deep feature extraction ability of the encoder, and finally obtain the fault diagnosis result of the rotating component through the decoder.
[0013] Optionally, obtaining the time-frequency features includes: performing data reconstruction, zero-mean normalization processing, and continuous wavelet transform on the sensor signal to extract the time-frequency features.
[0014] Optionally, performing zero-mean normalization processing on the sensor signal includes:
[0015]
[0016] where x is the original data, and x * is the normalized data, and μ and σ are the mean and variance of x, respectively.
[0017] Optionally, optimizing the process of training the Siamese neural network model using a small-sample training set through the form of a dictionary query task includes:
[0018] Regarding contrastive learning as a dictionary query task, creating a dictionary queue, which is used to store key samples to enable query samples to obtain more contrasts.
[0019] Randomly select samples from the small-sample training set as query samples. The query samples are encoded using a single encoder, and the query sample data reconstruction is used as key samples and loaded into the dictionary queue for storage; among them, the data reconstruction randomly selects the cropping size of the query samples between 0.2 and 1.0 as key samples.
[0020] Encode the key samples using another encoder. There is a 50% probability that the key samples are randomly applied with Gaussian blur, with the blur radius between 0.1 and 2.0, an 80% probability of randomly applying color jitter, with the range of color jitter being ±0.4 for brightness, contrast, and saturation, ±0.1 for hue, and a 20% probability that the samples are randomly converted to grayscale images.
[0021] In the process of storing the dictionary queue, the key sample input this time is used as the positive sample, and other samples in the queue are used as negative samples to train the siamese neural network model. When the number of samples stored in the queue reaches the upper limit, the earliest sample will be dequeued.
[0022] Optionally, the process of optimizing the training of the siamese neural network model with a small sample training set by momentum encoding includes:
[0023] Update the encoder by momentum encoding:
[0024] y t = m·y t-1 + (1 - m)·y t
[0025] where y t represents the encoder in the current state, m represents the momentum update coefficient, and its magnitude affects the update speed of the encoder. y t-1 represents the encoder in the previous state.
[0026] Optionally, encoding using the encoder includes: inputting the input data into a convolutional layer and a residual block for feature extraction, and inputting the feature extraction result into a fully connected layer for integration.
[0027] Optionally, the expression of the FocalLoss function is:
[0028] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0029] where p t is the probability of the model predicting the similarity between two samples, α t is the weight coefficient, and γ is the adjustment parameter.
[0030] Optionally, calculating the similarity metric for the input features includes:
[0031] The cosine similarity is expressed as:
[0032]
[0033] Then the cosine distance is:
[0034]
[0035] where D w (X 1 , X 2 ) is the cosine distance, ||X 1|| is the feature encoded for the query sample, || X 2 || is the feature encoded for the key sample, and finally the Sigmoid function is used to output the relative distance between 0 and 1. The closer to 0, the more similar, and vice versa;
[0036] The mathematical form of the Sigmoid function is:
[0037]
[0038] where x is the relative distance.
[0039] Optionally, obtaining the fault diagnosis result includes:
[0040] Using the query encoder trained by the siamese neural network to extract features from the sample, and inputting the obtained deep feature information into the decoder to complete fault diagnosis, and the decoder is implemented using the Softmax method;
[0041] The mathematical form of the Softmax function is:
[0042]
[0043] where x is the output vector of the fully connected layer, i is the subscript corresponding to the element of the output vector of the fully connected layer, and N is the number of elements of the output vector of the fully connected layer.
[0044] The beneficial effects of the present invention are:
[0045] The present invention adopts a learning strategy based on MOCO, creates a key queue, conducts sample comparison through dictionary query, and slowly updates the key encoder using momentum encoding, reducing the problem of sudden changes in the extracted features caused by encoder updates, significantly improving the anti-robustness of the model, and can effectively solve the fault diagnosis problem under small samples. In the case of 50 samples, the accuracy can also reach 75%.
[0046] Aiming at the problem of large measurement error, the present invention selects the cosine distance to measure the similarity of the features of key and query, which can better distinguish the differences between key and query and improve the accuracy of the model.
[0047] To solve the problem that it is difficult to distinguish the output of sample pairs hovering around 0.5 in the measurement method, the present invention selects the FocalLoss loss function to reduce the relative weight of the simple discrimination part, enhance the attention to difficult sample pairs, better train difficult sample pairs, accelerate the calculation process, and the accuracy is improved by about 5% compared with the traditional method. Description of the Drawings
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 Flowchart of a fault diagnosis method for rotating components in a nuclear power plant based on a MOCO Siamese neural network according to an embodiment of the present invention;
[0050] Figure 2 Structural schematic diagram of the Siamese neural network according to an embodiment of the present invention;
[0051] Figure 3 Structural diagram of the MoCo-SINN model according to an embodiment of the present invention;
[0052] Figure 4 Schematic diagram of the principle of the encoder according to an embodiment of the present invention;
[0053] Figure 5 Schematic diagram of the principle of the decoder according to an embodiment of the present invention;
[0054] Figure 6 Schematic diagram of the accuracy rate of the MoCo-SINN model according to an embodiment of the present invention;
[0055] Figure 7 t-SNE visualization image of the sample features without extraction according to an embodiment of the present invention;
[0056] Figure 8 t-SNE visualization image of the sample features extracted by the encoder according to an embodiment of the present invention. Detailed implementation manners
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0058] To make the above objects, features, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the drawings and specific implementation manners.
[0059] As Figure 1As shown in the figure, this embodiment discloses a fault diagnosis method for rotating components in a nuclear power plant based on a MOCO Siamese neural network, including: obtaining the sensor signals of the rotating components, preprocessing the sensor signals, and obtaining time-frequency features containing rich spatial position information; inputting the time-frequency features into the MoCo-SINN model to obtain input features; the MoCo-SINN model trains the Siamese neural network model using a small sample training set, and updates the model parameters by combining the FocalLoss loss function; the small sample training set includes: the original time-frequency diagram; training the Siamese neural network model using the small sample training set includes: optimizing the process of training the Siamese neural network model with the small sample training set through the form of a dictionary query task and the momentum encoding method; performing similarity measurement calculation on the input features to realize the deep feature extraction ability of the encoder, and finally obtaining the fault diagnosis result of the rotating components through the decoder.
[0060] Specifically: The present invention provides a fault diagnosis method for rotating components in a nuclear power plant based on a MOCO Siamese neural network, including the following steps: Data preprocessing: performing wavelet transform on the sensor signals of the rotating components to perform denoising processing and extract time-frequency features; Construction of a Siamese neural network model based on MOCO: encoding the training samples using the Siamese neural network method, and updating the encoder of the Siamese network using the momentum contrast method; Distance measurement: performing similarity measurement on the features after sample encoding using the cosine distance strategy to complete the fault diagnosis of the rotating components in the nuclear power plant.
[0061] Further, obtaining the time-frequency features includes: performing zero-mean normalization processing and continuous wavelet transform on the sensor signals to extract time-frequency features.
[0062] Further, performing zero-mean normalization processing on the sensor signals includes:
[0063]
[0064] where x is the original data, and x * is the normalized data, and μ and σ are the mean and variance of x, respectively.
[0065] As a preferred solution of the present invention, in the wavelet transform, the cmor wavelet is selected as the wavelet basis function for continuous wavelet transform, and the window size parameter is an array from 1, 1 / 2,..., 1 / 127. Finally, a time-frequency diagram with a size of 228×228 is obtained through continuous wavelet transformation.
[0066] Furthermore, contrastive learning is regarded as a dictionary query task, and a dictionary queue is created. The dictionary queue is used to store key samples, enabling query samples to obtain more contrasts. Samples in the small sample training set are randomly selected as query samples. The query samples are encoded using a single encoder, and the query sample data is reconstructed as key samples and loaded into the dictionary queue for storage. Among them, data reconstruction randomly selects the cropping size of the query samples between 0.2 and 1.0 as key samples. The key samples are encoded using another encoder. Among them, there is a 50% probability that the key samples are randomly applied with Gaussian blur, with the blur radius between 0.1 and 2.0, an 80% probability of randomly applying color jitter, where the range of color jitter is ±0.4 for brightness, contrast, and saturation, ±0.1 for hue, and a 20% probability that the samples randomly convert the image to grayscale. During the storage process of the dictionary queue, the current input key sample is used as the positive sample, and all other samples in the queue are used as negative samples to train the siamese neural network model. When the number of samples stored in the queue reaches the upper limit, the earliest sample will be removed from the queue.
[0067] Furthermore, the process of optimizing the training of the siamese neural network model for the small sample training set by momentum encoding includes: updating the encoder by momentum encoding:
[0068] y t = m·y t-1 +(1 - m)·y t
[0069] where y t represents the encoder in the current state, m represents the momentum update coefficient, whose magnitude affects the encoder update speed, and y t-1 represents the encoder in the previous state.
[0070] Furthermore, encoding using the encoder includes: inputting the input data into the convolutional layer and the residual block for feature extraction, and inputting the feature extraction result into the fully connected layer for integration.
[0071] Furthermore, the expression of the FocalLoss function is:
[0072] FL(p t ) = -α t (1 - p t ) γ log(p t )
[0073] where p t is the probability of the model predicting the similarity between two samples, α t is the weight coefficient, and γ is the adjustment parameter.
[0074] Further, the similarity measurement calculation of the input features includes:
[0075] The cosine similarity is expressed as:
[0076]
[0077] Then the cosine distance is:
[0078]
[0079] Where D w (X 1 , X 2 ) is the cosine distance, ||X 1 || is the feature after encoding the query sample, ||X 2 || is the feature after encoding the key sample. Finally, the Sigmoid function is used to output the relative distance between 0 and 1. The closer to 0, the more similar, and vice versa;
[0080] The mathematical form of the Sigmoid function is:
[0081]
[0082] Where x is the relative distance.
[0083] Further, obtaining the fault diagnosis result includes:
[0084] Using the query encoder trained by the Siamese neural network to extract features from the sample, and inputting the obtained deep feature information into the decoder to complete the fault diagnosis. The decoder is implemented using the Softmax method, as Figure 5 shown.
[0085] The mathematical form of the Softmax function is:
[0086]
[0087] Where x is the output vector of the fully connected layer, i is the subscript of the element corresponding to the output vector of the fully connected layer, and N is the number of elements in the output vector of the fully connected layer.
[0088] As a preferred solution of the present invention, in the Siamese neural network model based on MOCO, the Siamese neural network requires paired samples (x 1 , x 2 , y) as inputs, where the value of y is the label value. The encoder is used to extract features from the data, the extracted features are measured to obtain the similarity of the sample pair, and finally, a judgment is made according to the similarity; after each input batch is completed, the encoder is updated. The two encoders of the Siamese neural network have the same structure and equal weights.
[0089] As a preferred embodiment of the present invention, the improvement method of the twin neural network model based on MOCO is as follows: regarding contrastive learning as a dictionary query task, creating a dictionary queue, and randomly selecting a sample as a query and using a single encoder alone, and loading the sample into the queue for storage and using another encoder as the key. The positive samples in the queue are the others are all negative samples of t y = m·y t-1 +(1 - m)·y t where y t represents the encoder of the current state, m represents the momentum update coefficient, and its magnitude affects the update speed of the encoder. y t-1 represents the encoder of the previous state; ensuring that the features of all keys in the dictionary are extracted by similar encoders, and slowly updating the encoder while maintaining feature consistency as much as possible.
[0090] As a preferred embodiment of the present invention, the encoder selects the ResNet-152 model. The data first passes through a convolutional layer with a convolutional kernel size of 7x7, a stride of 2x2, and a pooling of 3x3, and then uses four residual blocks of the ResNet model to extract features from the graphic data. Among them, the first and fourth layers use 3 bottleneck modules, the second layer uses 4 bottleneck blocks, the third layer uses 23 bottleneck modules, and finally uses a fully connected layer to integrate the feature data.
[0091] As a preferred embodiment of the present invention, the loss function of the twin neural network model based on MOCO selects FocalLoss, that is, FL(p t ) = -α t (1 - p t ) γ log(p t ), enabling the model to focus on optimizing difficult-to-distinguish sample pairs with a difference hovering around 0.5.
[0092] As a preferred embodiment of the present invention, in the distance metric method, the cosine similarity can be expressed as: Then the cosine distance is expressed as: The smaller the angle between the two vectors (the closer to 0), the smaller the cosine distance, and the closer their feature distances are, that is, the more similar they are, and vice versa.
[0093] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those specifically described herein. Those skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0094] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other from other embodiments.
[0095] The present invention adopts a fault diagnosis method for rotating components in nuclear power plants based on a MOCO Siamese neural network, and its process is as Figure 1 shown. The diagnosis method consists of three parts: data preprocessing, a Siamese neural network model based on MOCO (MoCo-SINN), and distance measurement.
[0096] Embodiment 1
[0097] Figure 2 is a schematic diagram of the structure of the Siamese neural network. To verify the effectiveness of the MoCo-SINN model in fault diagnosis, the rotating component dataset of the Machinery Failure Prevention Technology Society (MFPT) is selected for experiments. The test bearing model of the MFPT rotating component dataset is NICE. Single-point faults are set on the inner and outer rings of the rotating component, and data measurements are carried out under different loads. Seven state datasets are selected, as shown in Table 1. Table 1 is a detailed description of the MFPT rotating components.
[0098] Table 1
[0099]
[0100] Embodiment 2
[0101] As a preferred solution of the present invention, data preprocessing includes data reconstruction, data normalization, and wavelet transform.
[0102] As a preferred solution of the present invention, data reconstruction in data preprocessing: each sample contains a one-dimensional time series signal with a length of 1024. Each sample is converted into a two-dimensional feature map through wavelet transform, and its length and width are both 64 as samples. Then, the cropping size of the samples is randomly selected between 0.2 and 1.0 as samples, where There is a 50% probability that the sample randomly applies Gaussian blur with a blur radius between 0.1 and 2.0, an 80% probability that it randomly applies color jitter with a range of ±0.4 for brightness, contrast, and saturation, and ±0.1 for hue, and a 20% probability that the sample randomly converts the image to grayscale.
[0103] As a preferred embodiment of the present invention, data normalization in data preprocessing: To eliminate the influence of gradient disappearance or gradient explosion during fault diagnosis, zero-mean normalization processing is performed on the signal data obtained from rotating component equipment, that is:
[0104]
[0105] In the above formula, x is the original data; x * is the normalized data, and μ and σ are the mean and variance of x respectively. 70% of the samples are selected as the training set, and the rest are used as the test data set.
[0106] As a preferred embodiment of the present invention, wavelet transform in data preprocessing: Among the same type of faults, different intercepted signal periods may lead to differences in the spatial positions of faults. Therefore, continuous wavelet transform (CWT) is introduced to convert the sample data into a time-frequency diagram containing rich spatial position information, thereby retaining most of the fault feature information. At the same time, continuous wavelet transform can also effectively remove the noise during data acquisition. By convolving the original data with the wavelet basis function, the wavelet transform coefficients of the signal are obtained to measure the matching degree between the signal and the scale position of the wavelet basis function. Convolution is performed at different window scales to obtain the local feature information of the signal in different frequency ranges, thereby realizing the extraction of the time-frequency characteristics of the signal.
[0107] As a preferred embodiment of the present invention, the parameter selection in the wavelet transform in data preprocessing is as follows:
[0108] The cmor wavelet is selected as the wavelet basis function for continuous wavelet transform; the window size parameter is an array from 1, 1 / 2,..., 1 / 127, and finally a time-frequency diagram with a size of 228×228 is obtained through continuous wavelet transform.
[0109] As a preferred embodiment of the present invention, in the Siamese neural network model based on MOCO, to solve the problem that there is a jump in the samples before and after the update of the encoder during feature extraction, which will affect the diagnostic accuracy, and to solve the problem that the encoder parameters change after each backpropagation, and there is a mutation in the sample features extracted by the encoder before and after. The present invention selects the MOCO method to improve the Siamese neural network, such as Figure 3As shown in the figure. The MOCO method regards contrastive learning as a dictionary query task, creates a dictionary queue, randomly selects a sample as the query, and uses a single encoder. The sample is loaded into the queue for storage and used as the key, and another encoder is used. The positive samples in the queue are the others are all negative samples, thereby increasing the number of training times for small samples. Each time positive and negative samples are judged, it can be regarded as finding the key identical to the query in the dictionary queue. If the key of the current batch is obtained from the current encoder and the subsequent key is obtained from the updated encoder, it is impossible to ensure the consistency of the keys in the dictionary. Therefore, the momentum encoding method of Equation 2 is used.
[0110] y t = m·y t-1 +(1 - m)·y t
[0111] In the formula, y t represents the encoder in the current state, m represents the momentum update coefficient, and its magnitude affects the update speed of the encoder. y t-1 represents the encoder in the previous state. Ensure that the features of all keys in the dictionary are extracted by similar encoders, and the encoder is updated slowly while maintaining feature consistency as much as possible.
[0112] As a preferred solution of the present invention, in the Siamese neural network model based on MOCO, the ResNet-152 model is selected as the encoder to prevent the problem of gradient disappearance caused by the network being too deep. The data first passes through a convolutional layer with a convolutional kernel size of 7x7, a stride of 2x2, and a padding of 3x3, and then uses four residual blocks of the ResNet model to extract features from the graphic data. Among them, the first and fourth layers use 3 bottleneck modules, the second layer uses 4 bottleneck blocks, the third layer uses 23 bottleneck modules, and finally a fully connected layer is used to integrate the feature data, as Figure 4 shown.
[0113] As a preferred solution of the present invention, the loss function of the model is selected as follows: The output form of the SINN model is binary classification, and the training goal is that the difference degree between similar samples is as close to 0 as possible, and the difference degree between different samples is as close to 1 as possible. Therefore, the FocalLoss is selected as the loss function of the model, so that the model can focus on optimizing the difficult-to-distinguish sample pairs with a difference hovering around 0.5, that is:
[0114] FL(p t ) = -α t (1 - p t ) γlog(p t )
[0115] where p t represents the probability of the model predicting the similarity between two samples, α t is the weight coefficient used to balance the contributions between different categories, γ is a tuning parameter used to reduce the relative weight of easy samples and enhance the attention to hard samples, which accelerates the calculation process and increases the model accuracy. After determining the loss function, the encoder of MoCo-SINN is trained using the gradient descent algorithm to achieve backpropagation of errors.
[0116] As a preferred embodiment of the present invention, in the distance metric method, the MoCo-SINN model determines whether samples are from the same category through the distance metric method. Common methods include Euclidean distance, cosine similarity, and Manhattan distance, etc. The present invention selects cosine similarity to judge the differences between samples because cosine similarity measures the differences between samples by calculating the cosine value of the included angle between two vectors, paying more attention to the differences in the directions of the two vectors rather than the distance or length, and can better distinguish the differences in output features. Therefore, using cosine similarity to judge the differences between samples can perform better similarity measurement on output features and improve the classification accuracy. Cosine similarity can be expressed as:
[0117]
[0118] Then the cosine distance is expressed as:
[0119] D w (X 1 ,X 2 ) = 1 - cos(X 1 ,X 2 )
[0120]
[0121] Therefore, the smaller the included angle between two vectors (the closer to 0), the smaller the cosine distance, the closer their feature distances are, that is, the more similar they are, and vice versa.
[0122] Example 3
[0123] First, set the number of iterations and batch size to 100 and 32 respectively. The optimizer selects SGD (stochastic gradient descent), and the learning rate is dynamically adjusted during the training process. The MoCo-SINN model is pre-trained through the training dataset. Then, the deep features of the training dataset samples are extracted through the pre-trained query encoder, and the deep features are then trained through the fully connected layer. After training, the performance of the MoCo-SINN model is evaluated using the test set data. To reduce the random effect, the model is tested 10 times, and the results on the test set are asFigure 6 As shown. The highest accuracy rate in multiple tests is 98.45%, and the average accuracy rate is 98.21%, indicating that the MoCo-SINN model has stable diagnostic performance and can accurately identify the state types of rotating components.
[0124] To more vividly demonstrate the feature extraction ability of the MoCo-SINN model, the t-distributed stochastic neighbor embedding (t-SNE) method is used for dimensionality reduction to visualize the sample features. Among them Figure 7 is the t-SNE image of the features extracted without the query encoding layer. It is impossible to distinguish between different categories and the distribution of samples in the same category is sparse. Figure 8 is the image after the features are extracted by the query encoding layer. It can be seen that the relative distances between different categories increase, the classification effect is obvious, and the distribution of samples in the same category is more concentrated. This further shows that the constructed model has good feature extraction ability.
[0125] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for fault diagnosis of rotating components of nuclear power plants based on MOCO twin neural network, characterized in that: include: Acquire sensor signals of the rotating component, preprocess the sensor signals, and acquire time-frequency features containing rich spatial position information; The time-frequency features are input into the MoCo-SINN model to obtain input features; the MoCo-SINN model uses a small sample training set to train a twin neural network model, and updates the model parameters in combination with the FocalLoss loss function; The small sample training set includes: original time-frequency graph; Using a small sample training set to train a twin neural network model includes: optimizing the process of training a twin neural network model using a small sample training set in the form of a dictionary query task and a momentum encoding method; The input features are subjected to similarity measurement calculation to realize the deep feature extraction capability of the encoder, and finally the fault diagnosis result of the rotating component is obtained through the decoder.
2. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 1 is characterized in that: Acquiring the time-frequency features includes: performing data reconstruction, zero-mean normalization processing and continuous wavelet transform on the sensor signal to extract the time-frequency features.
3. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 2 is characterized in that: The zero-mean normalization processing of the sensor signal includes: Among them, x is the original data, x * is the normalized data, μ and σ are the mean and variance of x respectively.
4. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 1 is characterized in that: The process of training the twin neural network model by optimizing the small sample training set in the form of a dictionary query task includes: Take contrastive learning as a dictionary query task and create a dictionary queue, which is used to store key samples so that query samples can obtain more comparisons. Randomly select a sample from the small sample training set as a query sample, encode the query sample using an encoder alone, and reconstruct the query sample data as a key sample and load it into the dictionary queue for storage; wherein the data reconstruction is to randomly select a crop size of the query sample between a ratio of 0.2 and 1.0 as the key sample; The key sample is encoded using another encoder, where the key sample has a 50% probability of randomly applying Gaussian blur with a blur radius between 0.1 and 2.0, an 80% probability of randomly applying color jitter with a range of ±0.4 for brightness, contrast, and saturation, and ±0.1 for hue, and a 20% probability sample randomly converts the image into a grayscale image; During the storage process of the dictionary queue, the key sample input this time is used as a positive sample, and other samples in the queue are used as negative samples to train the twin neural network model. When the queue storage samples reaches the upper limit, the earliest sample will be removed from the queue.
5. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 4 is characterized in that: The process of training the twin neural network model by optimizing the small sample training set through momentum encoding includes: Update the encoder using momentum encoding: and t =m·y t-1 +(1-m)·y t Among them, y t Indicates the encoder of the current state, m represents the momentum update coefficient, whose size affects the encoder update speed, y t-1 The encoder representing the previous state.
6. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 4 is characterized in that: Encoding using the encoder includes: inputting input data into a convolutional layer and a residual block for feature extraction, and inputting the feature extraction result into a fully connected layer for integration.
7. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 1 is characterized in that: The FocalLoss loss function expression is: FL(p t )=-a t (1-p t ) γ log(p t ) Among them, p t The model predicts the probability of similarity between two samples, α t is the weight coefficient, and γ is the adjustment parameter.
8. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 1 is characterized in that: Calculating the similarity measure of the input features includes: Cosine similarity is expressed as: Then the cosine distance is: Among them, D w (X1,X2) is the cosine distance, ||X1|| is the feature after query sample encoding, ||X2|| is the feature after key sample encoding, and finally the Sigmoid function is used to output the relative distance between 0 and 1. The closer to 0, the more similar, and vice versa; The mathematical form of the Sigmoid function is: Wherein, x is the relative distance.
9. The method for fault diagnosis of rotating parts of nuclear power plants based on MOCO twin neural network according to claim 1 is characterized in that: Obtaining the fault diagnosis result includes: The query encoder trained by the twin neural network is used to extract features from the sample, and the acquired deep feature information is input into the decoder to complete fault diagnosis. The decoder is implemented using the Softmax method; The mathematical form of the Softmax function is: Among them, x is the output vector of the fully connected layer, i is the subscript number corresponding to the element of the output vector of the fully connected layer, and N is the number of elements of the output vector of the fully connected layer.
Citation Information
Patent Citations
Bearing fault diagnosis method and system based on small samples
CN115081490A
MoCo convolutional neural network model detection device for classification and classification method
CN116206147A
Model training method and device used in industrial scene
CN117934919A
Semi-supervised medical image segmentation method and system based on image dense block contrast learning
CN118115514A
Apparatus for providing service that allow users to experience various affiliated facilities
KR1020250066185A