A small sample fault detection method based on double-branch contrastive learning

CN122734483APending Publication Date: 2026-09-11BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610940658.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-27
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

构建了名为MamCLR的检测框架,旨在重点解决工业场景中样本稀缺的问题,实现高精度、高鲁棒性的小样本故障检测

Benefits of technology

(1)从判别性与重构性两个角度学习故障特征,获得更强的特征表达能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734483A_ABST
    Figure CN122734483A_ABST
Patent Text Reader

Abstract

This invention discloses a few-sample fault detection method based on dual-branch contrastive learning, applicable to the interdisciplinary fields of industrial intelligence, fault diagnosis, and deep learning. The method includes: acquiring the support set and query set for a few-sample task; employing a dual-branch structure sharing the same Mamba encoder; wherein the dual-branch structure includes a feature discrimination branch and a feature reconstruction branch, used to calculate contrastive loss and reconstruction loss respectively; performing double-loop training using a first-order model-independent meta-learning algorithm; wherein the inner loop, based on classification loss, combines contrastive and reconstruction losses to adaptively update model parameters; the outer loop calculates losses using the same method as the inner loop, used to optimize the initial parameters of the meta-model based on the total loss of the outer loop; after training, the shared encoder is retained as an independent classifier. This invention can efficiently extract fault features under few-sample conditions, significantly improving detection accuracy and generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of industrial intelligence, fault diagnosis, and deep learning, and more specifically to a small-sample fault detection method based on dual-branch contrastive learning. Background Technology

[0002] Industrial equipment is widely used in modern production and daily life. However, most industrial equipment operates in harsh environments such as high temperature, high pressure, and high load for extended periods, making it prone to wear and tear and aging of components, leading to various malfunctions. Once equipment fails, it can range from affecting production efficiency and causing economic losses to causing serious safety accidents, threatening the lives of on-site personnel, and resulting in severe consequences. Therefore, fault detection and health management of industrial equipment are of great importance.

[0003] In recent years, various fault detection methods have emerged in the industrial sector. Among them, fault detection methods based on deep learning are currently a research hotspot. Deep learning methods can directly identify and capture fault information from complex monitoring data, achieving automated detection. This approach not only reduces reliance on expert experience but also improves the efficiency of fault detection, providing an effective solution for intelligent equipment monitoring.

[0004] Deep learning methods typically rely on large amounts of data. When data is insufficient, trained deep learning models often exhibit weak generalization ability and low accuracy in fault detection. In real-world production environments, industrial equipment operates normally most of the time. Once a fault occurs, the system is usually shut down immediately for repair and maintenance, rather than continuing to operate in a faulty state. This results in a scarcity of fault samples that can be collected, thus affecting the effectiveness of fault detection. Therefore, research on fault detection methods that utilize small sample data is of great significance.

[0005] For the problem of small sample sizes, existing research mainly employs two approaches: Generative Adversarial Networks (GANs) and Transfer Learning. GANs learn the distribution of real data through a generator and automatically generate simulated fault samples to expand the dataset. However, when the distributions of real and simulated samples differ, the sample quality deteriorates, affecting detection accuracy. Transfer Learning achieves fault detection through cross-domain knowledge transfer, but when the data distributions of the source and target domains differ significantly, the transfer effect is significantly weakened, making it difficult to meet the needs of complex industrial scenarios.

[0006] It is evident that existing small-sample fault detection methods suffer from problems such as insufficient feature representation capabilities, weak model generalization, and poor adaptability to complex working conditions. They are unable to efficiently extract the intrinsic features of fault signals and are difficult to achieve high-precision and robust small-sample fault discrimination.

[0007] Therefore, how to provide a small sample fault detection method based on dual-branch contrastive learning that can effectively solve the problem of sample scarcity in industrial scenarios and achieve high-precision and high-robust small sample fault detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] In view of this, this invention provides a few-shot fault detection method based on dual-branch contrastive learning. A detection framework called MamCLR is constructed to address the problem of scarce samples in industrial scenarios, achieving high-precision and robust few-shot fault detection. This invention employs a dual-branch architecture, with both branches sharing the same Mamba encoder. The first branch is a Mamba-based contrastive learning feature discrimination branch, using Mamba as the encoder for SimCLR to extract highly discriminative fault features. The second branch is a bidirectional Mamba-based signal reconstruction branch, reconstructing the signal using a bidirectional Mamba decoder based on the shared encoder. This reconstruction process forces the model to delve deeper into the intrinsic structure and details of the data, thereby obtaining more expressive features, which, in conjunction with the first branch, achieve a full representation of fault information. Furthermore, this invention jointly optimizes the contrastive loss, reconstruction loss, and classification loss, and uses first-order model-independent parameter learning to complete the two-layer cyclic parameter update, significantly improving the few-shot fault detection performance.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: A few-sample fault detection method based on dual-branch contrastive learning includes: Step 1: Obtain the support set and query set for the few-sample task; the support set is used for adaptive updating of model parameters in the inner loop; the query set is used for initial parameter optimization of the meta-model in the outer loop. Step 2: Adopt a dual-branch structure that shares the same Mamba encoder; wherein, the dual-branch structure includes: a feature discrimination branch and a feature reconstruction branch, which are used to calculate the contrastive loss based on the Mamba contrast discrimination mechanism and the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, respectively; Step 3: A first-order model-independent meta-learning algorithm is used to perform two-layer loop training. The inner loop uses the support set as the model input and is used to adaptively update the model parameters for the current task by combining the contrastive loss and reconstruction loss based on the original feature vectors output by the Mamba encoder. The outer loop uses the query set as the model input after parameter update and calculates the loss in the same way as the inner loop. It is used to optimize the initial parameters of the meta-model based on the total loss of the outer loop. Step 4: After training is complete, the shared Mamba encoder is retained as an independent classifier and used directly for fault category detection and discrimination.

[0010] Optionally, in step 2, the feature discrimination branch is used to calculate the contrastive loss based on the Mamba contrastive discrimination mechanism, specifically as follows: The original samples are processed by a random combination data augmentation strategy to generate two data augmentation samples. The samples are then input into a shared Mamba encoder to extract fault features and obtain the corresponding feature representation vectors. The feature vectors are then fed into a nonlinear projection head consisting of two linear layers and mapped to the contrast learning space. Finally, L2 normalization is performed on the projected features to convert them into vectors of unit length. The augmented data samples from all the original samples after the above encoding, projection, and normalization processes are integrated and concatenated to construct a representation matrix, as follows:

[0011] The representation matrix has the following number of rows: , representing the total number of data augmentation samples among all original samples; the number of columns in the representation matrix is This represents the feature dimension obtained by mapping to the contrast learning space via a nonlinear projection head; and They represent the first The first and second data augmented samples of the original samples, after the above encoding, projection, and normalization processes, output dimensions are: eigenvectors; and They represent the first The feature vectors corresponding to the first and second data augmentation samples of the original samples are... Dimensional value; Calculate the dot product among all eigenvectors in the representation matrix to obtain the similarity matrix, as follows:

[0012] The similarity matrix has 10 rows and 10 columns. The first similarity matrix Line number The elements of the column represent the output after the above encoding, projection, and normalization processes. The eigenvector and the eigenvector The dot product of eigenvectors, after L2 normalization, is equivalent to the cosine similarity between eigenvectors. Construct a mask matrix based on the sample labels; where the mask matrix is ​​a... A two-dimensional square matrix, the mask matrix of the th Line number The elements of the column represent the data augmentation samples from all the original samples. The augmented sample and the first If each augmented sample belongs to the same category, the element value is 1 if it is, and 0 if it is not, and the diagonal elements of the mask matrix are all set to 0. Based on the similarity matrix and mask matrix, a contrastive loss is constructed using the log probabilities of similar positive sample pairs: Based on the mask matrix, the probability of positive sample pairs of the same type in each row of the similarity matrix is ​​calculated as follows:

[0013] in, For the first The probability of positive sample pairs of the same type in a row; For the first The similarity of positive sample pairs of the same type in the same row, i.e., in the mask matrix at the _th ... The similarity corresponding to the element with a value of 1 in the row; The first element in the similarity matrix Line number Sample similarity of columns; It is a natural exponential function; Temperature coefficient; For each row, the probability of positive samples of the same type is negatively logarithmic, and the average value is calculated to obtain the contrastive loss, as follows:

[0014] in, To compare the losses.

[0015] Optionally, the original samples can be processed using a random combination data augmentation strategy to generate two data-augmented samples, specifically: The sampling random combination data augmentation strategy randomly selects two preset signal processing methods and applies them simultaneously to an original sample to generate two different data augmentation samples.

[0016] Optionally, in step 2, the feature reconstruction branch is used to calculate the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, specifically as follows: The original samples are input into a shared Mamba encoder and mapped to low-dimensional latent space feature vectors. Then, the feature vectors are reversed in time domain and fed into the forward and reverse Mamba decoders respectively to obtain forward and reverse reconstructed signals. The forward and reverse reconstructed signals are fused by residual connection to obtain the final reconstructed signal, and the mean square error between the original signal and the final reconstructed signal is used as the reconstruction loss.

[0017] Optionally, in step 3, the classification loss is calculated based on the original feature vector output by the Mamba encoder, specifically as follows: The original feature vector output by the Mamba encoder is input into the linear classification head to obtain the fault category prediction result, and the classification loss between the prediction result and the true label is calculated using the cross-entropy loss function.

[0018] Optionally, in step 3, based on the classification loss, contrastive loss and reconstruction loss are combined to adaptively update the model parameters for the current task, specifically as follows: The contrast loss, reconstruction loss, and classification loss are weighted and fused to obtain the total inner loop loss for the current small sample task. Based on the total inner loop loss, the model parameters are quickly adjusted through backpropagation to complete the adaptive update for the current task.

[0019] Optionally, in step 3, the initial parameters of the meta-model are optimized based on the total loss of the outer loop, specifically as follows: The gradient of the adaptively updated model parameters is obtained by calculating the query set loss on multiple tasks and directly accumulated into the gradient cache of the original meta-model.

[0020] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a small-sample fault detection method based on dual-branch contrastive learning, which has the following beneficial effects: (1) Learn fault features from both discriminative and reconstructive perspectives to obtain stronger feature representation capabilities; (2) Seven types of data enhancement simulation of industrial site noise, variable speed and other interference, to improve the model’s anti-interference ability; (3) Bidirectional Mamba can efficiently capture long-term dependencies in timing signals and avoid the loss of local information; (4) Meta-learning enables the model to adapt quickly with only a very small number of samples, solving the problem of scarce industrial fault samples. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the MamCLR method framework provided by the present invention.

[0023] Figure 2 This is a schematic diagram illustrating the data enhancement effect provided by the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Example 1: Embodiment 1 of this invention discloses a small-sample fault detection method based on dual-branch contrastive learning, named MamCLR, as follows: Figure 1 As shown, it includes: Step 1: Obtain the support set and query set for the few-sample task; the support set is used for adaptive updating of model parameters in the inner loop; the query set is used for initial parameter optimization of the meta-model in the outer loop.

[0026] Each few-sample task consists of a support set and a query set. The support set is used for adaptive updating of model parameters in the inner loop. The total number of samples is obtained by multiplying the number of randomly selected fault categories by the number of samples in each category, and the number of samples in each category is usually set to 1 or 5. The query set is used for initial parameter optimization of the meta-model in the outer loop, aiming to enable the model to generalize across tasks. The total number of samples in the query set is also obtained by multiplying the number of randomly selected fault categories by the number of samples in each category, but the number of samples in each category in the query set is not limited.

[0027] Step 2: Adopt a dual-branch structure that shares the same Mamba encoder; wherein, the dual-branch structure includes: a feature discrimination branch and a feature reconstruction branch, which are used to calculate the contrastive loss based on the Mamba contrast discrimination mechanism and the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, respectively.

[0028] The feature discrimination branch is used to calculate the contrastive loss based on the Mamba contrastive discrimination mechanism, specifically: The original samples are processed by a random combination data augmentation strategy to generate two data augmentation samples. The samples are then input into a shared Mamba encoder to extract fault features and obtain the corresponding feature representation vectors. The feature vectors are then fed into a nonlinear projection head consisting of two linear layers and mapped to the contrast learning space. Finally, L2 normalization is performed on the projected features to convert them into vectors of unit length. The augmented data samples from all the original samples after the above encoding, projection, and normalization processes are integrated and concatenated to construct a representation matrix, as follows:

[0029] The representation matrix has the following number of rows: , representing the total number of data augmentation samples among all original samples; the number of columns in the representation matrix is This represents the feature dimension obtained by mapping to the contrast learning space via a nonlinear projection head; and They represent the first The first and second data augmented samples of the original samples, after the above encoding, projection, and normalization processes, output dimensions are: eigenvectors; and They represent the first The feature vectors corresponding to the first and second data augmentation samples of the original samples are... Dimensional value; Calculate the dot product among all eigenvectors in the representation matrix to obtain the similarity matrix, as follows:

[0030] The similarity matrix has 10 rows and 10 columns. The first similarity matrix Line number The elements of the column represent the output after the above encoding, projection, and normalization processes. The eigenvector and the eigenvector The dot product of eigenvectors, after L2 normalization, is equivalent to the cosine similarity between eigenvectors. Construct a mask matrix based on the sample labels; where the mask matrix is ​​a... A two-dimensional square matrix, the mask matrix of the th Line number The elements of the column represent the data augmentation samples from all the original samples. The augmented sample and the first If each augmented sample belongs to the same category, the element value is 1 if it is, and 0 if it is not, and the diagonal elements of the mask matrix are all set to 0. Based on the similarity matrix and mask matrix, a contrastive loss is constructed using the log probability of similar positive sample pairs, driving the model to aggregate similar fault features and separate dissimilar fault features: Based on the mask matrix, the probability of positive sample pairs of the same type in each row of the similarity matrix is ​​calculated as follows:

[0031] in, For the first The probability of positive sample pairs of the same type in a row; For the first The similarity of positive sample pairs of the same type in the same row, i.e., in the mask matrix at the _th ... The similarity corresponding to the element with a value of 1 in the row; The first element in the similarity matrix Line number Sample similarity of columns; It is a natural exponential function; Temperature coefficient; For each row, the probability of positive samples of the same type is negatively logarithmic, and the average value is calculated to obtain the contrastive loss, as follows:

[0032] in, To compare the losses.

[0033] The original samples are processed using a random combination data augmentation strategy to generate two data-augmented samples, specifically: The sampling random combination data augmentation strategy randomly selects two preset signal processing methods and applies them simultaneously to an original sample to generate two different data augmentation samples.

[0034] In this embodiment of the invention, the preset signal processing method includes: ① Additive white Gaussian noise: Calculate signal power and add Gaussian noise within a specified signal-to-noise ratio range to simulate noise interference in industrial environments; ② Time-domain masking: Randomly mask local segments of the signal in the time domain by setting them to zero to simulate instantaneous signal loss from the sensor; ③ Random amplitude scaling: Adjust the signal amplitude using a random gain coefficient of 0.8 to 1.2 to eliminate dimensional differences in the acquisition equipment; ④ Cyclic shift: Cyclicly shift the signal in the time domain to simulate the difference between the sensor installation position and the sampling trigger time; ⑤ Random trimming and interpolation: Trim local segments of the signal and restore the original length through linear interpolation to enhance global feature learning; ⑥ Frequency domain masking: Transform the signal to the frequency domain and randomly mask specific frequency intervals, then inversely transform it back to the time domain; ⑦ Random resampling: Scale the signal length by 0.9 to 1.1 times and reconstruct it through interpolation to simulate frequency changes under variable speed conditions.

[0035] The data augmentation effect of each signal processing method, such as Figure 2 As shown.

[0036] The feature reconstruction branch is used to calculate the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, specifically: The original samples are input into a shared Mamba encoder and mapped to low-dimensional latent space feature vectors. Then, the feature vectors are reversed in time domain and fed into the forward and reverse Mamba decoders respectively to obtain forward and reverse reconstructed signals. The forward and reverse reconstructed signals are fused by residual connection to obtain the final reconstructed signal. The mean square error between the original signal and the final reconstructed signal is used as the reconstruction loss. By minimizing this loss, the Mamba encoder is forced to learn the intrinsic structure and detailed features of the signal, thereby improving its feature representation ability.

[0037] Step 3: A first-order model-independent meta-learning algorithm is used to perform two-layer loop training. The inner loop uses the support set as the model input. Based on the classification loss calculated from the original feature vectors output by the Mamba encoder, the contrastive loss and reconstruction loss are combined to adaptively update the model parameters for the current task. The outer loop uses the query set as the model input after parameter update, and the loss is calculated in the same way as the inner loop. It is used to optimize the initial parameters of the meta-model based on the total loss of the outer loop.

[0038] The classification loss is calculated based on the original feature vector output by the Mamba encoder, specifically as follows: The original feature vector output by the Mamba encoder is input into the linear classification head to obtain the fault category prediction result, and the classification loss between the prediction result and the true label is calculated using the cross-entropy loss function.

[0039] Based on classification loss, contrastive loss and reconstruction loss are combined to adaptively update the model parameters for the current task, specifically as follows: The contrast loss, reconstruction loss, and classification loss are weighted and fused to obtain the total inner loop loss for the current small sample task. Based on the total inner loop loss, the model parameters are quickly adjusted through backpropagation to complete the adaptive update for the current task.

[0040] The initial parameters of the meta-model are optimized based on the total loss of the outer loop, specifically as follows: The gradient of the adaptively updated model parameters is obtained by calculating the query set loss on multiple tasks and directly accumulated into the gradient cache of the original meta-model.

[0041] Step 4: After training is complete, the shared Mamba encoder is retained as an independent classifier and used directly for fault category detection and discrimination.

[0042] Example 2: Embodiment 2 of the present invention discloses the effectiveness verification of a small-sample fault detection method based on dual-branch contrastive learning, as follows: To verify the effectiveness of the method of this invention, experiments were conducted using the Case Western Reserve University (CWRU) dataset. The CWRU dataset uses electrical discharge machining (EDM) to create single-point damage on bearings, simulating bearing wear in industrial scenarios. Data was acquired through a motor-driven test bench, and raw vibration signals were collected using an accelerometer mounted on the motor housing. The dataset contains 10 categories of faults: in addition to the normal state, it includes fault states of the inner ring, outer ring, and rolling elements at three damage diameters (0.007 inches, 0.014 inches, and 0.021 inches). Specific fault types and their corresponding labels are shown in Table 1. Each category contains 500 samples, each sample being 2048 bytes long. All data were collected under four load conditions: 0, 1, 2, and 3 hp, at a sampling frequency of 12 kHz.

[0043] Table 1. Fault Description of CWRU Dataset

[0044] The specific methods for dividing the training and test sets are shown in Table 2.

[0045] Table 2 CWRU Dataset Scene Division

[0046] The accuracy, precision, recall, and F1 score of each method in different scenarios on the CWRU dataset are shown in Tables 3-6.

[0047] Table 3. Accuracy of each method in different scenarios on the CWRU dataset.

[0048] Table 4. Accuracy of each method in different scenarios on the CWRU dataset.

[0049] Table 5. Recall rates of various methods in different scenarios on the CWRU dataset.

[0050] Table 6. F1 scores of each method in different scenarios on the CWRU dataset.

[0051] As can be seen, the MamCLR proposed in this invention outperforms the comparison methods in all five test scenarios. In scenario 2, MamCLR achieves an accuracy of 99.95%, and even in the worst-performing scenario 5, its accuracy remains at 97.23%, higher than MoCo's 85.20% and ResNet's 77.37%. This demonstrates the effectiveness of MamCLR in fault diagnosis and detection tasks. Furthermore, MamCLR performs well in all scenarios, proving its generalization ability. MamCLR combines Mamba's long sequence modeling capabilities with a contrastive learning mechanism, enabling it to capture key features in signals and improve the model's fault identification performance. In addition, reconstructing network branches forces the model to retain key signal features and filter out redundant information, thus ensuring that the extracted features are the most effective. In summary, MamCLR can handle small-sample fault detection tasks and achieve fault identification and detection.

[0052] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0053] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A small-sample fault detection method based on dual-branch contrastive learning, characterized in that, include: Step 1: Obtain the support set and query set for the few-sample task; wherein, the support set is used for adaptive updating of model parameters in the inner loop; and the query set is used for initial parameter optimization of the meta-model in the outer loop. Step 2: Adopt a dual-branch structure sharing the same Mamba encoder; wherein, the dual-branch structure includes: a feature discrimination branch and a feature reconstruction branch, which are used to calculate the contrast loss based on the Mamba contrast discrimination mechanism and the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, respectively; Step 3: A first-order model-independent meta-learning algorithm is used to perform two-layer loop training. The inner loop uses the support set as the model input and is used to adaptively update the model parameters for the current task by combining the contrastive loss and reconstruction loss based on the classification loss calculated from the original feature vectors output by the Mamba encoder. The outer loop uses the query set as the model input after parameter update and calculates the loss in the same way as the inner loop. It is used to optimize the initial parameters of the meta-model based on the total loss of the outer loop. Step 4: After training is complete, the shared Mamba encoder is retained as an independent classifier and used directly for fault category detection and discrimination.

2. The small-sample fault detection method based on dual-branch contrastive learning according to claim 1, characterized in that, In step 2, the feature discrimination branch is used to calculate the contrastive loss based on the Mamba contrastive discrimination mechanism, specifically as follows: The original samples are processed by a random combination data augmentation strategy to generate two data augmentation samples. The samples are then input into a shared Mamba encoder to extract fault features and obtain the corresponding feature representation vectors. The feature vectors are then fed into a nonlinear projection head consisting of two linear layers and mapped to the contrast learning space. Finally, L2 normalization is performed on the projected features to convert them into vectors of unit length. The augmented data samples from all the original samples after the above encoding, projection, and normalization processes are integrated and concatenated to construct a representation matrix, as follows: Wherein, the number of rows in the representation matrix is , representing the total number of data-augmented samples from all original samples; the number of columns in the representation matrix is ​​. This represents the feature dimension obtained by mapping to the contrast learning space via a nonlinear projection head; and They represent the first The first and second data augmented samples of the original samples, after the above encoding, projection, and normalization processes, output dimensions are: eigenvectors; and They represent the first The feature vectors corresponding to the first and second data augmentation samples of the original samples are... Dimensional value; Calculate the dot product among all eigenvectors in the representation matrix to obtain the similarity matrix, as follows: The similarity matrix has both rows and columns. The first element in the similarity matrix Line number The elements of the column represent the output after the above encoding, projection, and normalization processes. The eigenvector and the eigenvector The dot product of eigenvectors, after L2 normalization, is equivalent to the cosine similarity between eigenvectors. A mask matrix is ​​constructed based on the sample labels; wherein, the mask matrix is ​​a... A two-dimensional square matrix, wherein the mask matrix contains the first... Line number The elements of the column represent the data augmentation samples from all the original samples. The augmented sample and the first If each augmented sample belongs to the same category, the element value is 1 if it is, and 0 if it is not, and the diagonal elements of the mask matrix are all set to 0. Based on the similarity matrix and mask matrix, a contrastive loss is constructed using the log probabilities of similar positive sample pairs: Based on the mask matrix, the probability of positive sample pairs of the same type in each row of the similarity matrix is ​​calculated as follows: in, For the first The probability of positive sample pairs of the same type in a row; For the first The similarity of positive sample pairs of the same type in the same row, i.e., in the mask matrix at the _th ... The similarity corresponding to the element with a value of 1 in the row; The first element in the similarity matrix Line number Sample similarity of columns; It is a natural exponential function; Temperature coefficient; The contrast loss is obtained by taking the negative logarithm of the probabilities of positive samples of the same type in each row and averaging them, as follows: in, To compare the losses.

3. The small-sample fault detection method based on dual-branch contrastive learning according to claim 2, characterized in that, The original samples are processed using a random combination data augmentation strategy to generate two data-augmented samples, specifically: The sampling random combination data augmentation strategy randomly selects two preset signal processing methods and applies them simultaneously to an original sample to generate two different data augmentation samples.

4. The small-sample fault detection method based on dual-branch contrastive learning according to claim 1, characterized in that, In step 2, the feature reconstruction branch is used to calculate the reconstruction loss based on the bidirectional Mamba reconstruction mechanism, specifically as follows: The original samples are input into a shared Mamba encoder and mapped to low-dimensional latent space feature vectors. Then, the feature vectors are reversed in time domain and fed into the forward and reverse Mamba decoders respectively to obtain forward and reverse reconstructed signals. The forward and reverse reconstructed signals are fused by residual connection to obtain the final reconstructed signal, and the mean square error between the original signal and the final reconstructed signal is used as the reconstruction loss.

5. The small-sample fault detection method based on dual-branch contrastive learning according to claim 1, characterized in that, In step 3, the classification loss is calculated based on the original feature vector output by the Mamba encoder, specifically as follows: The original feature vector output by the Mamba encoder is input into the linear classification head to obtain the fault category prediction result, and the classification loss between the prediction result and the true label is calculated using the cross-entropy loss function.

6. The small-sample fault detection method based on dual-branch contrastive learning according to claim 1, characterized in that, In step 3, based on the classification loss, and combined with the contrastive loss and reconstruction loss, the model parameters for the current task are adaptively updated, specifically as follows: The contrast loss, reconstruction loss, and classification loss are weighted and fused to obtain the total inner loop loss for the current small sample task. Based on the total inner loop loss, the model parameters are quickly adjusted through backpropagation to complete the adaptive update for the current task.

7. The small-sample fault detection method based on dual-branch contrastive learning according to claim 1, characterized in that, In step 3, the initial parameters of the meta-model are optimized based on the total loss of the outer loop, specifically as follows: The gradient of the adaptively updated model parameters is obtained by calculating the query set loss on multiple tasks and directly accumulated into the gradient cache of the original meta-model.