Small sample cross-domain fault diagnosis method based on multistage feature alignment

By constructing a diffusion model to generate enhanced samples of the target domain feature distribution and combining it with a multi-level feature alignment strategy, the problem of insufficient migration performance in small sample cross-domain fault diagnosis is solved, efficient diagnosis in complex industrial scenarios is achieved, and the adaptability and diagnostic accuracy of the model are improved.

CN120611234AActive Publication Date: 2025-09-09NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510627892.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-09
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Existing cross-domain fault diagnosis methods have difficulty achieving effective feature transfer and diagnosis under small sample conditions, especially in complex industrial scenarios. They suffer from negative transfer problems and insufficient diagnostic accuracy, and are unable to effectively distinguish fine-grained fault states. In addition, the scarcity of target domain data leads to model overfitting and degradation of diagnostic performance.

Method used

A small-sample cross-domain fault diagnosis method based on multi-level feature alignment is adopted. By constructing a diffusion model to generate enhanced samples of the target domain feature distribution, a multi-head self-attention mechanism and a normalization strategy are combined to design a joint loss function, including source domain classification loss, dual MMD distribution alignment loss, cross-domain triplet loss and consistency regularization term, to achieve unified modeling and accurate diagnosis of fault features.

Benefits of technology

Under small sample conditions, it significantly improves the ability to distinguish fault features and the cross-domain adaptability of the model, improves the diagnostic performance and stability under complex working conditions, overcomes the migration performance bottleneck caused by the scarcity of target domain samples, and has good industrial application potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611234A_ABST
    Figure CN120611234A_ABST
Patent Text Reader

Abstract

The invention provides a small sample cross-domain fault diagnosis method based on multistage feature alignment, and the method comprises the steps: collecting rotating machinery monitoring data in a source domain and a target domain, and constructing a cross-domain small sample fault diagnosis task data set with scarce target domain samples; constructing a diffusion model combining a multi-head self-attention mechanism and a normalization strategy, and generating an enhanced sample with target domain feature distribution; a shared feature extraction network fusing a source domain, a target domain and a generated sample is built, and unified modeling of diagnosis features is achieved; and designing a joint loss function which comprises source domain classification loss, dual-MMD distribution alignment loss, cross-domain triple loss and enhanced consistency regular terms, constructing a multi-stage feature alignment mechanism, and realizing fault feature migration and accurate diagnosis under the condition of small samples of a target domain. According to the method, source domain diagnosis knowledge can be effectively migrated under the condition of target domain data scarcity, the fault feature discrimination capability and the cross-domain adaptability of the model are enhanced, and the method is suitable for intelligent diagnosis tasks of the rotating machinery under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of small sample cross-domain fault diagnosis, and in particular to a small sample cross-domain fault diagnosis method based on multi-level feature alignment. Background Art

[0002] Rotating machinery is widely used in critical industrial applications such as energy, power generation, and manufacturing. Failure of core components such as bearings and gears can easily lead to unplanned equipment downtime, resulting in significant economic losses and safety hazards. Therefore, intelligent fault diagnosis technology has been given a crucial role in the health management of industrial systems.

[0003] In recent years, deep learning-based fault identification methods have made significant progress, demonstrating excellent diagnostic accuracy, particularly when large-scale annotated data is available. However, in real industrial environments, due to factors such as variable equipment operating conditions, limited sensor layout, and high annotation costs, the number of samples in the target domain (i.e., the scenario to be diagnosed) is often extremely scarce. Furthermore, significant differences in the distribution of signal characteristics between different devices, operating conditions, or measurement points can result in poor direct diagnostic results from the model, severely restricting the widespread application of deep diagnostic models in real-world scenarios across multiple devices and operating conditions.

[0004] To overcome the above challenges, scholars have proposed a variety of cross-domain fault diagnosis methods based on transfer learning in recent years, but the following technical bottlenecks still exist:

[0005] First, existing methods often rely on adversarial training (such as domain discriminators) or statistical distribution alignment strategies (such as maximum mean difference (MMD)) to narrow the feature distribution differences between the source and target domains. While these methods achieve a certain degree of consistency in inter-domain features, they typically focus only on alignment at the overall distribution level, neglecting the modeling and utilization of category structure information. This makes it difficult to ensure close aggregation between similar samples and sufficient separation between heterogeneous samples after cross-domain classification. This alignment method, with its fuzzy inter-class boundaries, is prone to negative transfer problems. This is particularly true in complex diagnostic tasks with multiple fault types and large inter-sample differences. It cannot effectively distinguish fine-grained fault states, severely limiting the diagnostic accuracy and generalization capabilities of the model.

[0006] Second, given the extreme scarcity of target domain data, existing methods often employ methods such as direct transfer or pseudo-label guidance to assist the model in learning target domain features. However, due to the limited number of target samples and the limited accuracy of pseudo-labels, the model is prone to falling into a cycle of false supervision during training. Furthermore, feature extraction networks are prone to overfitting, resulting in a lack of discriminativeness in the extracted target domain features, ultimately leading to degradation in diagnostic performance. Furthermore, the lack of effective feature enhancement mechanisms for small sample sizes makes it difficult for existing methods to fully exploit the potential structural information in limited data, resulting in insufficient stability and reliability in actual industrial deployments.

[0007] Therefore, there is an urgent need to propose a new cross-domain diagnosis method with multi-level feature alignment capability, which can achieve accurate transfer of diagnostic knowledge from the source domain to the target domain under small sample conditions, thereby improving the adaptability and diagnostic performance of the model in complex open environments. Summary of the Invention

[0008] To address the aforementioned technical issues, a small-sample cross-domain fault diagnosis method based on multi-level feature alignment is provided. This method effectively transfers source domain diagnostic knowledge when target domain data is scarce, enhancing the ability to discriminate fault features and the cross-domain adaptability of the model. It is suitable for intelligent diagnosis of rotating machinery under complex operating conditions.

[0009] The technical means adopted in the present invention are as follows:

[0010] A small sample cross-domain fault diagnosis method based on multi-level feature alignment, comprising:

[0011] S1. Collect rotating machinery monitoring data in the source domain and target domain, and construct a cross-domain small sample fault diagnosis task dataset where target domain samples are scarce.

[0012] S2. Construct a diffusion model that combines a multi-head self-attention mechanism with a normalization strategy to generate enhanced samples with the characteristic distribution of the target domain;

[0013] S3. Build a shared feature extraction network that integrates the source domain, target domain, and generated samples to achieve unified modeling of diagnostic features;

[0014] S4. Design a joint loss function, including source domain classification loss, dual MMD distribution alignment loss, cross-domain triplet loss and enhanced consistency regularization term, to build a multi-level feature alignment mechanism to achieve fault feature migration and accurate diagnosis in the case of small samples in the target domain.

[0015] Furthermore, step S1 specifically includes:

[0016] S11. Collect the monitoring signals of the source domain and the target domain on the rotating mechanical platform as the original signals, where the monitoring signals of the source domain are labeled data. The monitoring signal of the target domain is unlabeled data And the target domain obtained has no labeled data For very small amounts;

[0017] S12. Normalize the collected original signal and scale it to the interval [0,1]. The normalization formula is as follows:

[0018]

[0019] Among them, x norm represents the standardized signal, x represents the original signal point, min(x) represents the minimum value of the original signal, and max(x) represents the maximum value of the original signal. In this embodiment, the collected original signal is normalized to eliminate the influence of the signal amplitude difference under different working conditions on the training and improve the convergence stability of the model.

[0020] S13. Use the sliding window algorithm to segment the original long time series, with each segment containing 2048 sampling points, and construct a standard input sample format.

[0021] Furthermore, in step S2, a diffusion model is constructed based on the improved U-Net to enhance the data samples and generate high-quality signal samples. Includes forward process, reverse process and noise prediction.

[0022] Furthermore, in the forward process, noise is gradually added to a data sample until it approaches a Gaussian distribution, specifically including:

[0023] Assuming the data is x0, the noise process generates a series of samples x1, x2, ···, x through a Markov chain of multiple time steps. T , x T is Gaussian distributed noise, and the mathematical expression of this process is as follows:

[0024]

[0025] Among them, q(x 1:T |x0) means that given the initial sample x0, a series of intermediate samples x are generated. 1:T The joint probability distribution of x 1:T represents a series of intermediate samples, q(x t |x t-1 ) means that at a given x t-1 Under the conditions, generate x t The probability distribution of x t-1 represents the sample at time step t-1, β trepresents a hyperparameter, representing the amplitude of added noise, T represents the total number of diffusion steps, this process gradually adds noise from the original data to a set of pure noise, and β1<β2···<β T , indicating that the added noise is getting larger and larger, and I represents a matrix with the same shape as the input data.

[0026] Furthermore, in the reverse process, the true distribution of the data is gradually restored by denoising, and a neural network with parameters is used to estimate the conditional probability. The calculation formula is as follows:

[0027]

[0028] Among them, θ represents the learnable parameters of the neural network, μ θ and Σ θ represents the mean and variance; u θ (x t ,t) means that at time step t, the model is based on the current noisy sample x t The estimated amount of denoising, α t represents the scaling factor, x t-1 represents the sample at the previous moment, z θ represents the noise at time step t obtained from the prediction.

[0029] Furthermore, in the noise prediction, an improved Unet network architecture is used to predict the noise distribution in the reverse process, specifically including:

[0030] The multi-head attention mechanism is introduced in the encoder and decoder so that the model can capture long-range dependencies in the signal. The multi-head attention mechanism module includes one-dimensional convolution, batch normalization and multi-head attention mechanism. The mathematical formula is as follows:

[0031]

[0032] Among them, Q j represents the query matrix of the jth attention head, x represents the original signal point, K j represents the key matrix of the j-th attention head, V j represents the value matrix of the j-th attention head, and Represents a trainable projection matrix, Attention(Q,K,V) j represents the attention mechanism, represents the transpose of the key matrix of the j-th attention head, d k Represents the dimension of the key vector, MSA(Q,K,V) j Represents the multi-head attention mechanism, head jrepresents the output of the jth attention head, head j =Attention(Q,K,V) j ,j=1,2,...,H, H represents the number of attention heads, Represents a trainable projection matrix.

[0033] Furthermore, during the diffusion model training process, the noise prediction network based on U-Net needs to be constrained to calculate the loss value between the predicted noise and the actual noise. The calculation formula is as follows:

[0034]

[0035] in, represents the expected operation, z t represents the true noise added to the data at time step t, represents the cumulative noise coefficient at time step t.

[0036] Furthermore, step S3 specifically includes:

[0037] S31. Construct a shared parameter feature extraction network with three types of inputs: source domain samples, target domain samples, and target domain generated samples. The network structure includes multiple convolutional layers, pooling layers, and normalization layers, and is ultimately mapped to the fault category space.

[0038] S32. After feature extraction, all samples are sent to the classifier, and the prediction results are output and used for subsequent joint loss function calculation.

[0039] Furthermore, step S4 specifically includes:

[0040] S41. To ensure that the model accurately identifies the category features in the source domain labeled samples and learns the fault difference features in the original data, the cross entropy classification loss function is introduced. The mathematical formula is as follows:

[0041]

[0042] Among them, n s represents the number of samples in the source domain, x i represents the i-th source domain sample, Represents x i The test tag, Represents x i The true label of

[0043] S42. To improve the model's sensitivity and adaptability to cross-domain distribution changes, the present invention extends the traditional MMD alignment strategy to a dual alignment objective: on the one hand, aligning the overall feature distribution between the source domain samples and the true target domain samples; on the other hand, aligning the distribution between the source domain samples and the target samples generated by the diffusion model to ensure that the generated data can truly reflect the characteristics of the target domain. The specific mathematical expression is:

[0044]

[0045] Among them, φ(·) represents mapping the input features into the Reproducing Kernel Hilbert Space (RKHS), represents the source domain sample, represents the target domain sample, represents the target domain generated sample, n s Indicates the number of source domain samples, n t Indicates the number of target domain samples, n′ t Indicates the number of generated target domain samples;

[0046] S43. In order to enhance the intra-class compactness and inter-class separability of cross-domain samples in the feature space, a cross-domain triplet loss function is designed. The mathematical expression is as follows:

[0047]

[0048] Among them, the anchor sample a, positive sample p and negative sample n come from the source domain sample or target domain samples Positive samples and anchor samples belong to the same category, while negative samples come from different categories. Function f(·) represents the feature extraction representation of the sample, f(a) represents the feature representation of the anchor, f(p) represents the feature representation of the positive sample, f(n) represents the feature representation of the negative sample, and δ represents the boundary interval hyperparameter, which is used to control the minimum distinction between positive and negative samples.

[0049] S44. To further improve the robustness of the model, the present invention introduces a consistency regularization loss mechanism to constrain the prediction consistency of the original sample of the target domain and its enhanced sample in the feature extraction results. The expression of the consistency regularization loss function is as follows:

[0050]

[0051] in, Represents the original sample of the target domain Enhanced samples of

[0052] S45. In the present invention, both cross-domain triplet loss and consistency regularization loss involve pseudo-label prediction for target domain samples. Although this type of pseudo-label has a positive effect on improving the model's discrimination ability and generalization performance in the later stages of training, in the early stages of training, since the model has not yet converged, the generated pseudo-labels often have large errors. If they are directly involved in loss optimization, it may lead to training instability or even negative transfer, thereby weakening the effects of triplet loss and consistency loss. To this end, the present invention introduces a dynamic weight adjustment parameter α in the optimization process to gradually control the influence of the loss term corresponding to the pseudo-label sample. The adjustment function of the parameter α is defined as follows:

[0053]

[0054] Where i represents the current training epoch, E represents the total training epochs, and γ represents the growth rate control factor, which is set to 10.

[0055] S46. The cross-domain diagnosis network achieves efficient feature alignment and improved discrimination performance under the condition of limited target domain data through the combined optimization of multiple loss functions. The overall loss function expression is as follows:

[0056]

[0057] Among them, λ and α represent the weight adjustment parameters of the combination of MMD loss term, triple loss and consistency regularization loss term, respectively. represents the cross entropy classification loss function, represents the dual alignment objective loss function, represents the cross-domain triplet loss function, represents the consistency regularization loss function.

[0058] Compared with the prior art, the present invention has the following advantages:

[0059] The present invention provides a small-sample cross-domain fault diagnosis method based on multi-level feature alignment, which shows significant performance advantages in cross-domain tasks of bearing working conditions. It is significantly better than traditional non-enhancement or GAN enhancement methods, effectively overcomes the migration performance bottleneck caused by the scarcity of target domain samples, and has good industrial application potential.

[0060] Based on the above reasons, the present invention can be widely promoted in fields such as small sample cross-domain fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0062] Figure 1 Flow chart of the method of the present invention.

[0063] Figure 2 Schematic diagram of the division of source domain and target domain measurement point data sets in the present invention.

[0064] Figure 3 Schematic diagram of the diffusion model process of the present invention.

[0065] Figure 4 This is the network diagram of the multi-head attention mechanism of the present invention.

[0066] Figure 5 This is the overall structure diagram of the fault diagnosis network of the present invention.

[0067] Figure 6 This is a physical picture of the data acquisition device of the present invention

[0068] Figure 7 This is the fault diagnosis confusion matrix diagram of the present invention.

[0069] Figure 8 This is the time-frequency diagram of the generated signal and experimental signal of the bearing inner ring fault of the present invention. DETAILED DESCRIPTION

[0070] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0071] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.

[0072] like Figure 1As shown, the present invention provides a small sample cross-domain fault diagnosis method based on multi-level feature alignment, comprising:

[0073] S1. Collect rotating machinery monitoring data in the source domain and target domain, and construct a cross-domain small sample fault diagnosis task dataset where target domain samples are scarce.

[0074] S2. Construct a diffusion model that combines a multi-head self-attention mechanism with a normalization strategy to generate enhanced samples with the characteristic distribution of the target domain;

[0075] S3. Build a shared feature extraction network that integrates the source domain, target domain, and generated samples to achieve unified modeling of diagnostic features;

[0076] S4. Design a joint loss function, including source domain classification loss, dual MMD distribution alignment loss, cross-domain triplet loss and enhanced consistency regularization term, to build a multi-level feature alignment mechanism to achieve fault feature migration and accurate diagnosis in the case of small samples in the target domain.

[0077] In specific implementation, as a preferred embodiment of the present invention, step S1 specifically includes:

[0078] S11. Collect monitoring signals of the source domain (such as working condition A) and the target domain (such as working condition B) on the rotating machinery platform as original signals, where the monitoring signals of the source domain are labeled data. The monitoring signal of the target domain is unlabeled data And the target domain obtained has no labeled data For very small amounts;

[0079] S12. Normalize the collected original signal and scale it to the interval [0,1]. The normalization formula is as follows:

[0080]

[0081] Among them, x norm represents the standardized signal, x represents the original signal point, min(x) represents the minimum value of the original signal, and max(x) represents the maximum value of the original signal. In this embodiment, the collected original signal is normalized to eliminate the influence of the signal amplitude difference under different working conditions on the training and improve the convergence stability of the model.

[0082] S13, such as Figure 2 As shown in Figure 1, the sliding window algorithm is used to segment the original long time series, each segment contains 2048 sampling points, and the standard input sample format is constructed.

[0083] In specific implementation, as a preferred embodiment of the present invention, in step S2, a diffusion model is constructed based on the improved U-Net to enhance the data sample and generate a high-quality signal sample. like Figure 3 As shown, it includes the forward process, the reverse process and the noise prediction.

[0084] In a specific implementation, as a preferred embodiment of the present invention, the forward process gradually adds noise to a data sample until it approaches a Gaussian distribution, specifically including:

[0085] Assuming the data is x0, the noise process generates a series of samples x1, x2, ···, x through a Markov chain of multiple time steps. T , x T is Gaussian distributed noise, and the mathematical expression of this process is as follows:

[0086]

[0087] Among them, q(x 1:T |x0) means that given the initial sample x0, a series of intermediate samples x are generated. 1:T The joint probability distribution of x 1:T represents a series of intermediate samples, q(x t |x t-1 ) means that at a given x t-1 Under the conditions, generate x t The probability distribution of x t-1 represents the sample at time step t-1, β t represents a hyperparameter, representing the amplitude of added noise, T represents the total number of diffusion steps, this process gradually adds noise from the original data to a set of pure noise, and β1<β2···<β T , indicating that the added noise is getting larger and larger, and I represents a matrix with the same shape as the input data.

[0088] In specific implementation, as a preferred embodiment of the present invention, in the reverse process, the true distribution of the data is gradually restored by denoising, and a neural network with parameters is used to estimate the conditional probability. The calculation formula is as follows:

[0089]

[0090] Among them, θ represents the learnable parameters of the neural network. Currently, the Unet model is commonly used to estimate the original noise distribution. θ and Σ θ represents the mean and variance; u θ (x t ,t) means that at time step t, the model is based on the current noisy sample x t The estimated amount of denoising, α t represents the scaling factor, x t-1 represents the sample at the previous moment, z θ represents the noise at time step t obtained from the prediction.

[0091] In specific implementation, as a preferred embodiment of the present invention, in the noise prediction, an improved Unet network architecture is used to predict the noise distribution in the reverse process, specifically including:

[0092] The introduction of multi-head attention mechanism in encoder and decoder enables the model to capture long-range dependencies in the signal, such as Figure 4 As shown in the figure, the multi-head attention mechanism module includes one-dimensional convolution, batch normalization and multi-head attention mechanism. The mathematical formula is as follows:

[0093]

[0094] Among them, Q j represents the query matrix of the jth attention head, x represents the original signal point, K j represents the key matrix of the j-th attention head, V j represents the value matrix of the j-th attention head, and Represents a trainable projection matrix, Attention(Q,K,V) j represents the attention mechanism, represents the transpose of the key matrix of the j-th attention head, d k Represents the dimension of the key vector, MSA(Q,K,V) j Represents the multi-head attention mechanism, head j represents the output of the jth attention head, head j =Attention(Q,K,V) j ,j=1,2,...,H, H represents the number of attention heads, Represents a trainable projection matrix.

[0095] In specific implementation, as a preferred embodiment of the present invention, the noise prediction network based on U-Net needs to be constrained during the diffusion model training process, and the loss value between the predicted noise and the actual noise is calculated using the following formula:

[0096]

[0097] in, represents the expected operation, z t represents the true noise added to the data at time step t, represents the cumulative noise coefficient at time step t.

[0098] In specific implementation, as a preferred embodiment of the present invention, step S3 specifically includes:

[0099] S31. Construct a shared parameter feature extraction network with three types of inputs: source domain samples, target domain samples, and target domain generated samples. The network structure includes multiple convolutional layers, pooling layers, and normalization layers, and is ultimately mapped to the fault category space.

[0100] S32, all samples are sent to the classifier after feature extraction, and the prediction results are output and used for subsequent joint loss function calculation. Figure 5 As shown in Figure 1, the overall structure of the fault diagnosis network is trained by three types of data to achieve the generalization learning of the fault model for the fault mode of the small sample target domain.

[0101] In specific implementation, as a preferred embodiment of the present invention, step S4 specifically includes:

[0102] S41. To ensure that the model accurately identifies the category features in the source domain labeled samples and learns the fault difference features in the original data, the cross entropy classification loss function is introduced. The mathematical formula is as follows:

[0103]

[0104] Among them, n s represents the number of samples in the source domain, x i represents the i-th source domain sample, Represents x i The test tag, Represents x i The true label of

[0105] S42. To improve the model's sensitivity and adaptability to cross-domain distribution changes, the present invention extends the traditional MMD alignment strategy to a dual alignment objective: on the one hand, aligning the overall feature distribution between the source domain samples and the true target domain samples; on the other hand, aligning the distribution between the source domain samples and the target samples generated by the diffusion model to ensure that the generated data can truly reflect the characteristics of the target domain. The specific mathematical expression is:

[0106]

[0107] Among them, φ(·) represents mapping the input features into the Reproducing Kernel Hilbert Space (RKHS), represents the source domain sample, represents the target domain sample, represents the target domain generated sample, n s Indicates the number of source domain samples, n t Indicates the number of target domain samples, n t′ represents the number of generated target domain samples.

[0108] In this embodiment, MMD, a parameter-free, non-adversarial alignment method, offers the advantages of stable training and ease of implementation, making it particularly suitable for feature transfer in small-sample target domain scenarios. Jointly minimizing two MMD loss terms helps achieve global alignment between source and true target domain samples, as well as generated target domain samples, thereby promoting more effective domain adaptation.

[0109] S43. In order to enhance the intra-class compactness and inter-class separability of cross-domain samples in the feature space, a cross-domain triplet loss function is designed. The mathematical expression is as follows:

[0110]

[0111] Among them, the anchor sample a, positive sample p and negative sample n come from the source domain sample or target domain samples The positive samples and the anchor samples belong to the same category, while the negative samples come from different categories. The function f(·) represents the feature extraction representation of the sample, f(a) represents the feature representation of the anchor, f(p) represents the feature representation of the positive sample, f(n) represents the feature representation of the negative sample, and δ represents the boundary interval hyperparameter, which is used to control the minimum distinction between positive and negative samples. In this embodiment, the loss function optimizes the discriminative ability of the feature embedding space by minimizing the distance between samples of the same category (anchor samples and positive samples) and simultaneously expanding the interval between samples of different categories (anchor samples and negative samples).

[0112] S44. To further improve the robustness of the model, the present invention introduces a consistency regularization loss mechanism to constrain the prediction consistency of the original sample of the target domain and its enhanced sample in the feature extraction results. The expression of the consistency regularization loss function is as follows:

[0113]

[0114] in, Represents the original sample of the target domain enhanced samples; in this embodiment, the core idea of ​​the loss term is that when facing the same target sample after different data enhancement or perturbation, the model should maintain the consistency of feature output, thereby enhancing its stability to input perturbations and significantly improving its generalization ability under the condition of small samples in the target domain. Specifically, by performing the enhancement operation generated by the diffusion model on the target domain sample, its transformed form is constructed, and paired comparison is performed with the original sample in the feature space to constrain its output distance to be minimum. This strategy not only effectively suppresses the overfitting problem caused by the scarcity of target domain data, but also enhances the robustness of the model when facing input fluctuations in actual complex working conditions.

[0115] S45. In the present invention, both cross-domain triplet loss and consistency regularization loss involve pseudo-label prediction for target domain samples. Although this type of pseudo-label has a positive effect on improving the model's discrimination ability and generalization performance in the later stages of training, in the early stages of training, since the model has not yet converged, the generated pseudo-labels often have large errors. If they are directly involved in loss optimization, it may lead to training instability or even negative transfer, thereby weakening the effects of triplet loss and consistency loss. To this end, the present invention introduces a dynamic weight adjustment parameter α in the optimization process to gradually control the influence of the loss term corresponding to the pseudo-label sample. The adjustment function of the parameter α is defined as follows:

[0116]

[0117] Where i represents the current training epoch, E represents the total training epochs, and γ represents the growth rate control factor, which is 10. In this embodiment, the parameter α is set to 0 at the beginning of training. As the model performance improves, the α value gradually increases to 1, thereby gradually guiding the model to participate in training using pseudo-label information, enhancing robustness and training effect.

[0118] S46. The cross-domain diagnosis network achieves efficient feature alignment and improved discrimination performance under the condition of limited target domain data through the combined optimization of multiple loss functions. The overall loss function expression is as follows:

[0119]

[0120] Among them, λ and α represent the weight adjustment parameters of the combination of MMD loss term, triple loss and consistency regularization loss term, respectively. represents the cross entropy classification loss function, represents the dual alignment objective loss function, represents the cross-domain triplet loss function, Denotes the consistency regularization loss function. In this embodiment, by dynamically jointly optimizing the above loss function, the present invention achieves efficient alignment of multi-level features between source domain, true target domain, and generated target domain samples, providing strong support for small-sample intelligent fault diagnosis under complex working conditions.

[0121] Example

[0122] The data set used in this invention comes from the rolling bearing fault test platform, and its structure is as follows: Figure 6As shown in the figure, the test platform contains two rolling bearings, of which the right bearing serves as the bearing under test. Vibration signals are collected using an accelerometer mounted on the bearing housing, with a sampling frequency of 20kHz. To construct a variety of typical fault conditions, various types of faults are artificially introduced into the bearing, including rolling element fault (BF), inner race fault (IF), outer race fault (OF), combined fault (OF-BF), and healthy state (H). The above faults are imposed on the test bearings through precision machining to ensure the consistency and repeatability of the experiment. In addition, to simulate inter-domain differences, the test platform is set up with two different speed conditions: 1200 rpm and 1800 rpm, representing two independent operating domains. The vibration signals collected under these conditions constitute the data samples of the source and target domains, which are used for the subsequent verification and performance evaluation of the small sample cross-domain fault diagnosis method.

[0123] In order to fully verify the effectiveness and generalization ability of the small sample cross-domain fault diagnosis method proposed in this invention, this embodiment designs two cross-domain fault diagnosis tasks (T1 and T2) based on the rolling bearing dataset. In task T1, the bearing data with a rotation speed of 1800r / min is selected as the source domain, and the bearing data with a rotation speed of 1200r / min is selected as the target domain; in task T2, the source domain and the target domain are set oppositely, that is, the source domain is 1200r / min and the target domain is 1800r / min. In each of the above tasks, 300 labeled samples are selected for each health state in the source domain, and each state of the target domain initially contains only 20 unlabeled samples, simulating the cross-domain fault diagnosis scenario with scarce data in reality. In order to alleviate the problem of scarcity of target domain samples, the diffusion model is used to perform data enhancement on the target domain samples, so that each type of fault state is expanded to generate 280 enhanced samples, further supporting the feature alignment and model optimization process.

[0124] The cross-domain fault diagnosis method proposed in this invention adopts a time-domain vibration signal with an input dimension of 1×2048. The feature extraction network consists of 5 convolutional layers and 1 fully connected layer. The number of output channels of each convolutional layer is 4, 8, 16, 32 and 64 respectively, and all include convolution, batch normalization, LeakyReLU activation and maximum pooling operations. The output dimension of the fully connected layer is consistent with the number of fault categories, and the dropout ratio is set to 0.1. The optimizer uses SGD, the initial learning rate is 0.01, and it decays to half of the original value every 50 rounds. The total number of training rounds is 200, and the batch size is set to 32. In order to reduce the influence of randomness, all experiments are repeated 6 times and the average results are taken. The hyperparameter λ is set to 0.5. In order to verify the effectiveness of this method, three types of comparison methods are designed:

[0125] (1) Non-DA, which does not introduce any domain adaptation mechanism and relies only on source domain classification training;

[0126] (2) Non-AU, that is, no generated samples are used, and only a small number of target domain samples and MMD loss are used for alignment;

[0127] (3) DA-GAN, which uses a generative adversarial network (GAN) to generate target domain samples, is compared with the diffusion model used in this paper.

[0128] The above comparison methods all use the same network structure and parameter settings to ensure experimental fairness. Figure 7 The following is a time-frequency comparison of the generated signal and the experimental signal for a bearing inner race fault. The Fourier-transformed time-frequency plot shows that the spectrum of the generated signal is consistent with the true signal, particularly in terms of amplitude and phase matching at key frequency components. In the case of a ball fault in particular, the frequency components of the generated signal are nearly identical to those of the true signal. The low-frequency characteristics of the fault signals generated by the inner and outer rings match well. The signal generated by the diffusion model not only reconstructs the dynamic characteristics of the true signal but also accurately preserves its frequency information. This validates the effectiveness of the proposed DDPM and demonstrates its ability to provide high-quality data for fault diagnosis in small sample sizes.

[0129] To quantitatively compare the differences between the signals generated by different models and the true signals, root mean square error (RMSE) and spectral cosine similarity (FSCS) were introduced as evaluation metrics. Table 1 shows the evaluation results of the generated signals produced by different models. It can be observed that the proposed model achieves the best performance across all fault types. Across the five fault types, the proposed model achieves an average RMSE of 0.054, which is 0.01 lower than the original DDPM. The average PSNR is 25.67, an improvement of 1.127 over the original DDPM. The signal quality generated by the DA-GAN model is significantly lower than that of the DDPM, with average RMSE and PSNR values ​​of only 0.121 and 13.562, respectively. This is primarily because DA-GAN lacks the iterative denoising process and probabilistic modeling capabilities provided by the diffusion-based framework, which limits its ability to capture complex signal distributions and effectively suppress noise. In summary, the proposed improved DDPM model demonstrates superior feature modeling and denoising capabilities, thereby improving the reconstruction quality and robustness across various fault types.

[0130] Table 1 Comparison of signals generated by different models

[0131] Model index Roller failure Roller failure Inner race fault Mixed failure healthy average Method of the present invention RMSE 0.043 0.055 0.045 0.046 0.081 0.054 PSNR 27.432 25.208 26.985 26.857 21.867 25.670 Original diffusion model RMSE 0.054 0.065 0.049 0.059 0.095 0.064 PSNR 25.103 24.441 26.154 26.458 20.558 24.543 DA-GAN RMSE 0.109 0.113 0.098 0.116 0.168 0.121 PSNR 14.258 13.565 14.256 13.879 11.854 13.562

[0132] Small sample cross-domain fault diagnosis analysis:

[0133] Different comparison methods were used to test the fault diagnosis performance in these two tasks, and the statistical results are shown in Table 2. In task T1, the average diagnostic accuracy of the Non-DA method, which did not introduce any domain adaptation mechanism, was 72.12%, while the accuracy of the Non-AU method, which did not use data enhancement but adopted the MMD alignment strategy, was increased to 82.86%. Furthermore, the accuracy of the DA-GAN method, which used generative adversarial networks (GANs) for data enhancement, was increased to 89.25%. In comparison, the diffusion model-assisted method proposed in this invention performed best in the T1 task, with an accuracy of 93.83%. In task T2, the performance trends of each method were similar. The accuracy of the Non-DA method was 64.89%, the Non-AU method was 75.24%, the DA-GAN method was 83.56%, and the accuracy of the method of this invention was the highest, reaching 90.07%. The above results show that models trained solely on the source domain experience a significant decline in diagnostic performance when diagnosing the target domain under different speed conditions. While appropriate introduction of target domain data (such as Non-AU and DA-GAN) can improve performance, this is still limited by the number of samples and the quality of the generated data. In contrast, the present invention significantly improves cross-domain generalization capabilities and achieves higher-precision fault identification by constructing a diffusion model to generate realistic and diverse target domain enhanced samples. Furthermore, it combines a multi-level feature alignment strategy, including source domain classification loss, dual MMD alignment loss, cross-domain triplet loss, and consistency regularization loss.

[0134] Table 2 Statistics of fault diagnosis results

[0135]

[0136]

[0137] In addition, in order to further illustrate the diagnostic effect of the method of the present invention under various fault conditions, the confusion matrix of all methods under the T1 task is selected for visualization analysis, as shown in the following figure: Figure 8 The results show that the Non-DA and Non-AU methods have low recognition accuracy under rolling element fault (BF) and normal state (N), while the proposed method achieves relatively balanced and accurate diagnosis results under all fault types, further verifying the stability and generalization ability of the proposed method under small sample conditions.

[0138] In summary, the small-sample cross-domain fault diagnosis method based on diffusion model and multi-level feature alignment proposed in the present invention shows significant performance advantages in the cross-domain task of bearing working conditions, which is significantly better than traditional non-enhancement or GAN enhancement methods, effectively overcomes the migration performance bottleneck caused by the scarcity of target domain samples, and has good industrial application potential.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A small sample cross-domain fault diagnosis method based on multi-level feature alignment, characterized by: include: S1. Collect rotating machinery monitoring data in the source domain and target domain, and construct a cross-domain small sample fault diagnosis task dataset where target domain samples are scarce. S2. Construct a diffusion model that combines a multi-head self-attention mechanism with a normalization strategy to generate enhanced samples with the characteristic distribution of the target domain; S3. Build a shared feature extraction network that integrates the source domain, target domain, and generated samples to achieve unified modeling of diagnostic features; S4. Design a joint loss function, including source domain classification loss, dual MMD distribution alignment loss, cross-domain triplet loss and enhanced consistency regularization term, to build a multi-level feature alignment mechanism to achieve fault feature migration and accurate diagnosis in the case of small samples in the target domain.

2. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 1 is characterized in that: Step S1 specifically includes: S11. Collect the monitoring signals of the source domain and the target domain on the rotating mechanical platform as the original signals, where the monitoring signals of the source domain are labeled data. The monitoring signal of the target domain is unlabeled data And the target domain obtained has no labeled data For very small amounts; S12. Normalize the collected original signal and scale it to the interval [0,1]. The normalization formula is as follows: Among them, x norm represents the normalized signal, x represents the original signal point, min(x) represents the minimum value of the original signal, and max(x) represents the maximum value of the original signal; S13. Use the sliding window algorithm to segment the original long time series, with each segment containing 2048 sampling points, and construct a standard input sample format.

3. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 1 is characterized in that: In step S2, a diffusion model is constructed based on the improved U-Net to enhance the data samples and generate high-quality signal samples. Includes forward process, reverse process and noise prediction.

4. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 3 is characterized in that: In the forward process, noise is gradually added to a data sample until it approaches a Gaussian distribution, specifically including: Assuming the data is x0, the noise process generates a series of samples x1, x2, ···, x through a Markov chain of multiple time steps. T , x T is Gaussian distributed noise, and the mathematical expression of this process is as follows: Among them, q(x 1:T |x0) means that given the initial sample x0, a series of intermediate samples x are generated. 1:T The joint probability distribution of x 1:T represents a series of intermediate samples, q(x t |x t-1 ) means that at a given x t-1 Under the conditions, generate x t The probability distribution of x t-1 represents the sample at time step t-1, β t represents a hyperparameter, representing the amplitude of added noise, T represents the total number of diffusion steps, this process gradually adds noise from the original data to a set of pure noise, and β1<β2···<β T , indicating that the added noise is getting larger and larger, and I represents a matrix with the same shape as the input data.

5. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 3 is characterized in that: In the reverse process, the true distribution of the data is gradually restored by denoising, and a neural network with parameters is used to estimate the conditional probability. The calculation formula is as follows: Among them, θ represents the learnable parameters of the neural network, μ θ and Σ θ represents the mean and variance; u θ (x t ,t) means that at time step t, the model is based on the current noisy sample x t The estimated amount of denoising, α t represents the scaling factor, x t-1 represents the sample at the previous moment, z θ represents the noise at time step t obtained from the prediction.

6. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 3 is characterized in that: In the noise prediction, an improved Unet network architecture is used to predict the noise distribution in the reverse process, specifically including: The multi-head attention mechanism is introduced in the encoder and decoder so that the model can capture long-range dependencies in the signal. The multi-head attention mechanism module includes one-dimensional convolution, batch normalization and multi-head attention mechanism. The mathematical formula is as follows: Among them, Q j represents the query matrix of the jth attention head, x represents the original signal point, K j represents the key matrix of the j-th attention head, V j represents the value matrix of the j-th attention head, and Represents a trainable projection matrix, Attention(Q,K,V) j represents the attention mechanism, represents the transpose of the key matrix of the j-th attention head, d k Represents the dimension of the key vector, MSA(Q,K,V) j Represents the multi-head attention mechanism, head j represents the output of the jth attention head, head j =Attention(Q,K,V) j ,j=1,2,...,H, H represents the number of attention heads, Represents a trainable projection matrix.

7. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 3 is characterized in that: During the diffusion model training process, the noise prediction network based on U-Net needs to be constrained to calculate the loss value between the predicted noise and the actual noise. The calculation formula is as follows: in, represents the expected operation, z t represents the true noise added to the data at time step t, represents the cumulative noise coefficient at time step t.

8. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 1 is characterized in that: Step S3 specifically includes: S31. Construct a shared parameter feature extraction network with three types of inputs: source domain samples, target domain samples, and target domain generated samples. The network structure includes multiple convolutional layers, pooling layers, and normalization layers, and is ultimately mapped to the fault category space. S32. After feature extraction, all samples are sent to the classifier, and the prediction results are output and used for subsequent joint loss function calculation.

9. The small sample cross-domain fault diagnosis method based on multi-level feature alignment according to claim 1 is characterized in that: Step S4 specifically includes: S41. To ensure that the model accurately identifies the category features in the source domain labeled samples and learns the fault difference features in the original data, the cross entropy classification loss function is introduced. The mathematical formula is as follows: Among them, n s represents the number of samples in the source domain, x i represents the i-th source domain sample, Represents x i The test tag, Represents x i The true label of S42. Based on the traditional MMD alignment strategy, it is extended to a dual alignment goal: on the one hand, aligning the overall feature distribution between the source domain samples and the real target domain samples; on the other hand, aligning the distribution between the source domain samples and the target samples generated by the diffusion model to ensure that the generated data can truly reflect the characteristics of the target domain. The specific mathematical expression is: Among them, φ(·) represents mapping the input features into the reproducing kernel Hilbert space, represents the source domain sample, represents the target domain sample, represents the target domain generated sample, n s Indicates the number of source domain samples, n t Indicates the number of target domain samples, n t ′ represents the number of generated target domain samples; S43. In order to enhance the intra-class compactness and inter-class separability of cross-domain samples in the feature space, a cross-domain triplet loss function is designed. The mathematical expression is as follows: Among them, the anchor sample a, positive sample p and negative sample n come from the source domain sample or target domain samples Positive samples and anchor samples belong to the same category, while negative samples come from different categories. Function f(·) represents the feature extraction representation of the sample, f(a) represents the feature representation of the anchor, f(p) represents the feature representation of the positive sample, f(n) represents the feature representation of the negative sample, and δ represents the boundary interval hyperparameter, which is used to control the minimum distinction between positive and negative samples. S44. A consistency regularization loss mechanism is introduced to constrain the prediction consistency of the original sample in the target domain and its enhanced sample in the feature extraction results. The expression of the consistency regularization loss function is as follows: in, Represents the original sample of the target domain Enhanced samples of S45. A dynamic weight adjustment parameter α is introduced in the optimization process to gradually control the influence of the loss term corresponding to the pseudo-label sample. The adjustment function of the parameter α is defined as follows: Where i represents the current training round, E represents the total training rounds, and γ represents the growth rate control factor, which is set to 10. S46. The cross-domain diagnosis network achieves efficient feature alignment and improved discrimination performance under the condition of limited target domain data through the combined optimization of multiple loss functions. The overall loss function expression is as follows: Among them, λ and α represent the weight adjustment parameters of the combination of MMD loss term, triple loss and consistency regularization loss term, respectively. represents the cross entropy classification loss function, represents the dual alignment objective loss function, represents the cross-domain triplet loss function, represents the consistency regularization loss function.

Citation Information

Patent Citations

  • Small sample generation diffusion model multi-target neural architecture search method based on asynchronous gradient descent distributed parallel strategy

    CN117391147A

  • Gearbox cross-domain fault diagnosis method based on branch attention contrast transfer learning

    CN118758594A

  • Simulation data driven rolling bearing fault diagnosis method

    CN119509971A

  • Semi-supervised intelligent fault diagnosis method based on de-noising diffusion probability model and medium

    CN119848668A

  • Cross-domain mechanical fault diagnosis method based on multi-channel feature fusion of CBAM and use thereof

    US20240142342A1

Cited By

  • Power distribution network fault identification method combining deep learning and attention

    CN121388818A

  • Cross-tax-category finance and taxation inspection method, equipment and medium

    CN121937242A

  • A cross-tax financial and tax inspection method, device and medium

    CN121937242B

  • Variable working condition fault diagnosis method and device based on adaptive anchor point graph attention

    CN122310243B