Diffusion Schrodinger bridge and sample purification driven generalized zero sample fault diagnosis method

By generating multi-distribution samples through the conditional diffusion Schrödinger bridge model with adaptive group normalization, and combining boundary refinement and domain separation mechanisms, the problem of misjudgment of unknown faults in generalized zero-sample fault diagnosis is solved, achieving high-precision differentiation between known and unknown fault categories and improving diagnostic accuracy.

CN122020344APending Publication Date: 2026-05-12CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2026-01-27
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing generalized zero-sample fault diagnosis methods only use visible category samples to construct decision boundaries, which leads to unknown faults being misclassified as known categories and cannot effectively distinguish between known and unknown category faults.

Method used

The adaptive group normalization conditional diffusion Schrödinger bridge (AGCDSB) model is used to generate multi-distribution samples. A high-resolution decision boundary is constructed through the boundary refinement and purification mechanism (BRP) and the domain separation mechanism. Combined with the multi-distribution label-guided generation strategy (MDLG) and the local anomaly factor algorithm (LOF), the accurate identification of known and unknown categories of faults is achieved.

Benefits of technology

It improves the accuracy of fault diagnosis, effectively distinguishes between known and unknown fault categories, enhances the quality and diversity of generated samples, reduces the possibility of unknown samples being misclassified as known categories, and achieves high-precision generalized zero-sample composite fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020344A_ABST
    Figure CN122020344A_ABST
Patent Text Reader

Abstract

The invention relates to a generalized zero sample fault diagnosis method driven by a diffused Schrodinger bridge and sample purification, and belongs to the field of rotating machinery fault diagnosis. The method comprises the following steps: S1, collecting fault data; s2, carrying out signal analysis on the visible classes to obtain labels and distribution information, and constructing corresponding semantic prototypes at the same time; s3, the labels and the distribution information are input into a designed AGCDSB model, and unseen classes of different distributions are generated; s4, BRP is used for purifying the unseen species; s5, dividing the test samples into visible and non-visible classes through a domain separation mechanism; s6, using a visible class classifier to identify the fault which is predicted to be visible class; s7, aligning the fault semantics with the visible features; s8, extracting unseen class fault features, and correcting an unseen class fault prototype in combination with a prototype clustering matching technology; and S9, inferring by using nearest neighbor estimation and a correction prototype to obtain unseen class labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rotating machinery fault diagnosis and relates to a generalized zero-sample fault diagnosis method driven by diffusion Schrödinger bridge and sample purification. Background Technology

[0002] The increasing intelligence and complexity of industrial equipment is profoundly changing the operation and maintenance models in manufacturing, energy, and transportation sectors. Ensuring the safe and reliable operation of rotating machinery is a crucial aspect of industrial equipment operation and maintenance. As core components of rotating machinery, bearings and gears bear complex loads and harsh operating conditions, resulting in a high failure rate. With increasingly complex equipment structures and more diverse operating conditions, compound faults are also frequently occurring. Compound faults are not simply the superposition of single faults, thus posing a significant challenge under the assumptions of traditional single fault diagnosis. Therefore, developing high-precision compound fault diagnosis algorithms is of great importance for ensuring efficient equipment operation and reducing maintenance costs.

[0003] Traditional composite fault diagnosis algorithms are mainly divided into two categories. The first category is based on signal processing methods, aiming to decouple complex fault features into more easily identifiable single features. For example, Yi et al. achieved composite fault diagnosis by optimizing variational mode decomposition algorithms, and Ding et al. developed wavelet transform techniques suitable for wheelset-bearing composite fault identification. Although these methods have some effectiveness, they are highly dependent on expert knowledge and often face difficulties in decoupling complex signals. The second category employs traditional artificial intelligence techniques, which can achieve high diagnostic accuracy but require a large amount of data for model training. For example, Zhu et al. designed an advanced capsule network for composite fault decoupling through a fusion mechanism, while Wang et al. used multiple extreme learning machines for composite fault classification. Traditional AI-based fault diagnosis methods face challenges due to their reliance on large amounts of data. To address this, zero-shot learning methods have emerged, which can identify unseen faults using only seen category samples and semantic attributes. For example, Yang et al. proposed a zero-shot attribute description model for circuit breaker fault diagnosis, while Jiang et al. designed a zero-shot identification framework for chiller units using text attributes. However, the aforementioned zero-shot methods can only identify faults of unknown categories and cannot distinguish between known categories. In real-world scenarios, it is crucial to simultaneously identify both known and unknown category samples. Therefore, generalized zero-shot learning methods have been proposed, and related research is gradually unfolding.

[0004] Existing generalized zero-shot fault diagnosis methods only utilize visible class samples to construct decision boundaries, which may lead to semantic bias problems, i.e., unknown faults are misclassified as known categories. Therefore, it is particularly important to study a fault diagnosis method with higher accuracy. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a generalized zero-sample fault diagnosis method driven by a diffusion Schrödinger bridge and sample purification, which enables the identification of both single and compound faults using only a single fault sample during the training phase. To achieve this goal, the specific objectives of this invention include: 1) designing a conditionally controlled diffusion Schrödinger bridge model to control the labels of generated samples; 2) providing a high-resolution decision boundary construction method based on multi-distribution sample generation to effectively distinguish between known and unknown fault categories; and 3) constructing a boundary refinement and purification mechanism to optimize the decision boundary and filter out interference from invalid generated samples. Through the integration of the above technologies, high-precision generalized zero-sample compound fault diagnosis is ultimately achieved.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A generalized zero-sample fault diagnosis method driven by diffusion Schrödinger bridge and sample purification includes the following steps: S1: Collect fault data and divide the samples; S2: Perform signal analysis on visible fault samples to obtain labels and distribution information, and construct corresponding semantic prototypes. S3: Input the label and distribution information into the designed AGCDSB model to generate unseen fault samples with different distributions; where AGCDSB represents the adaptive group normalized conditional diffusion Schrödinger bridge; S4: Use the BRP strategy to purify unseen fault samples; where BRP stands for boundary purification. S5: Through the domain separation mechanism, the test samples are divided into visible fault samples and unseen fault samples; S6: Use a visible class classifier to identify samples predicted as visible class faults; S7: Align fault semantics with features of visible fault samples; S8: Extract features of unseen fault samples and correct the prototypes of unseen fault samples by combining prototype clustering matching technology; S9: Use nearest neighbor estimation and modified prototype inference to obtain the labels of unseen fault samples.

[0008] Furthermore, in step S3, the AGCDSB model achieves optimal transfer between the known data distribution and the given prior distribution under conditional control. Specifically, it includes: the diffusion Schrödinger bridge, i.e., DSB, decomposes the solution process into two steps, simplifying the complex joint distribution solution into a conditional distribution solution. (1) (2) The process described by formulas (1) and (2) transforms the complex path optimization problem into optimizing adjacent time steps. and The problem of transition probability; where, Optimize the path for odd-numbered steps. It is a reverse conditional transition distribution. For even-numbered steps, the conditional transition distribution estimate is... For the joint probability distribution path, It is the set of probability distribution paths in the state space from 0 to N steps. For the final distribution, For the prior distribution, Optimize the path for even-numbered steps. It is a positive conditional transition distribution. For the conditional transition distribution estimate of odd-numbered steps, For the initial distribution, For data distribution, Let KL divergence be a metric. Furthermore, the DSB model employs a method similar to the diffusion model, assuming that the transition probabilities follow a Gaussian distribution; its forward and backward processes can be represented as follows: (3) (4) in, Indicates step size, and Represents the offset item. and These are the joint densities of the forward and backward processes, respectively. Indicates the forward transition probability. Indicates the backward transition probability; The state variable represents the time step t. Represents the identity matrix. Indicates a Gaussian distribution; In actual training, AGCDSB uses two neural networks for learning and performs stepwise optimization: the forward network optimizes in even-numbered steps, and the backpropagation network optimizes in odd-numbered steps. (5) (6) in, These are the network parameters for forward and backward propagation. It is a feedforward neural network. For backward update function, For input parameters, It is a feedforward neural network. This is the forward update function; The loss function of the traditional diffusion Schrödinger bridge model is: (7) (8) However, since optimizing two networks simultaneously requires parameters from both networks for each loss calculation, the training process is extremely difficult. Therefore, AGCDSB employs a simplified loss function, which, under certain approximations, is equivalent to the training objective of DSB: (9) (10) in, For the backward network loss function, Forward network loss function, The mathematical expectation of the joint distribution of the forward process. This represents the mathematical expectation of the joint distribution of the backward process. The loss function described above calculates the network parameters only once for each training value, effectively reducing computation time.

[0009] The basic DSB model can only generate samples and is not easy to control conditions. To achieve controlled generation while maintaining training efficiency, the AGCDSB model employs a lightweight network that integrates an Adaptive Group Normalization (AdaGN) layer. In this network, conditional labels and prior distribution information are embedded through the AdaGN layer; the output of the AdaGN layer can be written as: (11) (12) in, c For embedded prior information; , These are the mean and standard deviation, respectively. , These are the trainable scale and shift parameters, respectively; x , y These are the input and output of the network layer, respectively.

[0010] Furthermore, in step S4, the BRP strategy purifies the sample by introducing LOF anomaly detection technology; when using the LOF algorithm to determine outliers, it is first necessary to calculate the first... k One reachable distance: (13) in, Let p be the k-th reachable distance from point p to point o. Let k be the k-th distance from point o. Let be the distance between points p and o; No. k Local Accessibility Density It can be calculated as: (14) Among them, point p of k Distance to Neighborhood Represented as: (15) in, For the sample set; Finally, the Local Outlier Factor (LOF) for each point is calculated; it is the ratio of the average local reachability density of all points in the k-distance neighborhood of point p to the local reachability density of that point itself. (16) in, For point p The local outlier factor; the magnitude of the outlier factor reflects the degree of abnormality of the sample.

[0011] Ultimately, having the largest front n Data points with a single LOF value were identified as data in the generated unknown domain.

[0012] Furthermore, in step S5, the domain separation mechanism specifically includes: first, treating the fault samples generated in step S3 as unknown categories and combining them with fault samples of known categories for training; then, employing a wide-mixed dilated convolutional neural network as a domain classifier; and using binary cross-entropy loss as the loss function for domain separation. (17) in, It is the loss function for domain separation. and These are real and predicted domain labels, respectively; The Youden index is used to determine the optimal classification threshold. (18) in, , , and These represent the number of true positives, false positives, true negatives, and false negatives, respectively; if the sample's predicted score is greater than the threshold... τ If the condition is met, the sample is determined to belong to an unknown category; otherwise, it is determined to belong to a known category.

[0013] Furthermore, in step S6, the visible class classifier adopts the same network architecture as the domain classifier and is trained using cross-entropy loss; samples identified as coming from a known domain during the domain separation stage will then be classified by this module to obtain their specific fault category labels.

[0014] Furthermore, in step S7, semantic alignment loss This can be expressed as: (19) in, , and These represent the extracted individual fault features, their corresponding semantic prototypes, and the number of training samples, respectively.

[0015] Furthermore, in step S8, the prototype clustering matching technique specifically includes: introducing unseen class features through Gaussian mixture clustering algorithm to further refine the semantic prototype constructed based on prior knowledge and data; its mean vector The mean vector, used to correct the prototype, is updated as follows: (20) in, This represents the posterior probability of the Gaussian mixture model. This indicates the number of Gaussian mixture models.

[0016] Furthermore, in step S9, the modified prototype can be described as follows: (twenty one) in, Indicates the first i A revised prototype Indicates the first i An initial semantic prototype, Indicates the first i A reconstructed prototype, Indicates the correction factor; Nearest neighbor estimation is used to establish the relationship between the modified prototype and the extracted features, thereby obtaining the corresponding unseen class labels: (twenty two) in, This indicates that no class tag was found. This indicates the number of categories with no observed samples. Dimensions representing semantic attributes Indicates the first i The first of the revised prototypes k One element, This represents the extracted semantic attribute features.

[0017] The beneficial effects of this invention are as follows: 1) The AGCDSB model proposed in this invention can control the labels of generated samples. This model compresses the decision boundary of the known domain by generating samples of unknown categories, thereby enhancing the fault differentiation capability between the two domains. High-precision diagnosis is achieved by combining traditional diagnostic methods with zero-shot diagnostic technology. This model effectively satisfies the optimal transport theory and can generate samples from different distributions, thus improving the quality and diversity of generated samples.

[0018] 2) The boundary purification mechanism proposed in this invention introduces the Local Anomaly Factor (LOF) algorithm to remove abnormal samples, thereby assisting in optimizing the decision boundary and effectively distinguishing known and unknown fault characteristics.

[0019] 3) This invention constructs a high-resolution decision boundary construction method based on multi-distribution sample generation, which can effectively distinguish between known and unknown categories of faults.

[0020] 4) The method of this invention proposes a multi-distribution label-guided generation strategy (MDLG) in the ASAP model. This strategy generates samples with multi-label information by fusing multiple Gaussian distributions and applying control conditions, thereby enhancing the diversity of generated samples. The ASAP model demonstrates superior performance in part-level and component-level experiments compared to six classic and state-of-the-art zero-shot and generalized zero-shot models. It can diagnose single and compound faults using only a single fault.

[0021] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 The diagram illustrates the comparison of four methods for constructing decision boundaries for visible and invisible domains, where (a) uses only the visible class; (b) uses the visible class and single-distributed invisible class samples; (c) uses the visible class and multi-distributed invisible class samples; and (d) uses the visible class and purified multi-distributed invisible class samples. Figure 2 This is a schematic diagram of the diagnostic process of the diffusion Schrödinger bridge and sample purification-driven generalized zero-sample method of the present invention; Figure 3 This is a diagram showing the main structure of the AGCDSB model; Figure 4Here is the confusion matrix of the ASAP model under four conditions in the part-level experiment, where (a) is condition 0; (b) is condition 1; (c) is condition 2; and (d) is condition 3. Figure 5 Here is the confusion matrix of the ASAP model under four component-level experimental conditions, where (a) is condition 0; (b) is condition 1; (c) is condition 2; and (d) is condition 3. Figure 6 To assess the diagnostic accuracy of three different methods in two experiments under eight different operating conditions, the component-level experiments were conducted as follows: (a) Operating condition 0; (b) Operating condition 1; (c) Operating condition 2; (d) Operating condition 3; and the part-level experiments were conducted as follows: (e) Operating condition 0; (f) Operating condition 1; (g) Operating condition 2; (h) Operating condition 3. Figure 7 The diagnostic process and performance validation of the ASAP model are shown in the following figures: (a) T-SNE visualization of the original samples; (b) classification results of visible and invisible categories; (c) classification results of all categories; (d) visible class samples and generated single-distribution samples; (e) visible class samples and generated multi-distribution samples; and (f) visible class samples and purified generated multi-distribution samples. Detailed Implementation

[0023] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0024] Existing generalized zero-shot fault diagnosis methods only utilize visible class samples to construct decision boundaries, which may lead to semantic bias problems, i.e., unknown faults are misclassified as known categories. To address this problem in generalized zero-shot learning, this invention proposes an adaptive group normalization conditional diffusion Schrödinger bridge (AGCDSB) driven sample augmentation and sanitization (ASAP) framework. This framework compresses the decision boundary of the known domain by generating unknown class samples, thereby enhancing the fault discrimination capability between the two domains. By combining traditional diagnostic methods with zero-shot diagnostic techniques, high-precision diagnosis is achieved.

[0025] However, regarding the construction of decision boundaries, such as... Figure 1As shown in (a), traditional methods rely solely on samples of known classes, resulting in loose and uncertain decision boundaries, which can lead to misclassification due to unknown faults. While generated samples can compress the boundary, samples generated from a single distribution or with a single label are difficult to compress accurately, such as... Figure 1 As shown in (b). Therefore, this invention employs a multi-distribution label-guided generation strategy to enhance the diversity of generated samples. Furthermore, some generated samples may overlap with the original samples, leading to misclassification, such as... Figure 1 As shown in (c), a boundary refining and purification mechanism is introduced to eliminate interference from such overlapping samples. Finally, as... Figure 1 As shown in (d), by integrating multi-distribution label-guided generation and boundary purification mechanisms, a high-resolution decision boundary can be constructed to achieve accurate diagnosis of known and unknown faults.

[0026] To construct high-resolution decision boundaries, this invention improves upon both the generation model and the decision boundary construction method. First, traditional generation methods (such as diffusion models) do not satisfy optimal transport theory and cannot generate samples from different distributions, affecting the quality and diversity of generated samples. Therefore, this invention proposes the AGCDSB model. This model effectively satisfies optimal transport theory and can generate samples from different distributions, thus improving the quality and diversity of generated samples. Furthermore, we introduce a conditional control mechanism into the traditional diffusion Schrödinger bridge to control the generation of diverse samples. Second, regarding decision boundary construction, samples generated from a single distribution and with a single label have homogeneous structures, easily leading to inaccurate decision boundaries. Therefore, this invention proposes a multi-distribution label-guided generation strategy (MDLG) within the ASAP framework. This strategy generates samples with multi-label information by fusing multiple distributions and applying control conditions, thereby enhancing the diversity of generated samples. Given that generated pseudo-unknown category samples may overlap with normal samples, this invention constructs a boundary refinement and purification mechanism, introducing a Local Anomaly Factor (LOF) algorithm to remove anomalous samples, thereby assisting in optimizing the decision boundary. The main innovations of this scheme can be summarized as follows: 1) To alleviate the semantic bias problem in generalized zero-shot fault diagnosis, a high-resolution decision boundary construction method is proposed. This method generates diverse unknown category samples through a multi-distribution label-guided generation strategy to compress the decision boundary and reduce the occurrence of unknown samples being misclassified as known categories. In addition, a boundary purification mechanism is constructed to optimize the decision boundary and reduce the interference of invalid generated samples, thereby improving the overall diagnostic accuracy.

[0027] 2) To improve the quality of generated samples, an adaptive group normalization conditional diffusion Schrödinger bridge model is proposed. This model satisfies optimal transport theory and can generate data from different distributions, thereby improving the quality and diversity of generated samples. By designing an adaptive group normalization module and introducing distribution and label information, effective conditional control of the generated signal is achieved.

[0028] 3) Utilizing the developed multi-distribution label-guided generation, boundary refinement and purification, and adaptive group normalization conditional diffusion Schrödinger bridge modules, combined with domain separation, known category identification, and zero-shot classification modules, a sample augmentation and purification model for generalized zero-shot fault diagnosis was constructed. This model was validated through a series of generalized zero-shot fault diagnosis experiments conducted at the component and part levels.

[0029] This embodiment provides a generalized zero-sample diagnostic method (ASAP) driven by diffusion Schrödinger bridge and sample purification, such as... Figure 2 As shown, the specific steps include: S1: Collect fault data; S2: Perform signal analysis on visible samples to obtain labels and distribution information, and construct corresponding semantic prototypes; S3: Input the label and distribution information into the designed AGCDSB model to generate unseen class samples with different distributions; S4: Use the BRP mechanism to purify unseen class samples; S5: Through the domain separation mechanism, test samples are divided into visible and invisible classes; S6: Use a visible class classifier to identify faults predicted to be visible; S7: Align fault semantics with visible class features; S8: Extract the features of unseen fault types and combine them with prototype clustering matching technology to correct the prototypes of unseen fault types; S9: Use nearest neighbor estimation and modified prototype inference to obtain the unseen class label.

[0030] In step S3, the main structure of the AGCDSB model is as follows: Figure 3 As shown. The core objective of this model is to achieve optimal transfer between a known data distribution and a given prior distribution under conditional control. The diffusion Schrödinger bridge decomposes the solution process into two steps, simplifying the complex joint distribution solution into a conditional distribution solution: (1) (2) The process described by formulas (1) and (2) transforms the complex path optimization problem into optimizing adjacent time steps. and The problem of transition probability.

[0031] Furthermore, the DSB model employs a method similar to the diffusion model, assuming that the transition probabilities follow a Gaussian distribution. Its forward and backward processes can be represented as follows: (3) (4) in, Indicates step size, and Represents the offset item. and These are the joint densities of the forward and backward processes, respectively.

[0032] In actual training, AGCDSB uses two neural networks for learning and performs stepwise optimization: the forward network optimizes in even-numbered steps, and the backpropagation network optimizes in odd-numbered steps. (5) (6) in, These are the network parameters for forward and backward propagation.

[0033] The loss function of the traditional diffusion Schrödinger bridge model is: (7) (8) However, since optimizing two networks simultaneously requires parameters from both networks for each loss calculation, the training process is extremely difficult. Therefore, AGCDSB employs a simplified loss function, which, under certain approximations, is equivalent to the training objective of DSB: (9) (10) The loss function described above calculates the network parameters only once for each training value, effectively reducing computation time.

[0034] The basic DSB model can only generate samples and is not easy to control conditions. To achieve controlled generation while maintaining training efficiency, the ASAP model employs a lightweight network that integrates an Adaptive Group Normalization (AdaGN) layer. A schematic diagram of this structure is shown below. Figure 3 The upper part is shown. In this network, conditional labels and prior distribution information are embedded through AdaGN layers. The output of the AdaGN layers can be written as... (11) (12) in, cFor embedded prior information; , These are the mean and standard deviation, respectively. , These are the trainable scale and shift parameters, respectively.

[0035] In step S4, the BRP strategy is used to purify the samples, i.e., to remove samples that overlap with known classes. In this strategy, we introduce LOF anomaly detection technology to purify the samples. Unlike traditional anomaly monitoring, our strategy eliminates the interference of non-anomaly samples on the construction of the decision boundary.

[0036] When using the LOF algorithm to identify outliers, the first step is to calculate the k-th reachable distance for each point within the input neighborhood: (13) No. k Local accessibility density can be calculated as: (14) Among them, point p of k Distance neighborhood is represented as: (15) Finally, the Local Outlier Factor (LOF) is calculated for each point. It is the ratio of the average local reachability density of all points within the k-distance neighborhood of point p to the local reachability density of the point itself. (16) The magnitude of the outlier factor reflects the degree of abnormality in the sample. Ultimately, the sample with the largest outlier factor... n Data points with a single LOF value were identified as data in the generated unknown domain.

[0037] In step S5, the domain separation mechanism is key to mitigating the semantic shift problem in generalized zero-shot learning. This technique effectively prevents unseen faults from being misclassified as seen faults by separating samples from seen and unseen domains. First, the samples generated by method S3 are treated as unknown categories and combined with samples from known categories for training. Then, a wide-mixed dilated convolutional neural network is used as the domain classifier. Binary cross-entropy loss is used as the loss function for domain separation. (17) in, and These are real and predicted domain labels, respectively.

[0038] The Youden index is used to determine the optimal classification threshold. (18) in, , , and These represent the number of true positives, false positives, true negatives, and false negatives, respectively; if the sample's predicted score is greater than the threshold... τ If the condition is met, the sample is determined to belong to an unknown category; otherwise, it is determined to belong to a known category.

[0039] In step S6, it can be seen that the known class classifier adopts the same network architecture as the domain classifier and is trained using cross-entropy loss. Samples identified as coming from the known domain during the domain separation stage will then be classified by this module to obtain their specific fault category labels.

[0040] In step S7, the semantic alignment loss can be expressed as: (19) in, , and These represent the extracted individual fault features, their corresponding semantic prototypes, and the number of training samples, respectively.

[0041] In step S8, this invention introduces unseen class features using a Gaussian mixture clustering algorithm to further refine the semantic prototype constructed based on prior knowledge and data. Its mean vector... The mean vector, used to correct the prototype, is updated as follows: (20) In step S9, the modified prototype can be described as follows: (twenty one) Nearest neighbor estimation is used to establish the relationship between the modified prototype and the extracted features, thereby obtaining the corresponding unseen class labels: (twenty two) Verification experiment: This invention validates the diagnostic performance of the proposed ASAP method through component-level and component-level composite fault diagnosis experiments. Component-level experiments were conducted on a self-built bearing composite fault experimental platform. Component-level experiments used the BJTU-RAO dataset. Ablation experiments verified the effectiveness of the proposed multi-distribution label-guided generation strategy and boundary refinement and purification mechanism. Furthermore, the classification process is illustrated in detail through visualization analysis, and the role of the proposed ASAP method in constructing the decision boundary is clarified.

[0042] Component-level experiments evaluated seven bearing conditions: normal condition, outer ring failure, inner ring failure, ball failure, combined outer and inner ring failure, combined ball and outer ring failure, and combined ball and inner ring failure. The rotational speed was set to 500 rpm, and different torque loads were applied to the bearing: 0 Nm (condition 0), 2 Nm (condition 1), 4 Nm (condition 2), and 6 Nm (condition 3). In the component-level experiments, the data was divided into training and test sets in an 8:2 ratio. A sliding window was used for splitting to increase the sample size.

[0043] Six benchmark zero-shot learning models were selected for comparative analysis: DAP, FDAT, SCE, ZSAECFD, AWSAE, and CGASNet. These methods maintain consistency with the proposed model in terms of data partitioning.

[0044] To compare the diagnostic performance of various methods in detail, this experiment evaluated three different diagnostic accuracies: known class accuracy (S), unknown class accuracy (U), and overall accuracy (A). Table 1 summarizes the results of various zero-shot learning methods on the BCF dataset. The proposed ASAP method achieved diagnostic accuracies ranging from 84.8% to 96.0% across four scenarios. When trained using only known class samples, the proposed ASAP method achieved significant overall diagnostic accuracy on both known and unknown classes, effectively demonstrating the great potential of the ASAP method. Compared to the six baseline methods, ASAP achieved the highest accuracy on unknown classes while maintaining high recognition accuracy on known classes, and also achieved the highest overall accuracy. Methods such as DAP, FDAT, SCE, and ZSAECFD are primarily designed for zero-shot learning tasks and therefore do not consider the semantic shift problem in generalized zero-shot learning. This leads to most samples being classified as known classes, resulting in lower overall diagnostic accuracy. Furthermore, FDAT and SCE are trained using traditional machine learning methods, resulting in the lowest overall diagnostic accuracy. AWSAE and CGASNet attempt to alleviate the semantic shift problem in generalized zero-shot learning, achieving some accuracy improvements. However, due to their limited ability to distinguish between known and unknown categories, neither method can match the superior diagnostic performance of ASAP. To further illustrate the performance of each category, Figure 4 Confusion matrices corresponding to the four operating conditions are presented. These matrices confirm that ASAP can effectively separate known and unknown categories, thereby achieving high diagnostic accuracy across all fault categories.

[0045] Table 1. Average diagnostic accuracy of the GZSFD method under four different conditions in part-level experiments.

[0046] Furthermore, the diagnostic performance of the proposed ASAP method was further explored through component-level experiments, which are closer to practical engineering applications. The component-level experiments used six health states of gears and bearings: normal state, tooth root crack, broken tooth, bearing inner ring failure, a combined fault of tooth root crack and inner ring failure, and a combined fault of broken tooth and inner ring failure. Similar to the component-level experiments, four different experimental conditions were considered: Condition 0: 20 Hz speed, lateral load +10 kN; Condition 1: 40 Hz speed, lateral load +10 kN; Condition 2: 60 Hz speed, lateral load +10 kN; Condition 3: 20 Hz speed, lateral load -10 kN. The sampling frequency of the vibration experimental data was set to 64 kHz. The component-level experiments used the same hyperparameters, data partitioning methods, and comparison methods as the component-level experiments.

[0047] Table 2 lists the diagnostic results of the proposed ASAP model and six different zero-shot methods under four experimental conditions. The overall diagnostic accuracy of the proposed ASAP method ranges from 84.6% to 92.6% under the four conditions, showing a significant improvement compared to the six comparison methods. Overall, ASAP demonstrates acceptable diagnostic results. The DAP and ZSAECFD methods exhibit extremely high diagnostic accuracy for known categories, reaching 99.9% under certain conditions. However, due to the lack of a mechanism to separate unknown samples, their classification accuracy for unknown categories is extremely low, resulting in relatively low overall accuracy. The AWSAE and CGASNet methods show better overall recognition accuracy in component-level experiments than in component-level experiments, but due to their poor classification performance in known and unknown domains, they still do not achieve very high overall accuracy. The confusion matrix of the component-level experiments is shown below. Figure 5 As shown, the experimental results for each category verify the diagnostic performance of the ASAP method proposed in this invention.

[0048] Table 2. Average diagnostic accuracy of the GZSFD method under four different conditions in component-level experiments.

[0049] Ablation experiments were conducted to further verify the effectiveness of the proposed MDLG mechanism and BRP strategy. Figure 6 The diagnostic accuracy of three different models was presented under eight operating conditions across two experimental categories. The ASAP model without the MDLG mechanism is referred to as ASAP w / o MDLG, and the model without the BRP mechanism is referred to as ASAP w / o BRP. Overall, both the MDLG and BRP modules contribute to improving diagnostic accuracy. However, the effect of MDLG is more pronounced compared to other modules. The performance degradation of ASAP w / o MDLG is more significant compared to ASAP w / o BRP. Figure 6As can be clearly seen in Figure (h), adding this module significantly improves the accuracy of identifying unknown categories, indicating that generating multi-type labels from multi-distribution data helps improve diagnostic accuracy. Regarding the accuracy of identifying known and unknown categories, both modules improve the accuracy of identifying both known and unknown categories by helping to construct a better decision boundary, thereby avoiding confusion between the two fault types. Overall, with the addition of the MDLG and BRP modules, diagnostic variability is reduced, further validating the superior performance of the proposed modules.

[0050] To illustrate the diagnostic process and performance of the proposed ASAP model, Figure 7 Figures (a)–(c) visualize the original data, the classification results for known and unknown categories, and the final classification result. Figure 7 Figures (b) and (c) show that the proposed ASAP can effectively distinguish between samples from known and unknown domains, thereby achieving high-precision identification of different types of faults. Furthermore, Figure 7 Figures (d) to (f) visualize the role of generated samples in constructing decision boundaries. Decision boundaries constructed from samples generated from a single distribution are limited, which restricts their classification performance; while samples generated from multiple distributions can construct decision boundaries from multiple perspectives, effectively improving the discriminative power of the decision boundaries. However, generated samples may overlap with normal samples, affecting the accurate identification of some samples. Figure 7 Figure (f) shows the decision boundary optimized by the BRP mechanism. By purifying the generated non-abnormal samples, the probability of unknown category samples being classified as known categories is reduced, thereby further improving diagnostic accuracy.

[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A generalized zero-sample fault diagnosis method driven by diffusion Schrödinger bridge and sample purification, characterized in that, The method specifically includes the following steps: S1: Collect fault data and divide the samples; S2: Perform signal analysis on visible fault samples to obtain labels and distribution information, and construct corresponding semantic prototypes. S3: Input the label and distribution information into the designed AGCDSB model to generate unseen fault samples with different distributions; where AGCDSB represents the adaptive group normalized conditional diffusion Schrödinger bridge; S4: Use the BRP strategy to purify the generated unseen fault samples; where BRP stands for boundary purification. S5: Through the domain separation mechanism, the test samples are divided into visible fault samples and unseen fault samples; S6: Use a visible class classifier to identify samples predicted as visible class faults; S7: Align fault semantics with features of visible fault samples; S8: Extract features of unseen fault samples and correct the prototypes of unseen fault samples by combining prototype clustering matching technology; S9: Use nearest neighbor estimation and modified prototype inference to obtain the labels of unseen fault samples.

2. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S3, the AGCDSB model achieves optimal transfer between the known data distribution and the given prior distribution under conditional control. Specifically, it includes: the diffusion Schrödinger bridge, i.e., DSB, decomposes the solution process into two steps, simplifying the complex joint distribution solution into a conditional distribution solution. in, Optimize the path for odd-numbered steps. It is a reverse conditional transition distribution. For even-numbered steps, the conditional transition distribution estimate is given. For the joint probability distribution path, It is the set of probability distribution paths in the state space from 0 to N steps. For the final distribution, As a prior distribution, Optimize the path for even-numbered steps. It is a positive conditional transition distribution. For the conditional transition distribution estimate of odd-numbered steps, For the initial distribution, For data distribution, Let KL divergence be a metric. Furthermore, the DSB model employs a method similar to the diffusion model, assuming that the transition probabilities follow a Gaussian distribution; its forward and backward processes are represented as follows: in, Indicates the step size. and Represents the offset item. and These are the joint densities of the forward and backward processes, respectively. Indicates the forward transition probability. Indicates the backward transition probability; The state variable represents the state at time step t. Represents the identity matrix. Indicates a Gaussian distribution; During training, AGCDSB uses two neural networks for learning and performs stepwise optimization: the forward network optimizes in even-numbered steps, and the backpropagation network optimizes in odd-numbered steps. in, These are the network parameters for forward and backward propagation. It is a feedforward neural network. For backward update function, For input parameters, It is a feedforward neural network. This is the forward update function; AGCDSB uses a simplified loss function, which, under approximate conditions, is equivalent to the training objective of DSB: in, For the backward network loss function, Forward network loss function, The mathematical expectation of the joint distribution of the forward process. Let be the mathematical expectation of the joint distribution of the backward process.

3. The generalized zero-sample fault diagnosis method according to claim 2, characterized in that, In step S3, the AGCDSB model employs a lightweight network that integrates an AdaGN layer, where AdaGN stands for Adaptive Group Normalization. In this network, conditional labels and prior distribution information are embedded through the AdaGN layer. The output of the AdaGN layer is written as: in, c For embedded prior information; , These are the mean and standard deviation, respectively. , These are the trainable scale and shift parameters, respectively; x , y These are the input and output of the network layer, respectively.

4. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S4, the BRP strategy cleanses the samples by introducing LOF anomaly detection technology; when using the LOF algorithm to determine outliers, it is first necessary to calculate the first LOF value of each point in the input neighborhood. k One reachable distance: in, Let p be the k-th reachable distance from point p to point o. Let k be the k-th distance from point o. Let be the distance between points p and o; No. k Local Accessibility Density The calculation is as follows: Among them, point p of k Distance to Neighborhood Represented as: in, For the sample set; Finally, the Local Outlier Factor (LOF) for each point is calculated; it is the ratio of the average local reachability density of all points in the k-distance neighborhood of point p to the local reachability density of that point itself. in, For point p Local outlier; Ultimately, having the largest front n Data points with a single LOF value were identified as data in the generated unknown domain.

5. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S5, the domain separation mechanism specifically includes: First, the fault samples generated in step S3 are treated as unknown categories and combined with fault samples of known categories for training; then, a wide-mixed dilated convolutional neural network is used as the domain classifier; binary cross-entropy loss is used as the loss function for domain separation. in, It is the loss function for domain separation. and These are real and predicted domain labels, respectively; The Youden index is used to determine the optimal classification threshold. in, , , and These represent the number of true positives, false positives, true negatives, and false negatives, respectively; if the sample's predicted score is greater than the threshold... τ If the condition is met, the sample is determined to belong to an unknown category; otherwise, it is determined to belong to a known category.

6. The generalized zero-sample fault diagnosis method according to claim 5, characterized in that, In step S6, the visible class classifier adopts the same network architecture as the domain classifier and is trained using cross-entropy loss; samples identified as coming from a known domain during the domain separation stage will then be classified by this module to obtain their specific fault category labels.

7. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S7, semantic alignment loss The expression is as follows: in, , and These represent the extracted individual fault features, their corresponding semantic prototypes, and the number of training samples, respectively.

8. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S8, the prototype clustering matching technique specifically includes: introducing unseen class features through Gaussian mixture clustering algorithm to further refine the semantic prototype constructed based on prior knowledge and data; its mean vector The mean vector, used to correct the prototype, is updated as follows: in, This represents the posterior probability of the Gaussian mixture model. This indicates the number of Gaussian mixture models.

9. The generalized zero-sample fault diagnosis method according to claim 1, characterized in that, In step S9, the modified prototype is described as follows: in, Indicates the first i A revised prototype Indicates the first i An initial semantic prototype, Indicates the first i A reconstructed prototype, Indicates the correction factor; Nearest neighbor estimation is used to establish the relationship between the modified prototype and the extracted features, thereby obtaining the corresponding unseen class labels: in, This indicates that no class tag was found. This indicates the number of categories with no observed samples. Dimensions representing semantic attributes Indicates the first i The first of the revised prototypes k One element, This represents the extracted semantic attribute features.