A zero-sample fault diagnosis method based on FE-FAGDM
Through the FE-FAGDM-based method, sensitive parameters are selected and combined with diffusion model and semantic guidance to generate high-quality zero-sample fault samples, the accuracy of fault diagnosis under zero-sample conditions is solved and the effectiveness of fault diagnosis is improved.
Patent Information
- Application Number
- CN202510694507.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-28
AI Technical Summary
In industrial equipment fault diagnosis, it is difficult for the prior art to achieve high-accurate fault diagnosis under zero sample conditions, especially due to insufficient diagnostic model performance and generalization capabilities caused by scarce and uneven distribution of fault data, and the existing DDPM methods have problems of proliferation of redundant information and coupled features in fault diagnosis.
Using the FE-FAGDM-based method, by collecting fault knowledge, selecting sensitive parameters, performing feature reconstruction, and combining diffusion model and semantic guidance generation method, high-quality zero-sample fault samples are generated and SVM is used for diagnosis.
It realizes high-quality generation of fault samples under zero sample conditions, improves the accuracy of fault diagnosis, and improves the effectiveness and practicality of zero sample fault diagnosis.
Smart Images

Figure CN120217120B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial equipment fault diagnosis, and relates to a fault diagnosis method under zero-sample conditions based on FE-FAGDM. Background Art
[0002] With the rapid advancement of technology, industrial systems are becoming increasingly complex, leading to an increase in the frequency and complexity of operational failures. These failures not only impact production efficiency but can also pose a threat to personnel safety. Therefore, the application of fault diagnosis technology in industrial systems is of great practical significance. By promptly and accurately detecting and diagnosing faults, system reliability and stability can be effectively improved, ensuring smooth production processes and ensuring personnel safety.
[0003] The diagnostic accuracy of fault diagnosis methods primarily depends on the quantity and quality of fault data. However, in industrial environments, acquiring fault data faces many challenges, such as the complexity of equipment structures, strict requirements for operational coordination, and the limitations of actual operating conditions. These challenges lead to the scarcity and uneven distribution of fault data, which greatly hinders the training and application of fault diagnosis models. When fault data is scarce, intelligent fault diagnosis models often struggle to extract sufficient diagnostic knowledge from the limited data, which directly impacts the model's diagnostic performance and generalization capabilities. When fault data is unevenly distributed, the diagnostic model may overemphasize healthy data during training, reducing its ability to identify faulty data.
[0004] Given the diversity of failure modes in complex systems, collecting data on all failure modes is impractical, and failure modes often have zero sample sizes. Therefore, achieving highly accurate fault diagnosis under zero-sample conditions is a pressing issue. Currently, data augmentation methods are widely used when dealing with limited sample sizes, and the Denoising Diffusion Probabilistic Model (DDPM) is a particularly powerful data augmentation method.
[0005] However, the DDPM method is currently rarely used in the field of fault diagnosis. Its primary application is to augment fault sample data to address the imbalanced fault diagnosis problem [Chinese Invention Patent 202411025618.3]. Research continues to address prominent issues, including redundant information in the original samples, difficulty in learning the coupled feature surge model, interference with diagnostic effectiveness, the difficulty of directly generating original samples, the unknown distribution of unseen fault samples after obtaining high-quality samples, and the inability of the DDPM model to independently infer the distribution of unseen fault samples. No research has yet applied DDPM or its variants to zero-shot fault diagnosis. Summary of the Invention
[0006] In response to the problems existing in the prior art, the present invention provides a zero-sample fault diagnosis method based on FE-FAGDM, which can achieve high-quality generation of fault samples of zero-sample fault modes in industrial processes and improve the accuracy of fault diagnosis.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is:
[0008] A fault diagnosis method based on FE-FAGDM under zero-sample conditions, the fault diagnosis method comprising the following steps:
[0009] Step 1: Collection of fault knowledge;
[0010] Raw fault data from the research object's operation is collected, including structured data on monitoring parameters and unstructured knowledge from fault descriptions and repair records. The start and end times of the monitoring parameter records are standardized. Based on the fault modes clearly identified in the fault descriptions and repair records, the corresponding fault attributes are extracted. Fault modes with samples are called seen fault modes, while those without samples are called zero-sample fault modes. These fault attributes include fault location, fault severity, and fault cause.
[0011] Step 2: Optimize monitoring parameters;
[0012] Based on the knowledge of fault modes and fault attributes obtained in step 1, a "fault mode-fault attribute" correlation matrix is constructed. Based on the fault attributes, an analysis of variance (ANOVA) is performed on the monitoring parameters. Parameters that are sensitive to changes in fault attributes are selected from all monitoring parameters and used as fault samples. The details are as follows:
[0013] Step 2.1, the “failure mode-failure attribute” association matrix Specifically:
[0014] (1)
[0015] in, Indicates the Failure Mode With the Fault attributes The correlation between them.
[0016] is a binary Boolean variable. If the failure mode and fault attributes If relevant, record If the failure mode and fault attributes Record if not relevant .
[0017] In step 2.2, the acquired fault modes all contain k parameters, and the parameter set is recorded as:
[0018] (2)
[0019] in, Represents a set of parameters, Indicates the monitoring parameters.
[0020] Step 2.3: Since the fault data of each fault mode contains all monitoring parameters, , according to the value of "fault mode-semantic attribute", the monitoring parameters in different fault modes are The data is divided into 2 groups. The grouping principles are:
[0021] If fault 1 with property The correlation value is "1", then the parameter of fault 1 The data were divided into Group A;
[0022] If fault 1 with property The correlation value is "0", then the parameter of fault 1 The data were grouped into Group B.
[0023] Step 2.4: Calculate monitoring parameters under two level groups based on analysis of variance (ANOVA) The within-group and between-group variances are:
[0024] The formula for calculating the Error Sum of Squares (SSE) is:
[0025] (3)
[0026] in, Indicates the The first in the group monitoring parameters, Indicates that it belongs to The average value of the parameters of the monitoring parameter group.
[0027] The calculation formula for the Across-Level Sum of Squares (SSA) is:
[0028] (4)
[0029] in, Represents a set of monitoring parameters The average value of the monitored parameters, Indicates the The total number of monitoring parameters contained in the group.
[0030] Step 2.5: Calculate the monitoring parameters based on the error sum of squares and cross-level sum of squares obtained in step 2.4. With attributes The sensitivity relationship between The value is calculated as follows:
[0031] (5)
[0032] in, The larger the value, the more obvious the difference between groups, and the smaller the difference within the group, which means that the monitoring parameters Attributes The higher the sensitivity of the change, and Have higher relevance.
[0033] Step 2.6: Analyze all monitoring parameters according to the process from step 2.3 to step 2.5, and analyze the sensitivity relationship of all parameters. The values are normalized to obtain the normalized sensitivity relationship Value. Select the value in descending order. The parameters are used as the fault semantic attributes The optimal parameters are then combined for all the optimal parameters of the fault semantic attributes, thereby discarding the parameters that are not sensitive to the fault semantic attributes, and finally obtaining The preferred parameter set for parameters:
[0034] (6)
[0035] Step 3: Fault sample feature reconstruction:
[0036] The optimal parameter set obtained in step 2 is used as the fault sample, and the fault sample is reconstructed based on principal component analysis (PCA) and kernel principal component analysis (KPCA). Specifically, linear and nonlinear transformations are performed, and the fault sample is finally reconstructed into a feature-reconstructed fault sample. This aims to reduce the coupling of fault features and reduce the difficulty of learning the coupling relationship between multiple fault attributes in the generated model in step 4. Specifically:
[0037] Step 3.1: Linear feature analysis of fault samples based on principal component analysis (PCA) Extraction of nonlinear features of fault samples based on kernel principal component analysis KPCA The extraction is shown below.
[0038] (7)
[0039] (8)
[0040] in, Represents the linear features extracted based on PCA, Represents the nonlinear features extracted based on KPCA.
[0041] In step 3.2, linear features and nonlinear features are geometrically concatenated to achieve feature reconstruction of the fault sample, thereby obtaining a feature-reconstructed fault sample for each fault mode, as shown in the following formula.
[0042] (9)
[0043] in, Represents feature reconstruction fault samples.
[0044] Step 4: Feature reconstruction sample guide generation:
[0045] A feature-enhanced and fault attribute-guided diffusion model (FE-FAGDM) is constructed. Fault samples and corresponding fault attribute labels are reconstructed using seen class features to train the FE-FAGDM. All fault attributes of the zero-shot fault mode are embedded into the FE-FAGDM as guiding conditions to obtain generated fault samples of the zero-shot fault mode. Specifically:
[0046] In step 4.1, we construct a feature-enhanced and fault attribute-guided diffusion model (FE-FAGDM). The FE-FAGDM model consists of three downsampling layers, two temporal embedding layers, two semantic embedding layers, three upsampling layers, and one output layer. The downsampling, upsampling, and output layers all use convolutional architectures, while the temporal and semantic embedding layers use fully connected architectures. The detailed structure of the model is as follows:
[0047] Table 1 shows the FE-FAGDM model structure.
[0048]
[0049] Where Conv(a, b), Linear(a, b) and ConvT(a, b) represent the convolutional layer, the fully connected layer and the deconvolution layer respectively. The input is a-dimensional and the output is b-dimensional. GELU represents the Gaussian error linear unit and ReLU represents the rectified linear unit. The number of fault attributes is Therefore, the input dimension of the semantic embedding layer is indivual.
[0050] Step 4.2: Construct the loss function of the FE-FAGDM model. The loss function of the FE-FAGDM is constructed based on the data generation process of the Denoising Diffusion Probabilistic Model (DDPM). The Denoising Diffusion Probabilistic Model is a data generation model that can generate the required data from random noise. During the training process, Gaussian noise is gradually added to the signal, so that the signal gradually transforms into pure random Gaussian noise. The Denoising Diffusion Probabilistic Model generates data by reversing the noise addition process. Given an original training signal and a final random Gaussian noise , defining the Markov state chain of its intermediate process ,in , T is the total number of time steps. Each state is obtained by adding Gaussian noise to the previous state:
[0051] (10)
[0052] in, represents the noise scale, represents standard normal random noise. Noise scale controls the amount of noise added at each step, The larger the value, the more noise is added. and The relationship between them is further expressed as follows:
[0053] (11)
[0054] in, ,and .
[0055] The main goal of DDPM is to learn a denoising model , according to the current state Predict previous state Since predicting the noise added from the previous state to the current state is easier for the model to learn than predicting the previous state itself. is reparameterized as:
[0056] (12)
[0057] in, represents the objective function when training the denoising diffusion probability model, is evenly distributed between 1 and Integers between represents the average value of the mean square error between the predicted noise and the true noise, The role of the neural network is to Prediction and solution Therefore, the loss function of DDPM is expressed as:
[0058] (13)
[0059] Although the denoising diffusion probability model has a strong data generation capability, it cannot achieve precise control over the generated data. To overcome this limitation, a classifier-free diffusion model guided generation method is used, which can control the generation of DDPM. Substitution , introduce fault guidance labels As a guiding condition, the loss function of the proposed FE-FAGDM is thus expressed as:
[0060] (14)
[0061] in, represents the fault label as the noise prediction result under the guidance condition;
[0062] Step 4.3, Fault Boot Label from Step 4.2 is the row vector of the fault semantic attribute matrix, Failure Boot tag As shown below.
[0063] (15)
[0064] in, For the Failure Mode The and Fault semantic attributes The correlation representation value of is the number of fault semantic attributes.
[0065] In step 4.4, the FE-FAGDM model loss function obtained in step 4.2 is used as the objective function of the training process. Based on the features obtained in step 3, the fault samples and corresponding fault semantic attribute labels are reconstructed, and the FE-FAGDM model constructed in step 4.1 is trained to obtain a trained FE-FAGDM model.
[0066] In step 4.5, the goal of the FE-FAGDM model is to generate high-quality fault samples for zero-sample fault modes. In this case, the fault semantic attributes of the zero-sample fault mode are required to be included in the fault semantic attributes used for model training. The FE-FAGDM model trained in step 4.4 generates fault samples under the guidance of the fault semantic attribute labels of the zero-sample fault mode. Based on the sampling principle of CFG-DDPM, the guidance labels of the zero-sample fault mode are used to generate high-quality fault samples. The generated fault samples that guide the model to generate zero-shot faults are based on the following formula:
[0067] (16)
[0068] in, represents the bootstrapped score estimate in the bootstrap model, It represents the guidance strength. The greater the guidance strength, the stronger the guidance of the guidance label to the model, and the generated samples are more consistent with the conditional information.
[0069] Step 5: Fault diagnosis:
[0070] SVM (Support Vector Machine, SVM) is selected as the fault diagnosis model. The generated fault samples of the zero-sample fault mode are used to train the SVM model. The feature-reconstructed fault samples of the actual zero-sample fault mode are input into the SVM model to achieve fault diagnosis under zero-sample conditions. Specifically:
[0071] The SVM model is a commonly used supervised learning algorithm that classifies data by finding an optimal hyperplane. The kernel function SVM can map feature vectors to a higher-dimensional space, making the originally linearly inseparable data linearly separable in the mapped space. The kernel function is expressed as:
[0072] (17)
[0073] in, Represents the kernel function, calculates the sample and Inner product in feature space; Represents a mapping function from input space to feature space; Represents vector transpose.
[0074] Furthermore, the hyperplane prediction model is implicitly constructed through the kernel function, which is expressed as:
[0075] (18)
[0076] in, Represents the classification decision function, and the prediction features reconstruct the fault category of the fault sample; represents the mapping function; represents the weight vector; Represents the bias term, which is used to adjust the position of the classification hyperplane. Represents vector transpose.
[0077] The beneficial effects of the present invention are:
[0078] (1) The present invention addresses the problem that the redundant information and coupling features in the original samples make it difficult to directly generate the original samples in the zero-sample generation task. The present invention introduces statistical methods based on the learning mechanism of the generation model to construct a fault sample feature reconstruction model that combines "parameter optimization" and "feature reconstruction". This model can realize the reconstruction of fault sample features for the zero-sample generation task and create good data conditions for zero-sample generation.
[0079] (2) In response to the DDPM model's inability to independently infer the distribution of unseen fault samples, this paper proposes a zero-shot fault diagnosis method based on the Feature-Enhanced and Fault Attributes-Guided Diffusion Model (FE-FAGDM). By combining the diffusion model's powerful data distribution learning capabilities with the semantically guided generation method, fault semantic attributes are extracted as guidance to achieve guided generation of zero-shot fault mode fault samples.
[0080] (3) Based on a chemical process fault dataset, this paper selected a support vector machine (SVM) as the fault diagnosis model for verification. The results demonstrated the effectiveness of the FE-FAGDM in zero-sample generation tasks and the strong practicality of the generated samples in implementing unseen fault diagnosis tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 It is a TEP dataset correlation matrix diagram provided by the present invention.
[0082] Figure 2 This is an example diagram of the original samples of the TEP dataset provided by the present invention.
[0083] Figure 3 This is a diagram showing the calculation results of the sensitivity F value of attribute 1 and each parameter provided by the present invention.
[0084] Figure 4 This is the reconstruction result diagram of feature 1 in fault mode 0 provided by the present invention.
[0085] Figure 5 This is a graph showing the model training loss value provided by the present invention.
[0086] Figure 6 This is an example diagram of sample generation of failure mode 6 provided by the present invention.
[0087] Figure 7 This is a confusion matrix diagram of the fault diagnosis results provided by the present invention. DETAILED DESCRIPTION
[0088] The present invention is further described below with reference to specific implementation cases.
[0089] This embodiment provides a fault diagnosis method based on FE-FAGDM under zero-sample conditions, comprising the following steps:
[0090] Step 1: Collection of fault knowledge;
[0091] Collect raw fault data from the operating system, including structured data on monitoring parameters and unstructured knowledge from fault descriptions and repair records. Unify the start and end times of the monitoring parameter records. Based on the fault modes identified in the fault descriptions and repair records, extract the corresponding fault attributes, including fault location, severity, and cause.
[0092] In this embodiment, the technical solution proposed in the present invention is verified using a public data set of the Tennessee Eastman Process (TEP process).
[0093] The TEP process is a chemical model simulation platform developed by Eastman Chemical Company in the United States. It simulates an actual chemical process, including multiple operating units such as a continuous stirred tank reactor, a partial condenser, a gas-liquid separation column, a stripping column, a reboiler, and a centrifugal compressor. The TEP process data is time-varying, strongly coupled, and nonlinear, making it widely used to test control and fault diagnosis models for complex industrial processes.
[0094] The TEP dataset includes 41 measurement variables and 11 operational variables, for a total of 52 monitoring parameters. The sampling frequency is 3 minutes per sample, and for each fault mode, a 24-hour sampling process is performed. This means that for each fault mode, each monitoring parameter contains 480 sampling points. Therefore, for each fault mode (including normal mode), the raw data size is 52 × 480.
[0095] The TEP dataset includes 21 fault modes, of which faults 1 to 15 have detailed fault knowledge descriptions, while faults 16 to 21 have almost no relevant fault descriptions. Therefore, the data of the first 15 fault modes and the normal data are selected as the case dataset, with sequence number 0 representing normal and sequence numbers 1-15 representing fault modes 1 to 15. The specific fault modes and related descriptions are shown in Table 2.
[0096] Table 2 is the description of the failure modes of the TEP dataset
[0097]
[0098] The entire TEP dataset contains 16 fault modes (including normal), so the data size of the entire TEP dataset is 16 × 52 × 480. The fault semantic attributes corresponding to the fault modes in the TEP dataset are then extracted, as shown in Table 3.
[0099] Table 3 is the description of fault attributes of the TEP dataset
[0100]
[0101] Step 2: Optimize monitoring parameters;
[0102] Based on the knowledge of fault modes and fault attributes obtained in step 1, a "fault mode-fault attribute" correlation matrix is constructed. Based on the fault attributes, an analysis of variance (ANOVA) is performed on the monitoring parameters. Parameters that are sensitive to changes in fault attributes are selected from all monitoring parameters and used as fault samples. The details are as follows:
[0103] Step 2.1, the “failure mode-failure attribute” correlation matrix is as follows:
[0104] (1)
[0105] in, Indicates the Failure Mode With the Fault attributes The correlation between them. is a binary Boolean variable. If the failure mode and fault attributes If relevant, record If the failure mode and fault attributes Record if not relevant .
[0106] Based on the fault modes obtained in step 1 and the extracted fault semantic attributes, combined with the inclusion relationship between each fault mode and fault attribute, the “fault mode-semantic attribute” association matrix on the TEP dataset is formed, as shown in Figure 1 As shown in the figure. For attribute 1, "Input A changes," faults 1, 2, 6, and 8 all cause changes in the input A content. However, the descriptions of the remaining faults do not indicate a change in the input A content. Therefore, the association between faults 1, 2, 6, and 8 and attribute 1 is recorded as "1," while the association between the remaining faults and attribute 1 is recorded as "0." Accordingly, the "fault mode-semantic attribute" association matrix can clearly express the association between other faults and all attributes.
[0107] To verify the applicability of the method proposed in this embodiment and demonstrate the complete technical process, 13 of the 15 fault modes were selected in this specific embodiment as observed fault types to represent the fault types with existing monitoring data in industrial processes, namely 0-5, 7, 8, 10-13, and 15. The remaining three fault modes were selected as zero-sample fault modes to represent fault modes for which there is no monitoring data due to difficulties in sample collection, namely 6, 9, and 14. Figure 2 This is the original parameter visualization result of parameter 1 corresponding to fault 0.
[0108] Step 2.2: The fault modes obtained in step 1 all contain k parameters, and the parameter set is recorded as:
[0109] (2)
[0110] in, Represents a set of parameters, Indicates the monitoring parameters.
[0111] Step 2.3: Since the fault data of each fault mode contains all monitoring parameters, , according to the value of "fault mode-semantic attribute", the monitoring parameters in different fault modes are The data is divided into 2 groups. The grouping principles are:
[0112] If fault 1 with property The correlation value is "1", then the parameter of fault 1 The data were divided into Group A;
[0113] If fault 1 with property The correlation value is "0", then the parameter of fault 1 The data were grouped into Group B.
[0114] For attribute 1, the fault modes with a correlation value of "1" with attribute 1 are 1, 2, 6, and 8, so the parameters of these four faults are grouped into group A, and the parameters of the remaining faults are grouped into group B.
[0115] In step 2.4, the observed fault samples and their corresponding fault attribute labels are used as input to calculate the monitoring parameters under two level groups based on variance analysis (ANOVA). within-group and between-group variance;
[0116] The formula for calculating the Error Sum of Squares (SSE) is:
[0117] (3)
[0118] in, Indicates the The first in the group monitoring parameters, Indicates that it belongs to The average value of the parameters of the monitoring parameter group.
[0119] The calculation formula for the Across-Level Sum of Squares (SSA) is:
[0120] (4)
[0121] in, Represents a set of monitoring parameters The average value of the monitored parameters, Indicates the The total number of monitoring parameters contained in the group.
[0122] Step 2.5: Calculate the monitoring parameters based on the error sum of squares and cross-level sum of squares obtained in step 2.4. With attributes The sensitivity relationship between The value is calculated as follows:
[0123] (5)
[0124] in, The larger the value, the more obvious the difference between groups, and the smaller the difference within the group, which means that the monitoring parameters Attributes The higher the sensitivity of the change, and Have higher relevance.
[0125] Sensitivity relationship A larger value indicates that the inter-group difference between levels A and B is more significant than the intra-group difference, which means that parameter 1 is more sensitive to changes in attribute 1, that is, parameter 1 has a higher correlation with parameter 1.
[0126] Step 2.6: Analyze all monitoring parameters according to the process from step 2.3 to step 2.5, and analyze the sensitivity relationship of all parameters. The values are normalized to obtain the normalized sensitivity Value. Normalized sensitivity between attribute 1 and each parameter Value Figure 3 As shown. According to the normalized sensitivity The values are sorted in descending order. The parameters are used as the fault semantic attributes The optimal parameters are then combined for all the optimal parameters of the fault semantic attributes, thereby discarding the parameters that are not sensitive to the fault semantic attributes, and finally obtaining The preferred parameter set for parameters:
[0127] (6)
[0128] In this embodiment, it is set The value of 15 is 15. This means that the top 15 parameters ranked by sensitivity are selected as the preferred parameters for this attribute. The preferred parameter set for attribute 1 is [0, 4, 5, 6, 8, 9, 12, 15, 20, 27, 33, 39, 40, 43, 46]. Further calculations yield the preferred parameters for each fault semantic attribute, as shown in Table 4. 41, 42, 43, 44, 45, 46, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101].
[0129] Table 4 shows the preferred parameter numbers for each fault attribute
[0130]
[0131] After parameter optimization is completed, the data size for seen fault data is reduced from the original 13 (number of seen fault modes) × 52 (number of monitored parameters) × 480 (number of sampling points) to 13 (number of seen faults) × 42 (number of optimized parameters) × 480 (number of sampling points). The parameter optimization results for zero-sample fault modes remain consistent with those for seen faults, meaning the data size for zero-sample fault modes is reduced from the original 3 (number of zero-sample fault modes) × 52 (number of monitored parameters) × 480 (number of sampling points) to 3 (number of unseen faults) × 42 (number of optimized parameters) × 480 (number of sampling points).
[0132] Step 3: Fault sample feature reconstruction:
[0133] The optimal parameter set obtained in step 2 is used as the fault sample, and the features of the fault sample are reconstructed based on principal component analysis and kernel principal component analysis, aiming to reduce the coupling of fault features and reduce the difficulty of learning the coupling relationship between multiple fault attributes in the generated model in step 4. Specifically:
[0134] Step 3.1: Linear feature analysis of fault samples based on PCA Extraction of nonlinear features of fault samples based on KPCA The extraction is shown below.
[0135] (7)
[0136] (8)
[0137] in, Represents the linear features extracted based on PCA, Represents the nonlinear features extracted based on KPCA.
[0138] In step 3.2, linear features and nonlinear features are geometrically concatenated to achieve feature reconstruction of the fault sample, thereby obtaining a feature-reconstructed fault sample for each fault mode, as shown in the following formula.
[0139] (9)
[0140] in, Represents feature reconstruction fault samples.
[0141] For fault 0, the fault sample feature reconstruction is performed. First, based on the optimal parameters, the 16 linear features are extracted by PCA, and then the 16 nonlinear features are extracted by KPCA. Then, the two types of features are spliced to form the feature reconstruction fault sample of fault 0, where feature 1 is as follows: Figure 4As shown in Figure 2, the feature-reconstructed fault samples of all seen fault modes and zero-shot fault modes are obtained accordingly. The feature-reconstructed fault samples of seen fault modes are then used to train the subsequent fault sample-guided generation model, and the feature-reconstructed fault samples of zero-shot fault modes are used as input data for the fault diagnosis model.
[0142] Step 4: Feature reconstruction sample guide generation:
[0143] In step 4.1, we construct the FE-FAGDM model. The model structure includes three downsampling layers, two temporal embedding layers, two semantic embedding layers, three upsampling layers, and one output layer. The downsampling layers, upsampling layers, and output layers all use convolutional structures, while the temporal embedding layer and semantic embedding layer use fully connected structures. The FE-FAGDM model structure is as follows:
[0144] Table 1 shows the FE-FAGDM model structure.
[0145]
[0146] In this implementation, Conv(a, b), Linear(a, b), and ConvT(a, b) represent the convolutional layer, the fully connected layer, and the deconvolutional layer, respectively. The input is a-dimensional, and the output is b-dimensional. GELU represents the Gaussian Error Linear Unit, and ReLU represents the Rectified Linear Unit. Since there are 20 fault attributes, the input dimension of the semantic embedding layer is also 20.
[0147] Step 4.2: Construct the loss function of the FE-FAGDM model. The loss function of the FE-FAGDM is constructed based on the data generation process of the Denoising Diffusion Probabilistic Model (DDPM). The Denoising Diffusion Probabilistic Model is a data generation model that can generate the required data from random noise. During the training process, Gaussian noise is gradually added to the signal, so that the signal gradually transforms into pure random Gaussian noise. The Denoising Diffusion Probabilistic Model generates data by reversing the noise addition process. Given an original training signal and a final random Gaussian noise , where the gradually adding noise process is a Markov state chain, that is, Depends only on , Represents the time step The potential state of the moment, where , T is the total number of time steps. Each state can be obtained by adding Gaussian noise to the previous state:
[0148] (10)
[0149] in, represents the noise scale, represents standard normal random noise. express Potential state at a moment; noise scale controls the amount of noise added at each step, The larger the value, the more noise is added. and The relationship between can be further expressed as follows:
[0150] (11)
[0151] in , Indicates the The noise control rate of the step; and , represents the product of the noise control rate of the noise addition process, Represents the noise adding process The noise control rate at time , where Indicates from 1 to At the middle moment of the noise adding process, represents the time step of the noise adding process.
[0152] The main goal of DDPM is to learn a denoising model , according to the current state Predict previous state Since predicting the noise added from the previous state to the current state is easier for the model to learn than predicting the previous state itself. can be reparametrized as:
[0153] (12)
[0154] in, represents the objective function when training the denoising diffusion probability model, is the optimization parameter in the denoising process, is evenly distributed between 1 and Integers between represents the average value of the mean square error between the predicted noise and the true noise, The role of the neural network is to Prediction and solution Therefore, the loss function of DDPM can be expressed as:
[0155] (13)
[0156] Although the denoising diffusion probability model has a strong data generation capability, it cannot achieve precise control over the generated data. To overcome this limitation, a classifier-free diffusion model guided generation method is used, which can control the generation of DDPM. Substitution , introduce fault guidance labels As a guiding condition, the loss function of the proposed FE-FAGDM can be expressed as:
[0157] (14)
[0158] in, The fault label is used as the noise prediction result under the guidance condition;
[0159] Step 4.3, Fault Boot Label from Step 4.2 is the row vector of the fault semantic attribute matrix, Failure Boot tag As shown below.
[0160] (15)
[0161] in, For the Failure Mode The and Fault semantic attributes The correlation representation value of is the number of fault semantic attributes.
[0162] In step 4.4, the FE-FAGDM model loss function obtained in step 4.2 is used as the objective function of the training process. Based on the features of the seen fault class obtained in step 3, the fault samples and fault semantic attribute labels are reconstructed as the training set, and the FE-FAGDM model constructed in step 4.1 is trained. The batch size of the FE-FAGDM model in the training process is set to 32, the epoch is set to 150, and the guidance intensity is set to 0. Set to 1. The change of training loss value with training rounds is as follows Figure 5 shown.
[0163] In step 4.5, the goal of the FE-FAGDM model is to generate high-quality fault samples for zero-sample fault modes. In this case, the fault semantic attributes of the zero-sample fault mode are required to be included in the fault semantic attributes used for model training. The fault semantic attribute labels of the zero-sample fault mode are used as input to guide the FE-FAGDM model trained in step 4.4 to generate fault samples. Based on the sampling principle of CFG-DDPM, the guidance labels of the zero-sample fault mode are used. The generated fault samples that guide the model to generate zero-shot faults are based on the following formula:
[0164] (16)
[0165] in, represents the bootstrapped score estimate in the bootstrap model, It represents the guidance strength. The greater the guidance strength, the stronger the guidance of the guidance label to the model, and the generated samples are more consistent with the conditional information.
[0166] In particular, the guiding strength Set to 1, the number of fault samples generated for each type of fault is 480. Taking zero-sample fault mode 6 as an example, the visualization results of fault sample generation are shown as follows: Figure 6 shown.
[0167] Step 5: Fault diagnosis:
[0168] Currently, we have obtained zero-sample fault mode feature reconstruction fault samples based on real monitoring data, as well as fault samples generated by the FE-FAGDM model. SVM (Support Vector Machine, SVM) is selected as the fault diagnosis model.
[0169] The SVM model is a commonly used supervised learning algorithm that classifies data by finding an optimal hyperplane. The kernel function SVM can map feature vectors to a higher-dimensional space, making the originally linearly inseparable data linearly separable in the mapped space. The kernel function is expressed as:
[0170] (17)
[0171] in, Represents the kernel function, calculates the sample and Inner product in feature space; Represents a mapping function from input space to feature space; Represents vector transpose.
[0172] Furthermore, the hyperplane prediction model is implicitly constructed through the kernel function, which is expressed as:
[0173] (18)
[0174] in, Represents the classification decision function, and the prediction features reconstruct the fault category of the fault sample; represents the mapping function; represents the weight vector; Represents the bias term, which is used to adjust the position of the classification hyperplane. Represents vector transpose.
[0175] The SVM classifier uses a linear kernel function and the regularization parameter C is set to 0.1. The generated zero-shot fault mode fault samples are used to train the SVM model. The feature-reconstructed fault samples of the actual zero-shot fault mode are then input into the SVM model to achieve fault diagnosis under zero-shot conditions.
[0176] Since there are three types of zero-sample fault modes, if a random diagnosis method is adopted, the accuracy rate is 33.33%, which is considered to be the effective line of the method in this paper. The matrix of the diagnosis results of the SVM diagnosis model for the three zero-sample fault modes is as follows: Figure 7 The final fault diagnosis accuracy is 65.69%, which is 32.36% higher than the effective line. This proves the effectiveness of FE-FAGDM in zero-shot generation tasks and the strong practicality of the generated samples in achieving zero-shot fault diagnosis tasks.
[0177] The above-described embodiments merely express the implementation methods of the present invention, but should not be understood as limiting the scope of the present invention. It should be pointed out that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, which all fall within the scope of protection of the present invention.
Claims
1. A zero-sample fault diagnosis method based on FE-FAGDM, characterized in that: The fault diagnosis method under zero-sample conditions comprises the following steps: Step 1: Collect fault knowledge; Collect the original fault data, unify the start and end time of the monitoring parameter records, and extract the fault attributes corresponding to the fault mode. The fault mode with samples is called the seen fault mode, and the fault mode without samples is called the zero-sample fault mode. Step 2: Optimize monitoring parameters; Build a "fault mode-fault attribute" correlation matrix based on the fault mode and fault attribute, and select monitoring parameters that are sensitive to changes in fault attributes as fault samples; Step 3: Fault sample feature reconstruction: Perform feature reconstruction on the fault sample to obtain a feature-reconstructed fault sample; Step 4: Feature reconstruction sample guide generation: A diffusion model FE-FAGDM is constructed. Fault samples and corresponding fault attribute labels are reconstructed using the features of the seen classes to train the FE-FAGDM. All fault attributes of the zero-shot fault mode are embedded into the FE-FAGDM as guiding conditions to obtain generated fault samples of the zero-shot fault mode. Specifically: Step 4.1: Construct a feature-enhanced and fault attribute-guided diffusion model, referred to as the FE-FAGDM model. The FE-FAGDM model consists of three downsampling layers, two temporal embedding layers, two semantic embedding layers, three upsampling layers, and one output layer. Step 4.2: Construct the loss function of the FE-FAGDM model. The loss function of the FE-FAGDM model is constructed based on the noise reduction diffusion probability model DDPM. During the training process, Gaussian noise is gradually added to the signal. Step 4.3, Fault Boot Label from Step 4.2 is the row vector of the fault semantic attribute matrix, the i-th fault F i Boot tag As follows: Among them, A ij is the i-th failure mode F i and the j-th fault semantic attribute A j The correlation representation value of m is the number of fault semantic attributes; In step 4.4, the FE-FAGDM model loss function of step 4.2 is used as the objective function of the training process. The FE-FAGDM model is trained based on the feature-reconstructed fault samples and corresponding fault semantic attribute labels obtained in step 3 to obtain a trained FE-FAGDM model. Step 4.5, the goal of the FE-FAGDM model is to generate high-quality fault samples for zero-sample fault modes; generate fault samples through the trained FE-FAGDM model; based on the sampling of CFG-DDPM, the guide label l of the zero-sample fault mode is used. u Guide the FE-FAGDM model to generate fault samples that generate zero-sample faults; Step 5: Fault diagnosis: The SVM model is selected as the fault diagnosis model, and the generated fault samples are used to train the SVM model. The feature-reconstructed fault samples of the actual zero-sample fault mode are input into the SVM model to realize fault diagnosis under zero-sample conditions.
2. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 1, characterized in that: In step 1, original fault data is collected during the operation of the research object, including structured data of monitoring parameters and unstructured knowledge of fault descriptions and fault repair records; the fault attributes include fault location, fault severity, and fault cause.
3. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.1, obtain the "failure mode-failure attribute" association matrix; In step 2.2, each fault mode contains k parameters, and the parameter set is recorded as: D origin =[P1,P2,...,P k ], where D origin Represents a set of parameters, P1, P2, ..., P k represents the 1st, 2nd, ..., kth monitoring parameter; Step 2.3, according to the value of "fault mode-semantic attribute", the monitoring parameter P in different fault modes is j The data were divided into 2 groups; Step 2.4, calculate the monitoring parameter P under two level groups j The within-group and between-group variances are used to obtain the error sum of squares SSE and the cross-level sum of squares SSA; Step 2.5: Calculate the monitoring parameter P based on the error sum of squares and the cross-level sum of squares j With attribute A i The sensitivity relationship F value is calculated as follows: Step 2.6: Analyze all monitoring parameters according to the process of steps 2.3 to 2.5, and obtain the normalized sensitivity relationship F value; select the parameters ranked first in descending order as the fault semantic attribute A. i Optimal parameters, for all the optimal parameters of the fault semantic attributes, take the union set, discard the parameters that are not sensitive to the fault semantic attributes, and get the set containing k s The preferred parameter set for parameters:
4. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 3, characterized in that: In the step 2: In step 2.1, the "failure mode-failure attribute" association matrix M F-A Specifically: Among them, A nm Indicates the nth failure mode F n With the mth fault attribute A m The correlation between A nm is a binary Boolean variable. If the failure mode F n With fault attribute A m If relevant, record A nm =1, if the failure mode F n With fault attribute A m If not relevant, record A nm =0; The grouping principle in step 2.3 is: if fault 1 is related to attribute A i The correlation value is "1", then the parameter P of fault 1 j The data is grouped into group A; if fault 1 is associated with attribute A i The correlation value is "0", then the parameter P of fault 1 j The data were divided into Group B; The calculation formula for the sum of squared errors SSE in step 2.4 is: Among them, P ik represents the kth monitoring parameter in the i-th group, represents the average value of the parameters belonging to the i-th monitoring parameter group; The calculation formula for the cross-level sum of squares SSA in step 2.4 is: in, Denotes the monitoring parameter set D origin The average value of the monitoring parameters, k i represents the total number of monitoring parameters contained in the i-th group; in step 2.5, the larger the value of F, the more obvious the difference between groups and the smaller the difference within the group.
5. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 3, characterized in that: The step three is as follows: Step 3.1: Linear feature D is used to analyze the fault samples based on principal component analysis (PCA). select-PCA Extraction of nonlinear features of fault samples based on kernel principal component analysis KPCA select-KPCA Extraction; In step 3.2, linear features and nonlinear features are geometrically spliced to achieve feature reconstruction of the fault sample and obtain feature-reconstructed fault samples for each fault mode.
6. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 5, characterized in that: In the step three: The extraction formula of step 3.1 is shown below: in, Represents the linear features extracted based on PCA, Represents the nonlinear features extracted based on KPCA; In step 3.2, the characteristic reconstruction fault sample of each fault mode is shown as follows: Among them, D rec Represents feature reconstruction fault samples.
7. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 1, characterized in that: In the step 4: In step 4.2, The specific process of constructing the loss function of FE-FAGDM is: Given an original training signal x0 and a final random Gaussian noise x T , x t Depends only on x t-1 , x t represents the potential state at time step t, where t = 1, 2, ..., T-1, and T is the total number of time steps; each state is obtained by adding Gaussian noise to the previous state: Among them, β t represents the noise scale, ε represents the standard normal random noise; x t-1 represents the potential state at time t-1; x t The relationship between and x0 is expressed as follows: Among them, α t =1-β t , α t represents the noise control rate at step t; and represents the product of the noise control rate of the noise adding process, α s represents the noise control rate at time s of the noise addition process, where s = 1, 2, ..., t represents the intermediate time of the noise addition process from 1 to t, and t represents the time step of the noise addition process; According to the current state x, the denoising diffusion probability model DDPM is used t Predict the previous state x t-1 , q θ (x t-1 |x) t Parameterized as: Where L(θ) represents the objective function when training the denoising diffusion probability model, θ is the optimization parameter in the denoising process, t is an integer uniformly distributed between 1 and T, E(x0, t, ε) represents the average value of the mean square error between the predicted noise and the true noise, and ε θ The role is to start from x t Predict and solve ε;q θ (x t-1 |x t ) is represented as a denoising model; Then the loss function of DDPM is expressed as: Using ε θ (z λ ) instead of And introduce the fault guidance label As the guiding condition, the loss function of FE-FAGDM is obtained as shown in formula (14): in, represents the fault label as the noise prediction result under the guidance condition; In step 4.5, the fault semantic attributes of the zero-sample fault mode are included in the fault semantic attributes used in model training; the sampling is based on the following formula: in, It represents the score estimate after guidance, w represents the guidance strength, and the greater the guidance strength, the stronger the guidance of the guidance label to the FE-FAGDM model, and the generated samples are more consistent with the conditional information.
8. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 5, characterized in that: The step five is specifically as follows: The feature vector is mapped to a higher-dimensional space through the kernel function SVM, where the kernel function is expressed as: Among them, K(a i ,a j ) represents the kernel function, calculating sample a i and a j Inner product in feature space; Represents a mapping function from input space to feature space; T represents vector transpose.
9. The zero-sample fault diagnosis method based on FE-FAGDM according to claim 8, characterized in that: The hyperplane prediction model is further implicitly constructed through the kernel function, which is expressed as: Among them, f(x) represents the classification decision function, and the prediction features reconstruct the fault category of the fault sample; Represents the mapping function; ξ represents the weight vector; b represents the bias term, which is used to adjust the position of the classification hyperplane; T represents the vector transpose.
Citation Information
Patent Citations
Small sample data processing method and device for bearing fault diagnosis
CN118568570A
Steering frame unknown fault diagnosis method based on generative model zero sample learning strategy
CN119358375A
Cited By
Fault diagnosis method and system based on multi-domain collaborative data enhancement
CN121070661A