An Engine Fault Feature Extraction Method Based on Data Augmentation and Self-Distillation Model
By combining an improved wavelet threshold function with a self-distillation model, the challenges of data denoising, balancing, and feature extraction in engine fault diagnosis were solved, resulting in higher diagnostic accuracy and model generalization ability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-03
AI Technical Summary
Existing engine fault diagnosis technologies suffer from contradictions between accuracy and robustness, effectiveness dilemmas, and the curse of dimensionality in vibration signal denoising, data balancing, and feature extraction, resulting in weak model generalization ability and low diagnostic accuracy.
An improved wavelet threshold function is used for data denoising. Combined with a data balancing model and a self-distillation-based autoencoder feature extraction model, data augmentation is performed by synthesizing minority oversampling and conditional generative adversarial networks. Finally, the self-distillation mechanism is used to optimize the student model parameters for fault feature extraction.
It improves the accuracy and reliability of engine fault feature extraction, solves the problem of low diagnostic accuracy in complex scenarios with few fault data samples, class imbalance, and large data noise interference, and enhances the generalization ability of the model.
Smart Images

Figure CN121579996B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for extracting engine fault features based on data augmentation and self-distillation models, belonging to the field of engine fault detection technology. Background Technology
[0002] As a core component of various power equipment, the engine's accurate monitoring and early fault diagnosis are crucial for ensuring equipment safety and reducing maintenance costs. Current engine fault diagnosis technology is data-driven, encompassing core stages such as data preprocessing, feature extraction, and pattern recognition. In the data preprocessing stage, vibration signal denoising is a key step. Traditional methods primarily employ wavelet thresholding, using hard or soft thresholding functions to shrink wavelet coefficients and suppress noise. In the data balancing stage, oversampling, undersampling, or generative adversarial networks (GANs) are often used to address class imbalance issues, generating synthetic samples. Feature extraction relies heavily on techniques such as principal component analysis, deep autoencoders, or knowledge distillation to mine fault features from high-dimensional monitoring data. These technologies constitute the basic framework for existing engine fault diagnosis methods and are widely applied in industrial scenarios.
[0003] However, existing technologies face three major drawbacks in practical applications: First, vibration signal denoising presents a contradiction between accuracy and robustness. Hard thresholding functions cause signal oscillations due to discontinuities at the threshold point, while soft thresholding functions, although ensuring continuity, suffer from constant bias leading to signal distortion. Furthermore, fixed thresholding strategies do not consider the attenuation characteristics of noise energy with the number of decomposition layers, easily resulting in the loss of useful low-frequency signals or the retention of high-frequency noise, ultimately making fault feature extraction difficult. Second, data balancing faces an effectiveness dilemma. Oversampling introduces redundant noise by repeating minority class samples, leading to model overfitting. Undersampling, by randomly deleting majority class samples, may lose key information. Conventional generative adversarial networks suffer from training instability, gradient vanishing, and mode collapse when fault samples are extremely sparse, making it difficult to generate high-quality samples that conform to the true distribution and limiting the model's generalization ability. Third, feature extraction suffers from the curse of dimensionality and an imbalance in accuracy. Principal component analysis can only handle linear relationships and has limited ability to represent nonlinear features; deep autoencoder models have redundant parameters and low training efficiency; knowledge distillation relies excessively on the performance of a single teacher model, easily inheriting teacher bias and leading to a decrease in feature extraction accuracy. The aforementioned shortcomings collectively result in existing methods exhibiting weak generalization ability and low diagnostic accuracy in complex scenarios with few samples, class imbalance, and high data noise, making it difficult to meet the requirements for high-precision fault diagnosis. Summary of the Invention
[0004] The purpose of this invention is to provide an engine fault feature extraction method based on data augmentation and a self-distillation model. This method uses an improved wavelet threshold function to denoise the engine dataset, performs data augmentation based on a data balancing model, and trains a fault feature extraction model based on self-distillation for fault feature extraction. This addresses the problems in existing engine fault diagnosis, such as difficulty in fault feature extraction, weak model generalization ability, and low diagnostic accuracy caused by the limited number of available fault data samples, class imbalance, and large data noise interference.
[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.
[0006] This invention provides a method for extracting engine fault features based on data augmentation and a self-distillation model, comprising:
[0007] Acquire engine datasets that include both normal operation data and fault operation data;
[0008] The wavelet threshold function is improved by setting an adjustment factor so that the improved wavelet threshold function is continuous within the wavelet threshold range;
[0009] The engine dataset is denoised using an improved wavelet threshold function to obtain a denoised engine dataset.
[0010] Based on the denoised engine dataset, data augmentation is performed using a trained data balancing model to obtain a data-augmented engine dataset; the data balancing model includes a synthetic minority oversampling model and a conditional generative adversarial network model connected in sequence.
[0011] Based on the augmented engine dataset, a self-distillation-based autoencoder feature extraction model is trained to extract engine fault features.
[0012] Furthermore, the improved wavelet threshold function is expressed as:
[0013] ;
[0014] In the formula, Indicates the first The first decomposition level Wavelet estimation coefficients after position filtering Indicates the first The first decomposition level Wavelet estimation coefficients before filtering at each position Represents a symbolic function. Indicates the wavelet threshold range, , Indicates the standard deviation of noise. , This represents the absolute value of the median of the scaling factor. Represented by natural numbers The logarithmic function with base 1. Indicates signal length. Indicates the regulating factor;
[0015] when hour, ,Right now It exhibits continuity within the wavelet threshold range, where, This indicates that the wavelet threshold is trending towards , This indicates that the wavelet threshold tends negatively to , Indicates the limit;
[0016] when hour, , ,Right now It exhibits continuity within the wavelet threshold range. It represents positive infinity.
[0017] Furthermore, the training method for the data balancing model includes:
[0018] The fault operation data is augmented sequentially using a synthetic minority oversampling model and a conditional generative adversarial network model to obtain augmented fault operation data.
[0019] Based on the data-augmented fault operation data, the generator is constrained by Wasserstein distance, and a constrained data balance model loss function is constructed.
[0020] By using a preset first penalty function, a preset second penalty function, and a dual interpolation strategy to improve the data balance model loss function after constraints, the improved data balance model loss function is obtained.
[0021] Based on the data-augmented fault operation data and the improved data balancing model loss function, the data balancing model is trained to obtain a well-trained data balancing model.
[0022] Furthermore, the loss function used to augment the faulty operational data by sequentially employing a synthetic minority oversampling model and a conditional generative adversarial network model is expressed as:
[0023] ;
[0024] In the formula, Indicates the generator when performing data augmentation Take the minimum value Discriminator Take the maximum value The loss value of the loss function used. express The mathematical expectation, This represents the fault operation data after data augmentation. express The mathematical expectation, This represents the noise vector after data augmentation. This represents a real sample of the fault operation data after data augmentation. This indicates the category labels for the data augmented during fault operation. Represents a random noise vector. This represents a conditional discriminator. Represents a condition generator. This represents the logarithmic function with the natural number 10 as the base.
[0025] Furthermore, the constrained data balancing model loss function is expressed as:
[0026] ;
[0027] In the formula, This represents the loss value of the loss function in the constrained data balancing model. Represents the input random noise vector Category labels for data-enhanced fault operation data Condition generator Compared with data-enhanced engine data square Euclidean distance Mathematical expectation The weight, This represents the square of the Euclidean distance.
[0028] Furthermore, the improved data balancing model loss function is expressed as:
[0029] ;
[0030] In the formula, This represents the loss value of the improved data balancing model loss function. This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. The distribution of fault operation data before data augmentation and fault operation data generated by the generator must be obeyed. The mathematical expectation, This represents the interpolated sample of real fault operation data after data augmentation and the fault operation data generated by the generator. The distribution of fault operation data after data augmentation and fault operation data generated by the generator must conform to the data distribution of fault operation data. The mathematical expectation, This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. gradient operator, This represents the interpolated sample of real fault operation data after data augmentation and the fault operation data generated by the generator. gradient operator, This represents the weight of the penalty function term. and Represent the first penalty function respectively Mathematical expectation Weights and the second penalty function Mathematical expectation The weight.
[0031] Furthermore, the method for constructing the autoencoder feature extraction model based on self-distillation includes:
[0032] Step S1: Based on the trained data balancing model, threshold filtering is performed according to the probability values of each label in the student model to obtain the confidence prediction function;
[0033] Step S2: Correct the feature extraction error of the teacher model according to the confidence prediction function to obtain the error-corrected teacher model loss function;
[0034] Step S3: Based on the error-corrected teacher model loss function, construct an improved student model loss function through knowledge distillation;
[0035] Step S4: Update the parameters of the student model through backpropagation based on the improved student model loss function;
[0036] Step S5: Use the updated student model as the new teacher model, and repeat steps S2 to S4 for iterative training;
[0037] When the fault classification accuracy of the student model reaches the preset threshold, the iteration stops, and the current student model is used as the self-distillation-based autoencoder feature extraction model to obtain the constructed self-distillation-based autoencoder feature extraction model.
[0038] Furthermore, the confidence prediction function is expressed as:
[0039] ;
[0040] In the formula, Student model The value of the confidence prediction function, The superscript indicates the confidence prediction value used to distinguish the values after threshold filtering. Student model The Confidence prediction values before each label threshold filtering This represents the preset confidence threshold. To represent the empty set, Indicates intersection.
[0041] Furthermore, the error-corrected teacher model loss function is expressed as:
[0042] ;
[0043] In the formula, This represents the teacher model loss function after error correction. This represents the teacher model loss function before error correction. Mean square error function The weight, The teacher model before error correction is based on input features. The predicted vector.
[0044] Furthermore, the improved student model loss function is expressed as:
[0045] ;
[0046] In the formula, This represents the improved student model loss function. This represents the cross-entropy loss of the student model. This represents the ranking loss of the student model. The mean squared error represents the softened probability distribution of the corrected teacher and student models. The weights, where, , The predicted vector representing the softening probability of the student model before error correction. Represents the normalized exponential function, This represents the temperature scaling factor. Indicates the correction term The weighting coefficients, The teacher model after error correction is based on input features. The predicted vector.
[0047] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0048] 1. This invention improves the wavelet threshold function by setting an adjustment factor to ensure its continuity within the wavelet threshold range, and uses this to denoise the engine dataset. Combined with data augmentation of the data balancing model and training of the autoencoder feature extraction model based on self-distillation, it effectively improves the accuracy and reliability of engine fault feature extraction. It solves the problems of difficult fault feature extraction, weak model generalization ability, and low diagnostic accuracy caused by the limited number of available fault data samples, class imbalance, and large data noise interference in existing engine fault diagnosis.
[0049] 2. This invention employs a data balancing model combining a synthetic minority oversampling model and a conditional generative adversarial network model. Through complex training methods, including utilizing Wasserstein distance constraints, pre-defined penalty functions, and dual interpolation strategies to improve the loss function, data augmentation is performed on the fault operation data, resulting in a rich and high-quality augmented engine dataset. This high-quality data provides a solid foundation for subsequent fault feature extraction, helping to more accurately capture engine fault characteristics.
[0050] 3. This invention, based on a pre-trained data balancing model, uses threshold filtering based on the probability values of each label in the student model to obtain a confidence prediction function. This function then corrects the feature extraction error of the teacher model, constructing an improved loss function for the student model. The student model parameters are updated through backpropagation, and iterative training is performed. This self-distillation method allows the model to continuously optimize during training. Training stops when the student model's fault classification accuracy reaches a preset threshold. The resulting autoencoder feature extraction model based on self-distillation can extract engine fault features more accurately, significantly improving the accuracy of fault feature extraction. Attached Figure Description
[0051] Figure 1 This is a flowchart illustrating an engine fault feature extraction method based on a data augmentation and self-distillation model provided in an embodiment of the present invention.
[0052] Figure 2 This is a schematic diagram comparing the effects of different noise reduction methods provided in the embodiments of the present invention;
[0053] Figure 3 This is a schematic diagram of the F1 score of the data-enhanced engine dataset provided in an embodiment of the present invention;
[0054] Figure 4 This is a schematic diagram of the AUPRC (Area Under the Precision-Recall Curve) score of the data-enhanced engine dataset provided in this embodiment of the invention.
[0055] Figure 5This is a schematic diagram of the loss function curve of the improved student model provided in the embodiment of the present invention;
[0056] Figure 6 This is a comparative diagram of model evaluation metrics provided in the embodiments of the present invention, wherein the left side (a) is a comparative diagram of F1 scores and the right side (b) is a comparative diagram of Hamming loss. Detailed Implementation
[0057] It should be noted that the Non model refers to a baseline control group that does not receive any intervention or treatment, used to compare with the experimental group to verify the effect of the variable.
[0058] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0059] Example 1
[0060] like Figure 1 As shown in the figure, this embodiment introduces a method for extracting engine fault features based on data augmentation and self-distillation models, including:
[0061] Step 1: Obtain the engine dataset, which includes normal operation data and fault operation data.
[0062] This embodiment comprehensively collects data covering both normal and faulty engine operating states, providing a rich and complete foundational data sample for subsequent fault feature extraction. This dataset, encompassing different states, more realistically reflects various situations encountered by the engine during actual operation, facilitating the mining of fault features from multiple dimensions and scenarios. This improves the comprehensiveness and accuracy of fault feature extraction, laying a solid foundation for accurate engine fault diagnosis.
[0063] Step 2: Set adjustment factors to improve the wavelet threshold function so that the improved wavelet threshold function has continuity within the wavelet threshold range.
[0064] This embodiment improves the wavelet threshold function by setting an adjustment factor, making it continuous within the wavelet threshold range. This improvement effectively avoids the discontinuity problem that may occur at the threshold point in traditional wavelet threshold functions, thus processing the signal more smoothly during denoising and reducing additional noise or signal distortion caused by function discontinuities. The continuous wavelet threshold function can more accurately separate noise and effective components in the signal, improving the denoising effect and providing clearer and more reliable data for subsequent accurate extraction of engine fault features.
[0065] Step 3: Use the improved wavelet threshold function to denoise the engine dataset to obtain the denoised engine dataset.
[0066] This embodiment utilizes an improved, continuous wavelet threshold function to denoise the engine dataset, effectively removing various types of noise interference from the data. During engine operation, various factors generate different types of noise, which can mask the true fault characteristic signals. The engine dataset obtained after this denoising step shows significantly improved signal quality and a substantial reduction in noise levels, making fault characteristics more prominent and easier to identify, thus providing high-quality data input for subsequent data augmentation and fault feature extraction steps.
[0067] Step 4: Based on the denoised engine dataset, perform data augmentation using the trained data balancing model to obtain the augmented engine dataset.
[0068] In this embodiment, the data balancing model includes a synthetic minority oversampling model and a conditional generative adversarial network model connected in sequence.
[0069] This embodiment uses a pre-trained data balancing model, consisting of a synthetic minority oversampling model and a conditional generative adversarial network (GAN) model, to augment a denoised engine dataset. Engine fault data often suffers from class imbalance, meaning that the number of data samples for certain fault types is far less than that for normal or other fault types. This data balancing model initially increases the number of minority fault samples using the synthetic minority oversampling model, and then further generates more realistic and diverse fault samples using the GAN model, effectively solving the data imbalance problem. The augmented engine dataset has a more balanced number of fault samples across different types, a more reasonable data distribution, and better reflects the diversity and complexity of engine faults. This helps improve the identification and generalization capabilities of the subsequent self-distillation-based autoencoder feature extraction model used for fault feature extraction.
[0070] Step 5: Based on the data-augmented engine dataset, train a self-distillation-based autoencoder feature extraction model to extract engine fault features.
[0071] This embodiment utilizes a pre-constructed autoencoder feature extraction model based on self-distillation to train fault feature extraction on a data-augmented engine dataset. The self-distillation mechanism allows the model to fully leverage its own knowledge during training; the teacher model transfers knowledge to the student model, and through continuous iteration and optimization, the student model becomes more accurate and efficient in extracting fault features. The autoencoder structure automatically learns the inherent feature representation of the data, eliminating the need for manual feature extraction rule design. It accurately extracts engine fault features from complex data-augmented engine datasets. These features are highly representative and discriminative, providing crucial evidence for engine fault diagnosis and effectively improving the accuracy and reliability of fault diagnosis.
[0072] Example 2
[0073] Based on the same inventive concept as Example 1, this example introduces a method for extracting engine fault features based on data augmentation and self-distillation models, including:
[0074] Step 1: Obtain the engine dataset, which includes normal operation data and fault operation data.
[0075] Step 2: Set an adjustment factor to improve the wavelet threshold function so that the improved wavelet threshold function has continuity within the wavelet threshold range.
[0076] In this embodiment, the improved wavelet threshold function is expressed as:
[0077] ;
[0078] In the formula, Indicates the first The first decomposition level Wavelet estimation coefficients after position filtering Indicates the first The first decomposition level Wavelet estimation coefficients before filtering at each position Represents a symbolic function. Indicates the wavelet threshold range, , Indicates the standard deviation of noise. , This represents the absolute value of the median of the scaling factor. Represented by natural numbers The logarithmic function with base 1. Indicates signal length. Indicates the regulating factor;
[0079] when hour, ,Right now It exhibits continuity within the wavelet threshold range, where, This indicates that the wavelet threshold is trending towards , This indicates that the wavelet threshold tends negatively to , Indicates the limit;
[0080] when hour, , ,Right now It exhibits continuity within the wavelet threshold range. It represents positive infinity.
[0081] Step 3: Use the improved wavelet threshold function to denoise the engine dataset to obtain the denoised engine dataset.
[0082] like Figure 2 The diagram illustrates a comparison of the effects of different denoising methods provided in this embodiment of the invention. Specifically, it compares the denoising effects of other methods with the improved wavelet threshold function. In this embodiment, after acquiring engine datasets including normal operation data and fault operation data, the denoising effect is measured using the signal-to-noise ratio (SNR) and root mean square error (RMSE). Then, four different wavelet thresholding denoising methods—hard thresholding, soft thresholding, other wavelet thresholding functions, and the improved wavelet thresholding function—are applied to the noisy signal. Simultaneously, a particle swarm optimization algorithm is used to adjust the adjustment factor in the improved wavelet thresholding function of this invention. Perform an optimal value search.
[0083] from Figure 2 It can be seen that the signal-to-noise ratio (SNR) of the original engine dataset for both normal and fault operation is 5.91 dB, and the RMSE is 0.172174. The SNR of the engine dataset after denoising using a hard thresholding function is 5.91 dB, and the RMSE is 0.172174. The SNR of the engine dataset after denoising using a soft thresholding function is 2.17 dB, and the RMSE is 0.407113. The SNR of the engine dataset after denoising using other wavelet thresholding functions is 8.76 dB, and the RMSE is 0.091169. This invention utilizes an improved adjustment factor... The signal-to-noise ratio (SNR) of the engine dataset signal after denoising with a wavelet threshold function of 7.0194 is 10.50 dB, and the RMSE is 0.059886. Figure 2 The experimental results show that the improved wavelet threshold function of this invention achieves the best denoising effect.
[0084] Step 4: Based on the denoised engine dataset, perform data augmentation using the trained data balancing model to obtain the augmented engine dataset.
[0085] In this embodiment, the data balancing model includes a synthetic minority oversampling model and a conditional generative adversarial network model connected in sequence. The training method for the data balancing model includes:
[0086] Step 4.1: Use the synthetic minority oversampling model and the conditional generative adversarial network model sequentially to augment the fault operation data, and obtain the augmented fault operation data.
[0087] In this embodiment, the loss function used to augment faulty operational data by sequentially employing a synthetic minority oversampling model and a conditional generative adversarial network model is expressed as:
[0088] ;
[0089] In the formula, Indicates the generator when performing data augmentation Take the minimum value Discriminator Take the maximum value The loss value of the loss function used. express The mathematical expectation, This represents the fault operation data after data augmentation. express The mathematical expectation, This represents the noise vector after data augmentation. This represents a real sample of the fault operation data after data augmentation. This indicates the category labels for the data augmented during fault operation. Represents a random noise vector. This represents a conditional discriminator. Represents a condition generator. This represents the logarithmic function with the natural number 10 as the base.
[0090] Step 4.2: Based on the data-augmented fault operation data, constrain the generator using Wasserstein distance and construct the constrained data balance model loss function.
[0091] In this embodiment, the constrained data balancing model loss function is expressed as:
[0092] ;
[0093] In the formula, This represents the loss value of the loss function in the constrained data balancing model. Represents the input random noise vector Category labels for data-enhanced fault operation data Condition generator Compared with data-enhanced engine data square Euclidean distance Mathematical expectation The weight, This represents the square of the Euclidean distance.
[0094] Step 4.3: Improve the data balance model loss function after constraints using the preset first penalty function, the preset second penalty function, and the double interpolation strategy to obtain the improved data balance model loss function.
[0095] In this embodiment, the improved data balancing model loss function is expressed as:
[0096] ;
[0097] In the formula, This represents the loss value of the improved data balancing model loss function. This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. The distribution of fault operation data before data augmentation and fault operation data generated by the generator must be obeyed. The mathematical expectation, This represents the interpolated sample of real fault operation data after data augmentation and the fault operation data generated by the generator. The distribution of fault operation data after data augmentation and fault operation data generated by the generator must conform to the data distribution of fault operation data. The mathematical expectation, This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. gradient operator, This represents the interpolated sample of real fault operation data after data augmentation and the fault operation data generated by the generator. gradient operator, This represents the weight of the penalty function term. and Represent the first penalty function respectively Mathematical expectation Weights and the second penalty function Mathematical expectation The weight.
[0098] Step 4.4: Based on the data augmented fault operation data and the improved data balancing model loss function, train the data balancing model to obtain the trained data balancing model.
[0099] Figure 3This is a schematic diagram of the F1 score of the data-augmented engine dataset provided in this embodiment of the invention. In this embodiment, the denoised engine dataset is divided into a training set and a test set in an 8:2 ratio. Stratified sampling is used during the partitioning to ensure that the sample ratio in the training and test sets remains consistent. The test set is then used only for final evaluation. Five algorithm models are then designed and implemented sequentially. Simultaneously, the performance of the classifier on the validation set is monitored during training. If the performance does not improve for five consecutive epochs, training will automatically stop and revert to the optimal model, thereby ensuring that each model is evaluated at its best performance point and preventing overfitting.
[0100] from Figure 3 As can be seen, after five rounds of comparison, the final average F1 score is as follows: Non model: 0.103; SMOTE model (Synthetic Minority Over-sampling Technique): 0.783; GAN model: 0.794; SMOTEfied-GAN model (Synthetic Minority Over-sampling Technique-ified Generative Adversarial Network): 0.792; and the data balancing model trained in this invention (SMOTE-WCGAN-GP): 0.798. The experimental results show that the data balancing model trained in this invention has the best data augmentation effect on the engine dataset.
[0101] Figure 4 This is a schematic diagram of the AUPRC (Area Under the Precision-Recall Curve) score of the data-augmented engine dataset provided in this embodiment of the invention. Figure 4 As can be seen, after five rounds of comparison, the final average AUPRC score is 0.7099 for the Non model, 0.7970 for the SMOTE model, 0.8299 for the GAN model, 0.8265 for the SMOTE-fed-GAN model, and 0.8386 for the data balancing model trained in this invention (SMOTE-WCGAN-GP). The experimental results show that the data balancing model trained in this invention has the best data augmentation effect on the engine dataset.
[0102] Step 5: Based on the data-augmented engine dataset, train a self-distillation-based autoencoder feature extraction model to extract engine fault features.
[0103] In this embodiment, the method for constructing a self-distillation-based autoencoder feature extraction model includes:
[0104] Step 5.1: Based on the trained data balancing model, threshold filtering is performed according to the probability values of each label in the student model to obtain the confidence prediction function.
[0105] In this embodiment, the confidence prediction function is expressed as:
[0106] ;
[0107] In the formula, Student model The value of the confidence prediction function, The superscript indicates the confidence prediction value used to distinguish the values after threshold filtering. Student model The Confidence prediction values before each label threshold filtering This represents the preset confidence threshold. To represent the empty set, Indicates intersection.
[0108] Step 5.2: Correct the feature extraction error of the teacher model according to the confidence prediction function to obtain the error-corrected loss function of the teacher model.
[0109] In this embodiment, the error-corrected teacher model loss function is expressed as:
[0110] ;
[0111] In the formula, This represents the teacher model loss function after error correction. This represents the teacher model loss function before error correction. Mean square error function The weight, The teacher model before error correction is based on input features. The predicted vector.
[0112] Step 5.3: Based on the teacher model loss function after error correction, construct the improved student model loss function through knowledge distillation.
[0113] In this embodiment, the improved student model loss function is expressed as:
[0114] ;
[0115] In the formula, This represents the improved student model loss function. This represents the cross-entropy loss of the student model. This represents the ranking loss of the student model. The mean squared error represents the softened probability distribution of the corrected teacher and student models. The weights, where, , The predicted vector representing the softening probability of the student model before error correction. Represents the normalized exponential function, This represents the temperature scaling factor. Indicates the correction term The weighting coefficients, The teacher model after error correction is based on input features. The predicted vector.
[0116] Step 5.4: Update the parameters of the student model through backpropagation based on the improved student model loss function.
[0117] Step 5.5: Use the updated student model as the new teacher model and repeat steps 5.2 to 5.4 for iterative training.
[0118] When the fault classification accuracy of the student model reaches a preset threshold, the iteration stops, and the current student model is used as the self-distillation-based autoencoder feature extraction model to obtain the constructed self-distillation-based autoencoder feature extraction model.
[0119] Figure 5 This is a schematic diagram of the loss function curve of the improved student model provided in this embodiment of the invention. The teacher model and student model in the constructed self-distillation-based autoencoder feature extraction model are each trained for 20 rounds to obtain the initial student model. Then, a feedback loop of three iterations is initiated. In each loop, the student model selects confidence prediction functions with confidence scores higher than a threshold. These confidence prediction functions are used as additional supervision functions for the teacher model, which is then fine-tuned for 8 rounds. The corrected teacher model and the uncorrected teacher model together generate a more complex soft objective, which is then used to further optimize the student model for 10 rounds.
[0120] from Figure 5It can be seen that when the student model undergoes iterative retraining after twenty training rounds, the loss value of the student model in the self-distillation-feedback loop-teacher-encoder-decoder auto-encoder feature extraction model (SDFL-TEA) constructed in this invention is lower than that of the student model in the knowledge distillation-teacher-encoder-decoder auto-encoder model (KD-TEA). The experimental results show that the self-distillation-feedback-feedback-feedback-feedback-loop-teacher-encoder-decoder auto-encoder model constructed in this invention exhibits lower loss during fault feature extraction training.
[0121] Figure 6 This is a comparative diagram of model evaluation indicators provided in the embodiments of the present invention. Figure 6 As shown in Figure (a), the constructed self-distillation-based autoencoder feature extraction model (SDFL-TEA) outperforms the knowledge distillation-based teacher-student model (KD-TEA) in all three F1 score types: Micro-F1 (micro-averaged F1-score), Macro-F1 (macro-averaged F1-score), and Example-F1 (example-based F1-score). Figure 6 Figure (b) shows that the constructed autoencoder feature extraction model based on self-distillation performs significantly worse than the teacher-student model based on knowledge distillation in Hamming loss, indicating that the constructed autoencoder feature extraction model based on self-distillation in this invention makes significantly fewer independent errors when predicting labels for each sample. The experimental results show that the constructed autoencoder feature extraction model based on self-distillation is superior to the teacher-student model based on knowledge distillation in fault feature extraction.
[0122] Example 3
[0123] Based on the same inventive concept as other embodiments, this embodiment describes a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the methods of Embodiment 1 or 2 described above.
[0124] Example 4
[0125] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including computer instructions that, when executed by a processor, implement the steps of the methods described in Embodiment 1 or 2 above.
[0126] In summary, this invention improves the wavelet threshold function by setting an adjustment factor to ensure its continuity within the wavelet threshold range, and uses this to denoise the engine dataset. Combined with data augmentation from a data balancing model and training of a self-distillation-based autoencoder feature extraction model, it effectively improves the accuracy and reliability of engine fault feature extraction. This solves the problems of difficult fault feature extraction, weak model generalization ability, and low diagnostic accuracy in existing engine fault diagnosis, caused by a small number of available fault data samples, class imbalance, and large data noise interference.
[0127] This invention employs a data balancing model combining a synthetic minority oversampling model and a conditional generative adversarial network model. Through complex training methods, including utilizing Wasserstein distance constraints, pre-defined penalty functions, and dual interpolation strategies to improve the loss function, it augments fault operation data, resulting in a rich and high-quality augmented engine dataset. This high-quality data provides a solid foundation for subsequent fault feature extraction, helping to more accurately capture engine fault characteristics.
[0128] This invention utilizes a pre-trained data balancing model to obtain a confidence prediction function by thresholding the probability values of each label in the student model. This function then corrects the feature extraction error of the teacher model, constructing an improved loss function for the student model. The student model parameters are updated through backpropagation, and iterative training is performed. This self-distillation method allows the model to continuously optimize during training. Training stops when the student model's fault classification accuracy reaches a preset threshold. The resulting self-distillation-based autoencoder feature extraction model can extract engine fault features more accurately, significantly improving the accuracy of fault feature extraction.
[0129] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0133] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for extracting engine fault features based on data augmentation and a self-distillation model, characterized in that, include: Acquire engine datasets that include both normal operation data and fault operation data; The wavelet threshold function is improved by setting an adjustment factor so that the improved wavelet threshold function is continuous within the wavelet threshold range; The engine dataset is denoised using an improved wavelet threshold function to obtain a denoised engine dataset. Based on the denoised engine dataset, data augmentation is performed using a trained data balancing model to obtain a data-augmented engine dataset; the data balancing model includes a synthetic minority oversampling model and a conditional generative adversarial network model connected in sequence. Based on the augmented engine dataset, a self-distillation-based autoencoder feature extraction model is trained to extract engine fault features. The improved wavelet threshold function is expressed as: ; In the formula, Indicates the first The first decomposition level Wavelet estimation coefficients after position filtering Indicates the first The first decomposition level Wavelet estimation coefficients before filtering at each position Represents a symbolic function. Indicates the wavelet threshold range, , Indicates the standard deviation of noise. , This represents the absolute value of the median of the scaling factor. Represented by natural numbers The logarithmic function with base 1. Indicates signal length. Indicates the regulating factor; when hour, ,Right now It exhibits continuity within the wavelet threshold range, where, This indicates that the wavelet threshold is trending towards , This indicates that the wavelet threshold tends negatively to , Indicates the limit; when hour, , ,Right now It exhibits continuity within the wavelet threshold range. It represents positive infinity.
2. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 1, wherein the training method of the data balancing model includes: The fault operation data is augmented sequentially using a synthetic minority oversampling model and a conditional generative adversarial network model to obtain augmented fault operation data. Based on the data-augmented fault operation data, the generator is constrained by Wasserstein distance, and a constrained data balance model loss function is constructed. By using a preset first penalty function, a preset second penalty function, and a dual interpolation strategy to improve the data balance model loss function after constraints, the improved data balance model loss function is obtained. Based on the data-augmented fault operation data and the improved data balancing model loss function, the data balancing model is trained to obtain a well-trained data balancing model.
3. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 2, characterized in that, The loss function used to augment the faulty operational data by sequentially employing a synthetic minority oversampling model and a conditional generative adversarial network model is expressed as: ; In the formula, Indicates the generator when performing data augmentation Take the minimum value Discriminator Take the maximum value The loss value of the loss function used. express The mathematical expectation, This represents the fault operation data after data augmentation. express The mathematical expectation, This represents the noise vector after data augmentation. This represents a real sample of the fault operation data after data augmentation. This indicates the category labels for the data augmented during fault operation. Represents a random noise vector. This represents a conditional discriminator. Represents a condition generator. This represents the logarithmic function with the natural number 10 as the base.
4. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 3, characterized in that, The loss function of the constrained data balancing model is expressed as: ; In the formula, This represents the loss value of the loss function in the constrained data balancing model. Represents the input random noise vector Category labels for data-enhanced fault operation data Condition generator Compared with data-enhanced engine data square Euclidean distance Mathematical expectation The weight, This represents the square of the Euclidean distance.
5. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 4, characterized in that, The improved data balancing model loss function is expressed as follows: ; In the formula, This represents the loss value of the improved data balancing model loss function. This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. The distribution of fault operation data before data augmentation and fault operation data generated by the generator must be obeyed. The mathematical expectation, This represents the interpolated samples of real fault operation data after data augmentation and the fault operation data generated by the generator. The distribution of fault operation data after data augmentation and fault operation data generated by the generator must conform to the data distribution of fault operation data. The mathematical expectation, This represents the interpolated samples of real fault operation data before data augmentation and the fault operation data generated by the generator. gradient operator, This represents the interpolated samples of real fault operation data after data augmentation and the fault operation data generated by the generator. gradient operator, This represents the weight of the penalty function term. and Represent the first penalty function respectively Mathematical expectation Weights and the second penalty function Mathematical expectation The weight.
6. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 5, characterized in that, The method for constructing the autoencoder feature extraction model based on self-distillation includes: Step S1: Based on the trained data balancing model, threshold filtering is performed according to the probability values of each label in the student model to obtain the confidence prediction function; Step S2: Correct the feature extraction error of the teacher model according to the confidence prediction function to obtain the error-corrected teacher model loss function; Step S3: Based on the error-corrected teacher model loss function, construct an improved student model loss function through knowledge distillation; Step S4: Update the parameters of the student model through backpropagation based on the improved student model loss function; Step S5: Use the updated student model as the new teacher model, and repeat steps S2 to S4 for iterative training; When the fault classification accuracy of the student model reaches the preset threshold, the iteration stops, and the current student model is used as the self-distillation-based autoencoder feature extraction model to obtain the constructed self-distillation-based autoencoder feature extraction model.
7. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 6, characterized in that, The confidence prediction function is expressed as: ; In the formula, Student model The value of the confidence prediction function, The superscript indicates the confidence prediction value used to distinguish the values after threshold filtering. Student model The Confidence prediction values before each label threshold filtering This represents the preset confidence threshold. To represent the empty set, Indicates intersection.
8. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 7, characterized in that, The error-corrected teacher model loss function is expressed as follows: ; In the formula, This represents the teacher model loss function after error correction. This represents the teacher model loss function before error correction. Mean square error function The weight, The teacher model before error correction is based on input features. The predicted vector.
9. The engine fault feature extraction method based on data augmentation and self-distillation model according to claim 8, characterized in that, The improved student model loss function is expressed as follows: ; In the formula, This represents the improved student model loss function. This represents the cross-entropy loss of the student model. This represents the ranking loss of the student model. The mean squared error represents the softened probability distribution of the corrected teacher and student models. The weights, where, , The predicted vector representing the softening probability of the student model before error correction. Represents the normalized exponential function, This represents the temperature scaling factor. Indicates the correction term The weighting coefficients, The teacher model after error correction is based on input features. The predicted vector.
Citation Information
Patent Citations
Threshold-function control-type wavelet shrinkage noise eliminator and program
JP2010232892A
Method for denoising underwater acoustic signal on the basis of adaptive window filtering and wavelet threshold optimization
WO2021258832A1