A Collaborative Diagnosis Method for Multiple Faults in Power Systems Based on Dynamic Relative Advantage
By combining adaptive variational encoders and interactive residual automatic encoders, along with mean square error loss and a two-layer optimization algorithm, the problem of unbalanced missing data rates from multiple sensors is solved, achieving high-precision and robust multi-fault diagnosis and improving the operational reliability and maintenance efficiency of the electromechanical composite transmission system.
Patent Information
- Application Number
- CN202511361378.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-09-23
AI Technical Summary
In complex working conditions and extreme environments, the imbalance of missing data rates from multiple sensors prevents traditional methods from effectively utilizing missing data, resulting in a significant degradation in the model's diagnostic capabilities for specific fault types and insufficient robustness.
A multi-fault collaborative diagnosis method based on dynamic relative advantage is constructed. The optimal feature representation is extracted by an adaptive variational encoder, and cross-fault type feature fusion is combined with an interactive residual autoencoder. The feature expression capability is quantified by mean square error loss. A fault feature relative advantage evaluation technique is designed, and a two-layer collaborative optimization algorithm is used to balance training bias and form a closed-loop feedback mechanism.
By effectively utilizing missing data and balancing training biases, the fault diagnosis accuracy and robustness of electromechanical composite transmission systems under complex working conditions have been significantly improved, thereby enhancing the reliability of equipment operation and maintenance efficiency.
Smart Images

Figure CN120849810B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis technology for electromechanical composite transmission systems, and more specifically to a multi-fault collaborative diagnosis method for power units based on dynamic relative advantage drive. Background Technology
[0002] As a core component of aviation, military equipment, and high-end industrial equipment, the stable operation of the power system directly affects the reliability of the equipment and the safety of the mission. However, under complex operating conditions and extreme environments, problems such as sensor degradation, acquisition equipment failure, or transmission link interruption often lead to an imbalance in the missing data rate of multi-sensor data (e.g., the missing data rate for some fault types is as high as 80%, while that for other types is only 20%). Traditional methods, which directly mask missing data or employ inefficient reconstruction strategies based on a single feature space, not only waste the potential information hidden in the missing data but also cause the model's diagnostic ability for specific fault types to deteriorate significantly because strong channels (low missing rate) dominate training while weak channels (high missing rate) are ignored.
[0003] To address the aforementioned challenges, the deep integration of artificial intelligence and big data technologies provides an innovative path for the intelligent processing of multi-source heterogeneous data. Deep learning models, with their nonlinear feature hierarchical learning capabilities, exhibit unique advantages in fault diagnosis with imbalanced data missing rates: for example, variational information bottleneck theory compresses redundant features by minimizing information loss, enabling the extraction of generalized representations across fault types from incomplete samples of missing data; generative adversarial networks compensate for missing information through data generation mechanisms, but existing methods often ignore the differences in supervision between different fault types; while attention mechanisms can dynamically allocate feature weights, they struggle to quantify the training bias caused by differences in missing rates.
[0004] This invention focuses on the challenge of imbalanced multi-sensor data loss rates in power plants caused by data gaps. It constructs an intelligent diagnostic framework combining a "cross-fault type generalized joint diagnostic strategy + fault feature relative advantage assessment." This framework extracts the optimal feature representation for each fault type through an adaptive variational encoder, and combines label supervision to reconstruct and utilize missing data, overcoming the dependence of traditional methods on complete data. A fault feature relative advantage assessment technique is designed, quantifying feature representation capabilities based on mean square error loss to dynamically identify strong / weak advantage faults. A two-layer collaborative optimization algorithm with supervised weight adaptive adjustment is constructed. The inner layer trains the model using weighted loss, while the outer layer dynamically adjusts the supervision intensity based on relative advantage, forming a closed-loop feedback loop of "advantage assessment - weight adjustment - model optimization" to balance training bias caused by differences in loss rates. This framework overcomes the performance bottleneck of traditional methods under imperfect information. Through big data-driven feature fusion and adaptive adjustment of artificial intelligence algorithms, it significantly improves the accuracy and robustness of multi-fault diagnosis under complex operating conditions, providing a deep collaborative technical paradigm of "data-algorithm-model" for intelligent operation and maintenance of power plants. Summary of the Invention
[0005] In view of this, the present invention provides a collaborative diagnosis method for multiple faults of power units based on dynamic relative advantage driving, which improves the fault diagnosis accuracy of electromechanical composite transmission systems under the condition of unbalanced data loss rate of multiple sensor sources.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A collaborative diagnosis method for multiple faults in a power unit based on dynamic relative advantage includes the following steps:
[0008] Obtain the raw dataset of multiple fault types and multiple sensors, mark the missing state of each data point, and perform preprocessing.
[0009] The optimal feature representations of different sensors for each fault type in the preprocessed standardized dataset are extracted using an adaptive variational encoder.
[0010] Based on the optimal feature representation, it is determined whether the fault data of each type is missing. If it is missing, the information bottleneck loss is minimized based on the interactive residual autoencoder, and the cross-sensor feature fusion is used to make up for the missing fault data by using the optimal features of other fault types. The missing segments are completed by label supervision reconstruction, and the generalized feature of the fault type is output. If it is complete, the quality of feature extraction is constrained by mean square error supervision, and the generalized feature is output.
[0011] After obtaining the generalization features, the feature representation ability of each fault type is quantified by mean squared error loss, and the relative advantage is defined.
[0012] A two-layer collaborative optimization is performed. When optimizing the outer layer, the supervision weights of each fault type are adjusted based on the relative advantage. When optimizing the inner layer, a multi-fault diagnosis model is trained based on the current supervision weights to minimize the weighted loss. Dynamic adaptive gradient constraints are used to avoid abnormal parameter updates.
[0013] After optimized training, the generalized features are decoded by the decoder to reconstruct the original signal, and the fault label is predicted based on the reconstructed original signal to output the fault type.
[0014] Preferably, the adaptive variational encoder adopts a Gaussian variational form:
[0015]
[0016] in, Indicates a given and When, the optimal feature representation The conditional probability distribution, This represents the set of parameters for an adaptive variational encoder. express Follows the mean f μ (·), covariance is f Σ (·) Gaussian distribution, Let represent the mean function of the adaptive variational encoder, which outputs a Gaussian distribution mean vector with dimension d. Let represent the covariance function of the adaptive variational encoder, and let represent the covariance matrix of the Gaussian distribution, with dimensions d×d. This represents the optimal feature representation. Indicates original features.
[0017] Preferably, the information bottleneck loss function is approximated by a variational upper bound:
[0018]
[0019] in, Represents the information bottleneck loss function. express The variational upper limit, K represents the total number of fault types (including auxiliary faults), Represents the relationship between random variables (original features) (Optimal feature representation) Expected calculation of (label), The KL divergence represents the difference between two distributions, measuring the difference between the encoded feature distribution and the prior distribution. Denotes a b-dimensional spherical Gaussian, that is Among them 0 b and I b They are a b-dimensional all-zero vector and a b×b identity matrix, respectively. The function notation used to define the Gaussian distribution describes The distribution pattern, where γ represents the balance coefficient, adjusts the intensity of tag monitoring. Indicates the decoder's input to the tag. Conditional distribution, This represents the decoder parameters.
[0020] Preferably, after obtaining the generalized features, the feature representation ability of each fault type is quantified by mean squared error loss, and the relative advantage is defined, including:
[0021] The characteristic representation ability of each fault type is quantified by mean squared error loss, and the mean squared error loss of K types of faults is averaged as a benchmark to measure the performance of each fault type.
[0022]
[0023] in, This represents the mean squared error loss for the k-th fault type. This represents the average value of the mean squared error loss;
[0024] Based on the mean squared error loss average, the relative advantage of fault types is defined.
[0025]
[0026] Determining strong-dominance and weak-dominance faults based on relative advantage:
[0027] when At that time, fault k is considered a dominant fault relative to other faults; The interpolation loss of fault k is comparable to the average value, indicating an equilibrium state; when When the interpolation loss of fault k is greater than the average loss, it means that this fault type performs poorly during training and belongs to the weak-dominant fault category.
[0028] Preferably, during inner-layer optimization, a multi-fault diagnosis model is trained based on the current supervision weights to minimize the weighted loss:
[0029]
[0030] in, Represents the set of model parameters, θ k The supervisory weights represent different faults. Represents the information bottleneck loss function. This represents the mean square error loss for the k-th fault type.
[0031] Preferably, during inner-layer optimization, to ensure that each parameter update does not disrupt model training due to abnormally large gradients, gradient constraints are introduced into the inner-layer optimization process. Each time a gradient is calculated, dynamic gradient smoothing is performed first, followed by the application of gradient constraints.
[0032]
[0033] in, Represents the gradient operator, applied to the set of model parameters. Find the partial derivative, L clipped This represents the loss function after gradient clipping. The parameter subset of the k-th fault type iteration, ||·||2 represents the L2 norm, and the "length" of the gradient vector is calculated. MAL Indicates mixed loss c represents the gradient norm threshold, which prevents gradient explosion and ensures stable parameter updates.
[0034] Preferably, during outer layer optimization, the supervision weights for each fault type are adjusted based on relative advantage, and the objective function is:
[0035]
[0036] in, Let ξ1 denote the relative advantage vector, which represents the lower bound of the weights to ensure that each fault type has at least basic supervision and avoid completely ignoring a certain type of fault. Let ξ2 denote the upper bound of the weight norm to ensure weight normalization and maintain the rationality of the loss function. Let p denote L1 regularization to achieve sparse weight adjustment and highlight the supervision of key fault types.
[0037] Meanwhile, the updates to the weights θ follow these rules:
[0038] in, This means projecting the updated weights onto the constraint range, where α represents the learning rate and controls the update step size; θ (t) θ represents the supervision weight vector for each fault type at the t-th iteration, indicating the weight configuration used for fault diagnosis supervision in the current iteration step; (t+1) This represents the updated supervision weight vector at the (t+1)th iteration, which is based on the update rule for θ. (t) The result after adjustment and projection onto the constraint range.
[0039] Preferably, since the training process is affected by stochastic gradients, changes in error may lead to unstable model performance. Therefore, after adjusting the supervision weights for each fault type based on relative advantage during outer layer optimization, the method also includes smoothing error fluctuations using a moving average loss.
[0040]
[0041] in, This represents the mixed loss after moving average smoothing at the g-th iteration. This represents the mixed loss associated with the k-th type of fault in the (g-1)-th iteration, and is the historical error term used to calculate the moving average. β represents the mean squared error loss of the k-th type of fault in the g-th iteration, representing the error situation of this fault type in the current iteration, and is the current error term in the moving average; β represents the smoothing coefficient, which controls the weight of historical error and current error.
[0042] Preferably, the interactive residual auto encoder includes four residual blocks, each with the following structure:
[0043]
[0044] Among them, RA lThe representation is the output obtained by calculating the l-th residual in the interactive residual autoencoder. The MLP is a multilayer perceptron, and the output dimension is transformed layer by layer from 256 to 128 to 64 to 32 to 64 to 128 to 256, finally generating a 128-dimensional cross-fault type representation.
[0045] As can be seen from the above technical solution, compared with the prior art, the present invention has the following beneficial effects:
[0046] 1) Effective Utilization of Missing Data: Existing technologies often directly mask missing data or employ inefficient reconstruction strategies, wasting the potential information hidden within the missing data. This invention utilizes a cross-fault type generalized joint diagnosis strategy based on variational information bottleneck theory, combined with label supervision, to reconstruct incomplete samples of missing data. This achieves effective utilization of missing data, avoids the excessive reliance on complete data in traditional methods, and improves data utilization.
[0047] 2) Balancing Training Bias: Traditional methods often suffer from imbalanced missing data rates across multiple sensors, leading to strong channels dominating training while weak channels are neglected, resulting in a degradation in the model's diagnostic capabilities for specific fault types. The fault feature relative advantage evaluation technique designed in this invention quantifies feature representation capabilities based on mean squared error loss, dynamically identifies strong / weak advantage faults, and dynamically adjusts the supervision intensity of each fault type according to relative advantage through a two-layer collaborative optimization algorithm, forming a closed-loop feedback loop. This effectively balances the training bias caused by differences in missing data rates, making the model's diagnostic capabilities more balanced across different fault types.
[0048] 3) Improved diagnostic accuracy and robustness: Existing methods tend to suffer from decreased diagnostic accuracy and insufficient robustness under complex operating conditions with unbalanced data missing rates. This invention extracts the optimal feature representations for each fault type through an adaptive variational encoder, achieves cross-fault type feature fusion using an interactive residual autoencoder, and combines techniques such as dynamic adaptive gradient constraints and moving average loss smoothing to effectively compress redundant information, avoid abnormal parameter updates and model instability, and significantly improve the accuracy and robustness of multi-fault collaborative diagnosis under complex operating conditions, providing more reliable technical support for the intelligent operation and maintenance of electromechanical composite transmission systems.
[0049] In summary, the technical solution of this invention covers the entire process from data preprocessing and feature extraction to model optimization, forming a complete technical chain of "data reconstruction - feature evaluation - weight adjustment". Compared with existing technologies, this invention has significant technical advantages in data utilization, diagnostic balance, and model robustness. Even in complex application scenarios such as high missing rates, multiple fault coupling, and unbalanced data distribution, this invention can still maintain high diagnostic accuracy and robustness. Furthermore, this invention has important practical significance for improving equipment reliability, reducing maintenance costs, and ensuring safety in industrial scenarios. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0051] Figure 1 A flowchart of a multi-fault collaborative diagnosis method for a power unit based on dynamic relative advantage driving provided by the present invention;
[0052] Figure 2 The reasoning flowchart for the cross-fault type generalized joint diagnosis strategy provided by this invention;
[0053] Figure 3 This is a structural diagram of the interactive residual automatic encoder provided by the present invention;
[0054] Figure 4 The figure shows the experimental verification results of the ablation of the key model of the method of this invention.
[0055] Figure 5 This figure shows a comparison of the average F1 score between the method of this invention and existing methods for data missing reconstruction and data imbalance optimization. Detailed Implementation
[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0057] This invention discloses a collaborative diagnosis method for multiple faults in a power unit based on dynamic relative advantage, comprising four main parts: data preprocessing and labeling, construction of a cross-fault type generalized joint diagnosis strategy, design of fault feature relative advantage evaluation technology, and a two-layer collaborative optimization algorithm.
[0058] like Figures 1-2 As shown, it specifically includes:
[0059] Obtain the original dataset of multiple fault types and multiple sensors, including missing data. First, perform noise reduction and standardization preprocessing on the dataset, and mark the missing status of each data point ("incomplete" or "completely missing", where "incomplete" means that the sensor data is partially missing but retains valid fragments, and "completely missing" means that the sensor data is completely empty). Construct a standardized dataset of multiple fault types and multiple sensors, including K types of faults, with each fault type containing l sensor data.
[0060] For each type of fault, an adaptive variational encoder is used to generate the optimal feature representation for the corresponding fault type, completing the initial mapping from the original data to the feature space.
[0061] Based on the optimal feature representation, it is determined whether the fault data for each type is missing. If it is missing, the interactive residual autoencoder process is entered. The process of minimizing information bottleneck loss is executed in sequence to compress redundancy and retain key information. Cross-sensor feature fusion is used to make up for the missing fault data by using the optimal features of other fault types. The missing segments are completed by combining label supervision reconstruction and outputting the generalized features of the fault type. If it is complete, the quality is extracted directly by means of square error supervision constraint features and the generalized features are output. Finally, after all fault types are processed by the corresponding branches, a unified set of generalized features is output, which provides standardized input for subsequent fault relative advantage assessment and two-layer optimization diagnosis.
[0062] After obtaining the generalization features, the feature representation ability of each fault type is quantified by mean squared error loss, and the relative advantage is defined.
[0063] The system employs a two-layer collaborative optimization algorithm for model training. During outer-layer optimization, the supervision weights for each fault type are adjusted based on relative advantages to ensure balanced learning across different fault types during training. Inner-layer optimization, under the current supervision weights, trains the multi-fault diagnosis model by minimizing the weighted loss. To ensure training stability and effectiveness, the inner-layer optimization introduces dynamic adaptive gradient constraints to prevent gradient explosion and employs moving average loss smoothing to reduce the impact of error fluctuations. Through this collaborative approach, the model parameters are continuously optimized, enhancing the model's diagnostic capabilities for various fault types.
[0064] After optimized training, the trained model is used to decode the generalized features and reconstruct the original signal, including the completion results of missing data, thereby restoring the complete data space. Based on the reconstructed original signal, fault label prediction is performed, and the probability distribution of fault types is output through a classifier to achieve accurate diagnosis of multiple fault types.
[0065] The entire process of this invention forms a closed-loop feedback mechanism. After each training iteration, the relative advantages of fault features are reassessed based on the model performance, the supervision weights are adjusted, the model parameters are further optimized, and the fault diagnosis accuracy and robustness of the system under complex working conditions are continuously improved, providing a strong guarantee for the reliable operation of the electromechanical composite transmission system.
[0066] The steps described above in this invention will be further explained below.
[0067] I. Data Preprocessing and Labeling
[0068] 1. Multiple sensors (such as vibration sensors, temperature sensors, current sensors, etc.) collect raw data from different monitoring points of specialized equipment (such as electromechanical composite transmission systems, industrial motors, etc.). These data contain characteristic information of the equipment under normal operation and various fault conditions. Due to factors such as sensor hardware failure, transmission link interruption, or sudden changes in operating conditions, data loss is inevitable (e.g., vibration data at a certain moment is not collected, or the temperature sensor completely fails for a certain period of time).
[0069] 2. Noise Reduction and Standardization Preprocessing
[0070] Noise reduction: Wavelet denoising is used to remove noise interference from the original data. For example, for vibration sensor data, the signal is decomposed to different scales using wavelet transform, and high-frequency noise components are thresholded before reconstruction to restore the pure fault characteristic signal. The denoised data can more accurately reflect the true operating status of the equipment and provide reliable input for subsequent feature extraction.
[0071] Standardization: Z-score standardization is performed on the data from each sensor, mapping the data to a uniform scale with a mean of 0 and a variance of 1. This step eliminates the impact of sensor range and unit differences on model training, making data from different sensors comparable in the feature space and improving the model's ability to fuse multi-source data.
[0072] 3. Dataset Construction and Missing Type Labeling
[0073] Dataset Construction: A dataset containing K types of faults (such as bearing wear, gear cracks, motor overload, etc.) is constructed, with each type of fault corresponding to monitoring data from l sensors. The dataset is divided into training, validation, and test sets according to time series to ensure the model's generalization ability under different data distributions.
[0074] "Incomplete" data: Sensor data is partially missing (e.g., only the first 50% of the vibration signal was collected within a certain period, and the remaining 50% was lost due to transmission interruption), but valid segments are retained (e.g., key vibration peaks at the time of the fault). This type of data is recovered by an interactive residual autoencoder to recover the missing parts for feature learning.
[0075] "Completely missing" data: Sensor data is completely empty for a certain period of time (e.g., a temperature sensor has no output for 10 consecutive minutes due to hardware failure). This type of data requires cross-sensor inference based on the correlation features of other sensors to avoid interruption of model training due to missing data.
[0076] II. Construction of a Cross-Fault Type Generalized Joint Diagnosis Strategy
[0077] 1. Feature extraction from an adaptive variational encoder
[0078] For the original characteristics of the l-th sensor in the k-th type of fault First, the optimal feature representation for each fault type is extracted using an adaptive variational encoder with shared parameters. This encoder uses a Gaussian variational form to map the original data to a low-dimensional space, generating the optimal feature representation:
[0079]
[0080] in, Represents the given original features and encoder parameters When, the optimal feature representation The conditional probability distribution, This represents the set of parameters for an adaptive variational encoder. express Follows the mean f μ (·), covariance is f Σ (·) Gaussian distribution, Let represent the mean function of the adaptive variational encoder, which outputs a Gaussian distribution mean vector with dimension d. Let represent the covariance function of the adaptive variational encoder, which outputs a Gaussian-distributed covariance matrix with dimension d×d.
[0081] The mean function captures typical patterns of fault characteristics (such as the mean vibration frequency of bearing faults), while the covariance function reflects the range of characteristic fluctuations (such as the variance change of current when a motor is overloaded). This step compresses redundant information, retains key fault characteristics, provides a concise and robust input for cross-fault fusion, and reduces noise interference.
[0082] 2. Interactive residual autoencoder cross-fault fusion
[0083] The interactive residual autoencoder receives the outputs of the adaptive variational encoder for each fault type and achieves cross-fault feature interaction through residual connections (e.g., feature fusion of bearing and gear faults, and mining of composite fault associations). Redundancy is compressed using information bottle loss, where information divergence constrains the difference between the features and the prior distribution, and a balance coefficient adjusts the label supervision strength to ensure that the fused features retain key information related to the fault labels (e.g., filtering normal operating noise and enhancing fault discrimination). The specific expression is shown in the following formula:
[0084]
[0085] in, Represents the information bottleneck loss function. express The variational upper limit, where K represents the total number of fault types (including auxiliary faults), Represents the relationship between random variables (original features) (Optimal feature representation) Expected calculation of (label), The KL divergence represents the difference between two distributions, measuring the difference between the encoded feature distribution and the prior distribution. Denotes a b-dimensional spherical Gaussian, that is Among them 0 b and I b They are a b-dimensional all-zero vector and a b×b identity matrix, respectively. The function notation used to define the Gaussian distribution describes Distribution pattern, γ balance coefficient, and adjustment of tag monitoring intensity. Indicates the decoder's input to the tag. Conditional distribution, This represents the decoder parameters.
[0086] This process enhances the correlation between cross-fault features, enabling the model to learn the commonalities and unique characteristics of multiple faults, improving its generalization ability to unknown faults. The interactive residual autocoding structure, such as... Figure 3 As shown. The processing flow is as follows:
[0087] Feature Input and Adaptation: After the adaptive variational encoders for each fault type complete the initial feature extraction, they use the output feature data as the input to the interactive residual autoencoder to prepare for cross-fault feature interaction.
[0088] Layer-by-layer processing (taking a single residual block as an example):
[0089] LayerNorm normalization: for optimal feature representation Performing the LayerNorm operation eliminates differences in feature distribution, making the data distribution more stable, laying a solid foundation for subsequent training and feature processing, and avoiding the impact of distribution fluctuations on model learning.
[0090] MLP (Multilayer Perceptron) Transformation: Through dimensional transformation "256→128→64→32→64→128→256", it first compresses and then expands to mine multi-level features, extract key information from different dimensions and depths, and capture fault feature details.
[0091] LayerNorm normalization is performed again: the output features of the MLP are normalized a second time to further regularize the distribution, ensure feature consistency, and facilitate subsequent residual connections and feature fusion.
[0092] Residual connection: representing the optimal feature The residual connection is added to the normalized output features from the previous MLP step to construct a residual connection. This preserves the previous feature information, alleviates gradient vanishing, and helps the model learn more complex fault associations. Residual block output is generated: After the above operations, a single residual block output is obtained, which serves as the input to the next residual block, continuously passing and processing features.
[0093] Multi-residual block cascaded iteration: Four residual blocks are cascaded sequentially, with the output of the previous residual block directly serving as the input for the next, continuously iterating the features. Each round deepens cross-fault feature interactions (such as bearing and gear fault feature fusion) and gradually uncovers the correlations between complex faults.
[0094] Final output and fault fusion value: After processing through 4 residual blocks, the output is a 128-dimensional fault feature. In this process, redundancy is compressed by combining information bottleneck loss, the difference between the feature and the prior distribution is constrained by information divergence, and the label supervision strength is adjusted by the balance coefficient to ensure that the fused feature retains the key information of the fault label (filtering normal operating noise and enhancing fault discrimination power), which helps the model learn the commonalities and individualities of multiple faults and improves the generalization ability of unknown faults.
[0095] 3. Missing data reconstruction and complete data supervision
[0096] For missing data (incomplete samples), combine label supervision. (e.g., fault type) The interactive residual autoencoder reconstructs partially missing data. By minimizing information bottleneck loss and label error, it ensures that the completed features are consistent with the actual fault modes, maximizing the utilization of missing data.
[0097] For complete data (no missing samples), mean squared error (MSE) loss is introduced to supervise reconstruction accuracy. MSE calculates the error between the reconstructed features and the original data, ensuring that the optimal features retain key details, avoiding the loss of important information during compression, and guaranteeing feature quality. Mean squared error loss:
[0098]
[0099] in, L represents the mean squared error loss, used to supervise the reconstruction accuracy of complete data (no missing sensor features). k This represents the number of complete data samples in the k-th type of fault. This represents the generalization feature of the l-th complete sample in the k-th type of fault after decoding by an interactive residual autoencoder. It represents the square of the Euclidean norm, calculates the mean square value of the reconstruction error, and measures the difference between the optimal feature representation and the generalized feature.
[0100] III. Fault Feature Relative Advantage Assessment Technology Design
[0101] 1. The role and calculation of mean square error loss
[0102] First, for each type of fault (e.g., the kth type), complete data samples (i.e., sensor feature data without missing data) are reconstructed using an interactive residual autoencoder, and then the mean square error loss is calculated. The loss value directly reflects the feature representation ability of fault type k: the smaller the loss, the closer the reconstructed features are to the original data, the more complete the features are preserved, and the better the feature learning effect of this fault type (for example, a small reconstruction error of the vibration signal of a bearing fault indicates that the model has a strong ability to capture its features). Conversely, a larger loss indicates that there are defects in feature representation (such as severe missing sensor data leading to a large reconstruction error, indicating that the model has not learned the fault sufficiently).
[0103] 2. Definition of Relative Advantage Indicators
[0104] Mean square error calculation: The mean square error loss of all K types of faults is averaged to serve as a benchmark for measuring the performance of each fault type.
[0105]
[0106] Relative advantage indicators This indicator intuitively quantifies the strength of fault k relative to other faults. It is defined by comparing the MSE of the current fault k with the average MSE:
[0107]
[0108] Determining strong-dominance and weak-dominance faults based on relative advantage:
[0109] Strong advantage fault The MSE of fault k is lower than the average, indicating that its feature representation ability is outstanding (such as motor overload fault, the data is complete and the features are obvious, and the model is easy to learn).
[0110] Equilibrium MSE equals the mean, indicating that the feature representation ability is at a moderate level (e.g., slight wear on gears, low data missing rate, and stable reconstruction error).
[0111] Weak dominance fault The MSE is higher than the average, indicating weak feature representation ability (e.g., sensor failure leads to severe data loss, large reconstruction error, and difficulty in model learning).
[0112] This invention quantifies the feature representation effect of each fault type by calculating the mean squared error loss and defines a relative advantage index accordingly. This index can identify which fault types are advantageous and which are disadvantageous in feature representation. Based on the evaluation results, the system dynamically adjusts the supervision weights of each fault type. Fault types with strong feature representation capabilities are given higher supervision weights to receive more attention during training; for fault types with weak feature representation capabilities, the weights are adjusted appropriately to prevent them from being ignored, thus providing a reasonable weight allocation for subsequent optimization training.
[0113] IV. Construction of a Two-Layer Collaborative Optimization Algorithm
[0114] 1. Inner layer optimization:
[0115] Weighted loss training: The inner layer minimizes the mixed loss (information bottleneck loss + weighted mean squared error loss) based on the current supervision weights, and adjusts the model parameters accordingly.
[0116]
[0117] in, Represents the set of model parameters, θ k The supervisory weights represent different faults. Represents the information bottleneck loss function. This represents the mean square error loss for the k-th fault type.
[0118] Inner layer optimization enables the model to allocate learning resources more reasonably for different fault types. For example, it can focus on learning weak dominant faults (high weight) to make up for their insufficient data quality.
[0119] Gradient Constraints: A threshold c is introduced to control the gradient norm and avoid gradient explosion during parameter updates. When the gradient exceeds c, it is scaled proportionally to ensure training stability. This mechanism improves model convergence speed by 30% under high missing rate and noisy data, avoids the "gradient explosion" problem of traditional methods, and enhances training robustness.
[0120] Gradient constraint techniques are used to avoid abnormal parameter updates:
[0121]
[0122] in, Represents the gradient operator, applied to the set of model parameters. Find the partial derivative, L clipped This represents the loss function after gradient clipping. The parameter subset of the k-th fault type iteration, ||·||2 represents the L2 norm, L MAL Indicates mixed loss c represents the gradient norm threshold, which prevents gradient explosion and ensures stable parameter updates.
[0123] 2. Outer layer optimization:
[0124] Objective Function and Constraints: The outer layer dynamically updates the supervision weights by minimizing the objective function, combined with upper and lower bounds on the weights (ensuring basic supervision and normalization) and L1 regularization (sparse adjustment, highlighting key faults). The weights of strong-dominant faults are reduced (to avoid overfitting), while the weights of weak-dominant faults are increased (to improve learning priority), achieving a balanced allocation of training resources. The objective function is:
[0125]
[0126] in, Let ξ1 denote the relative advantage vector, which represents the lower bound of the weights to ensure that each fault type has at least basic supervision and avoid completely ignoring a certain type of fault. Let ξ2 denote the upper bound of the weight norm to ensure weight normalization and maintain the rationality of the loss function. Let p denote L1 regularization to achieve sparse weight adjustment and highlight the supervision of key fault types.
[0127] Weight update rules: The projection operator is used to constrain the weights within a reasonable range, and the learning rate α controls the update step size to ensure smooth weight adjustment. The update of weight θ follows these rules:
[0128]
[0129] in, This means projecting the updated weights onto the constraint range, where α represents the learning rate, controls the update step size, and θ... (t) θ represents the supervision weight vector for each fault type at the t-th iteration, indicating the weight configuration used for fault diagnosis supervision in the current iteration step; (t+1) This represents the updated supervision weight vector at the (t+1)th iteration, which is based on the update rule for θ. (t) The result after adjustment and projection onto the constraint range.
[0130] 3. Dynamic monitoring and error smoothing
[0131] Moving average loss: Since the training process is affected by stochastic gradients, changes in error may cause model instability. After adjusting the supervision weights for each fault type based on relative advantage during outer layer optimization, moving average loss is also used to smooth error fluctuations.
[0132]
[0133] in, This represents the mixed loss after moving average smoothing at the g-th iteration. This represents the mixed loss associated with the k-th type of fault in the (g-1)-th iteration, and is the historical error term used to calculate the moving average. β represents the mean squared error loss of the k-th type of fault in the g-th iteration, representing the error situation of this fault type in the current iteration, and is the current error term in the moving average; β represents the smoothing coefficient, which controls the weight of historical error and current error.
[0134] Closed-loop feedback: forming a closed loop of "advantage assessment → weight adjustment → model optimization" to continuously balance training bias.
[0135] To verify the practicality of the method of the present invention, an experiment was conducted using the fault diagnosis of the electromechanical composite transmission system of a certain equipment power unit as a case study: multi-channel signal data was collected through a simulation test bench, the data covered multiple operating conditions of the equipment under different health states, and five scenarios of information loss and imbalance were designed.
[0136] The experimental data acquisition and processing are as follows: The signal contains 9 channels, specifically the triaxial acceleration signal (6 channels) from the drive end (DE) and the fan end (FE) and the three-phase current signal (3 channels), with a sampling frequency of 25.6kHz; a non-overlapping sliding window of length 1024 is used to extract samples, and a total of 1200 samples are extracted for each motor operating speed, which are divided into training set and test set in an 8:2 ratio; the experiment simulates 8 motor health states, covering normal state and various typical fault types. Through data acquisition and design under multi-dimensional and imperfect information conditions, a real-world scenario is provided to verify the effectiveness of the method of this invention in handling the problem of data missing rate imbalance under complex working conditions.
[0137] To verify the superiority of the multi-fault collaborative diagnosis method for power units based on dynamic relative advantage, an ablation study was designed to compare the following three scenarios: 1) lack of a generalized joint diagnosis method across fault types (Method A); 2) lack of a relative advantage awareness monitoring mechanism (Method B); 3) the method of this invention. The three methods were validated under the same network structure and under three different imbalanced conditions of missing rates for outer ring faults, inner ring faults, and rolling element faults. Quantitative results are as follows: Figure 4 As shown.
[0138] Experimental results show that Method A, lacking a generalized joint diagnostic strategy across fault types, can only directly mask or inefficiently utilize missing data. It cannot extract optimal feature representations for each fault type through an adaptive variational encoder, nor can it effectively utilize incomplete samples with missing data through label supervision. Consequently, its diagnostic classification accuracy under different imbalanced missing data rates is significantly lower than that of the method in this study. This demonstrates that a generalized joint diagnostic strategy across fault types is crucial for addressing the problem of imbalanced missing data rates in multi-sensor data and effectively utilizing the potential information within missing data.
[0139] Method B, lacking a relative advantage awareness monitoring mechanism, cannot quantify the feature representation ability of different fault types through mean squared error loss, nor can it dynamically adjust the supervision intensity of each fault type based on relative advantage. It also struggles to balance training bias caused by differences in missing rates, resulting in a diagnostic classification accuracy that, while higher than Method A, is still lower than the method used in this study. This indicates that a relative advantage awareness monitoring mechanism plays a crucial role in balancing training bias and preventing overfitting of strong-advantage faults and neglect of weak-advantage faults.
[0140] This invention combines a generalized joint diagnosis strategy across fault types with a relative advantage awareness monitoring mechanism. Through a closed-loop feedback loop of "advantage assessment - weight adjustment - model optimization," it effectively overcomes the adverse effects of data missing rate imbalance on fault diagnosis of electromechanical composite transmission systems. Experimental data show that the diagnostic classification accuracy of this method is significantly higher than that of Method A and Method B under three different data missing rate imbalance conditions. This verifies the effectiveness of the synergistic effect of the generalized joint diagnosis strategy across fault types and the relative advantage awareness monitoring mechanism, as well as the significant improvement in the accuracy of multi-fault collaborative diagnosis under complex operating conditions.
[0141] The experiment uses a cross-fault type generalized joint diagnostic framework as the basic architecture. Its basic structure is shown in Table 1, and the training parameters are shown in Table 2.
[0142] Table 1. Summary of Basic Network Structure
[0143]
[0144] Table 2. Summary of Training-Related Parameters
[0145] Parameter name set up Number of iterations T=30 Learning rate η = 0.005 Optimizer Adam Loss weights α=0.2
[0146] To further verify the performance of the method of this invention, experiments were conducted to compare it with three existing fault diagnosis methods for imbalanced data missing rates. These three methods include: Method C, which integrates complete data with uncertain information through dynamic Bayesian networks to compensate for data integrity issues and minimize prediction errors; Method D, which proposes a novel fuzzy clustering framework that incorporates non-negative latent factor analysis and feature-weighted fuzzy double C-means clustering for high-precision clustering of incomplete data; and Method E, which proposes a novel data augmentation method based on a Deo diffusion probability model for fault diagnosis of data with imbalanced missing rates.
[0147] The comparison results show that... Figure 5 As shown, 1) When the average missing rate of multi-sensor data is low (i.e., missing rates of 0.0 and 0.1), the proposed method outperforms all comparative models. This is because other models are not fully exposed to missing data and therefore struggle to cope with such situations, while this method can continue to perform representation learning across fault types even without missing data. 2) When the average missing rate of multi-sensor data is at a moderate level (e.g., between 0.2 and 0.6), despite the less data-intensive supervised training mechanism used in this invention, its performance is at least comparable to other models, verifying the advantage of this method in supervising missing data. 3) When the average missing rate of multi-sensor data is high (e.g., between 0.7 and 0.9), this method still maintains relatively stable performance with only a slight to moderate performance decline, while the comparative models generally show a significant performance decline. The proposed method outperforms the three existing methods in terms of classification accuracy, macro-average F1 score, and generalization ability. This indicates that the proposed method can exhibit higher robustness and diagnostic efficiency under conditions of missing data and imbalanced missing rates.
[0148] Through the above experiments and comparative analyses, the practicality and superiority of the method of this invention in addressing complex fault diagnosis problems can be demonstrated. By combining a cross-fault type generalized joint diagnosis strategy with fault feature relative advantage evaluation technology, the method of this invention can effectively solve the common problem of data mismatch in practical applications, exhibiting high application value and technical advantages.
[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0150] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for multi-fault collaborative diagnosis of a power plant based on dynamic relative advantage driving, characterized in that, The method comprises the following steps: Obtain a multi-fault type multi-sensor original data set, and mark the missing state of each piece of data and perform preprocessing; Extract the optimal feature representation of each fault type different sensor in the preprocessed standardized data set through an adaptive variational encoder; Determine whether each class of fault data is missing based on the optimal feature representation, if missing, minimize the information bottleneck loss based on an interactive residual autoencoder, cross-sensor feature fusion compensates for the current fault data missing by means of other fault type optimal features, and the missing segment is reconstructed and completed in combination with label supervision, and the generalization features of the fault type are output; If complete, the quality of feature extraction is constrained by mean square error supervision, and the generalization features are output; After obtaining the generalization features, the feature expression ability of each fault type is quantified by mean square error loss, and the relative advantage is defined; Carry out double-layer collaborative optimization, adjust the supervision weight of each fault type based on the relative advantage during outer optimization, train the multi-fault diagnosis model based on the current supervision weight during inner optimization, minimize the weighted loss, and adopt dynamic adaptive gradient constraint to avoid parameter update anomaly; After optimization training, the generalization features are decoded by the decoder to reconstruct the original signal, and the fault label prediction is carried out based on the reconstructed original signal, and the fault type is output; After obtaining the generalization features, the feature expression ability of each fault type is quantified by mean square error loss, and the relative advantage is defined, including: Quantify the feature expression ability of each fault type by mean square error loss, and average the mean square error losses of K fault types as the benchmark for measuring the performance of each fault type; wherein, represents the mean square error loss of the kth fault type, represents the mean square error loss average value; Based on the mean value of the mean square error loss, define the relative advantage of the fault type Determine strong advantage faults and weak advantage faults based on the relative advantage: When the fault k is considered as a strong dominant fault with respect to the other faults; the interpolation loss of the fault k is comparable to the average value, belonging to the balanced state; when the interpolation loss of the fault k is greater than the average loss, belonging to the weak dominant fault; During outer optimization, adjust the supervision weight of each fault type based on the relative advantage, and the objective function is: wherein, denotes the relative advantage vector, denotes the lower bound of the weight norm, denotes the upper bound of the weight norm, and p denotes the L1 regularization. weights The update of the weights follows the following rules: wherein, denotes the projection of the updated weights into the constraint range, a denotes the learning rate, and controls the update step size; denotes the supervision weight vector of each fault type at the tth iteration, and represents the weight configuration used for fault diagnosis supervision in the current iteration step; denotes the updated supervision weight vector at the t+1th iteration, and is the result of adjusting and projecting into the constraint range according to the update rule adjusting and projecting into the constraint range according to the update rule 2. The method of claim 1, wherein, The adaptive variational encoder adopts a Gaussian distribution variational form: wherein, denotes the optimal feature representation and denotes the conditional probability distribution of the optimal feature representation denotes the parameter set of the adaptive variational encoder, denotes is subject to a Gaussian distribution with mean μ (·) and covariance Σ (·), denotes the mean function of the adaptive variational encoder, outputs the mean vector of the Gaussian distribution, dimension d, denotes the covariance function of the adaptive variational encoder, outputs the covariance matrix of the Gaussian distribution, dimension d x d, denotes the optimal feature representation, denotes the original feature.
3. The method of claim 1, wherein, The information bottleneck loss is calculated by variational upper bound approximation: wherein, represents the information bottleneck loss function, represents the variational upper bound, K represents the total number of failure types, represents the expectation calculation over the random variable represents the label, represents the optimal feature representation, represents the original feature, represents the KL divergence between two distributions, represents the b-spherical Gaussian, i.e. where 0 b and I b are the b-dimensional all-zero vector and b x b identity matrix, respectively, the function symbol used to define the Gaussian distribution, describes the distribution law of, γ represents the balance coefficient, adjusts the label supervision intensity, represents the conditional distribution of the label by the decoder, represents the decoder parameter. 4. The method of claim 1, wherein, During inner optimization, train the multi-fault diagnosis model based on the current supervision weight, and minimize the weighted loss: wherein, denotes a set of model parameters, denotes a supervision weight for different faults, denotes an information bottleneck loss function, denotes a mean square error loss for the kth fault type.
5. The method of claim 1, wherein, Dynamic adaptive gradient constraint is to perform dynamic smoothing of gradient at each time of calculating gradient, and then apply gradient constraint: wherein, denotes the gradient operator, and denotes the partial derivative with respect to the parameter set clipped denotes the loss function after gradient clipping, denotes the parameter subset of the kth failure type iteration, and MAL denotes the hybrid loss denotes the gradient norm threshold.
6. The method of claim 1, wherein, After adjusting the supervision weight of each fault type based on the relative advantage during outer optimization, it also includes smoothing error fluctuation by moving average loss: wherein, represents the mixed loss after moving average smoothing at the gth iteration, represents the mixed loss related to the kth fault at the (g-1)th iteration, is the historical error term used for calculating the moving average, represents the mean square error loss of the kth fault at the gth iteration, represents the error situation of the fault type in the current iteration, is the current error term in the moving average; β represents the smoothing coefficient, which controls the weight of the historical error and the current error.
7. The method of claim 1, wherein, The interactive residual autoencoder comprises four residual blocks, and the structure of each residual block is: wherein RA l represents the output calculated from the lth residual in the interactive residual autoencoder, MLP is a multi-layer perceptron, the output dimension is transformed layer by layer through 256-128-64-32-64-128-256, and finally a 128-dimensional cross-fault type representation is generated, represents the original feature.
Citation Information
Patent Citations
Gravitational wave detection link fault rapid diagnosis method based on machine learning
CN114330012A
Scene generation method for automatic driving
CN114926712A