Fault diagnosis method under imperfect information condition based on orthogonal spatial heterogeneous mapping
By employing the orthogonal space heterogeneous mapping method, and utilizing cross-domain attention mechanism and generative adversarial network for data reconstruction and label self-correction, the problems of incomplete data and missing labels in the fault diagnosis of key components of power systems are solved, achieving high-precision and high-reliability fault diagnosis.
Patent Information
- Application Number
- CN202511384031.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-26
AI Technical Summary
In the fault diagnosis of key components of the power system, there are problems such as incomplete data and missing or incorrect labels, which cause traditional diagnostic methods to fail and make it difficult to maintain high accuracy and reliability under complex conditions.
We employ a method based on orthogonal spatial heterogeneous mapping, which uses a cross-domain attention mechanism and generative adversarial networks to construct embedded multidimensional feature vectors for data reconstruction and label self-correction. By combining spatial feature information matching and frequency differential projection, we improve the comprehensiveness and accuracy of fault diagnosis.
It significantly improves the accuracy and reliability of power system fault diagnosis, maintains high performance in the event of missing data and incorrect labeling, reduces operation and maintenance costs, and enhances the safety and reliability of equipment operation.
Smart Images

Figure CN120873531A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault diagnosis technology for key components of power systems, and more specifically to a fault diagnosis method based on imperfect information conditions of orthogonal space heterogeneous mapping. Background Technology
[0002] As a core component of special equipment, the stability and safety of key parts of the power system are crucial for the operation of critical fields such as aviation, military, and special industries. However, in actual operation, fault diagnosis technology for key components of the power system faces many challenges. First, due to factors such as sensor degradation, data acquisition equipment failure, or multi-rate sampling, the acquired operational data is often missing or incomplete. This data incompleteness not only increases the complexity of fault diagnosis but may also cause traditional diagnostic methods to fail. Second, high-quality fault diagnosis models usually rely on a large amount of labeled data, but data labeling requires a lot of expert experience and resources, resulting in low efficiency and high cost. In addition, there are problems with missing or incorrect labels in some scenarios, which further limits the promotion and application of traditional data-driven methods.
[0003] To address the aforementioned issues, the rapid development of artificial intelligence and big data technologies has provided novel solutions for fault diagnosis of key components in power systems. In particular, intelligent algorithms based on deep learning, generative adversarial networks, and attention mechanisms have demonstrated significant advantages in handling high-dimensional data, incomplete information, and complex label correction. These methods can automatically extract deep-level features from multi-dimensional, multi-source data and reduce reliance on complete data and accurate labels through data reconstruction and label self-correction strategies. Fault diagnosis frameworks integrating these technologies can not only efficiently address practical problems such as missing data and label errors but also significantly improve diagnostic accuracy and reliability under complex operating conditions.
[0004] Therefore, how to provide a fault diagnosis method based on imperfect information under the condition of heterogeneous mapping of orthogonal space is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the present invention provides a fault diagnosis method under imperfect information conditions based on orthogonal spatial heterogeneous mapping. An enhanced orthogonal spatial heterogeneous network with embedded multidimensional feature vectors is constructed through a cross-domain attention mechanism and a generative adversarial network to effectively reconstruct incomplete information. Furthermore, spatial feature information matching and spatial feature frequency differential projection are used to intelligently correct the labels, which significantly improves the comprehensiveness and accuracy of power system fault diagnosis and provides strong technical support for the safe operation of special equipment.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] Fault diagnosis methods based on imperfect information conditions under orthogonal space heterogeneous mapping include:
[0008] S1: Determine the integrity of the faulty dataset, classifying it into complete dataset, incomplete dataset, and missing dataset;
[0009] S2: The complete dataset is input into the cross-domain attention mechanism network, and the output of the cross-domain attention mechanism network is input into the hierarchical fusion multidimensional fault feature extraction network. After processing, the data feature encoder output is formed.
[0010] S3: Input the incomplete dataset, missing dataset, and the output of the data feature encoder into the enhanced orthogonal spatial heterogeneous network with embedded multidimensional generator to efficiently reconstruct the data missing and incomplete problems, and obtain the reconstructed complete dataset. The reconstructed complete dataset and the complete dataset together constitute the standard dataset. The standard dataset is then input into the hierarchical fusion multidimensional fault feature extraction network based on the cross-domain attention mechanism to generate data features related to different fault types and sensors.
[0011] S4: Spatial information location encoding is performed on the data features obtained after processing the standard dataset to form a spatial relative location encoding system for different fault types and sensors;
[0012] S5: A label self-correction strategy based on spatial differentiation is proposed. By utilizing the spatial relative position coding system of different fault types and sensors, spatial information feature matching is performed on different fault types and sensors. Spatial feature projection is then performed on the matched fault types and sensor information to obtain the final fault classification result.
[0013] Preferably, the cross-domain attention mechanism network includes a channel attention network and a spatial attention network. The channel attention network includes three parallel branches, which are concatenated and then configured with convolutional layers, activation functions, and convolutional layers. The first branch includes a convolutional layer, an asymmetric convolutional block, batch normalization, and an activation function. The second branch includes a convolutional layer, two asymmetric convolutional blocks, batch normalization, and an activation function. The third branch includes a convolutional layer and an average pooling layer. The spatial attention network includes an average pooling layer and a max pooling layer, which are concatenated and then configured with multiple asymmetric convolutional blocks and a sigmoid activation function.
[0014] The cross-domain attention mechanism network processing procedure is as follows:
[0015] ;
[0016] in, This represents the output of the cross-domain attention mechanism network. This represents the learnable weight parameters. Represents sensor data, This represents element-wise multiplication. Represents the ReLU activation function. This represents an asymmetric convolution block with kernel size k. This represents a 1×1 convolution operation. This indicates the average pooling operation. This indicates a max pooling operation. This represents the Sigmoid activation function. This represents the concatenation operation of features. This indicates a summation operation across all feature channels. This represents a non-linear activation function.
[0017] Preferably, the output of the cross-domain attention mechanism network is fed into the hierarchical fusion multidimensional fault feature extraction network, specifically including:
[0018] The output of the cross-domain attention mechanism network is divided into four sensor data streams. Each sensor data stream enters the hierarchical fusion multidimensional fault feature extraction network, which includes multi-branch convolution, convolutional pooling, and multi-branch pooling.
[0019] The multi-branch convolution process is as follows:
[0020] ;
[0021] in, This represents the output of a multi-branch convolutional block. BN and BP represent the data normalization and batch normalization operations of the multi-branch convolutional layer, respectively. This indicates a max pooling operation. f = 1, 3, represents the standard convolution of f × f. Represents convolution sum, This represents an asymmetric convolution operation with a kernel size of 3 and g repetitions.
[0022] The output after convolutional pooling is calculated as follows:
[0023] ;
[0024] in, It is the output after convolutional pooling. This represents an asymmetric convolution operation with a kernel size of 1, repeated 3 times.
[0025] The multi-branch pooling process is as follows:
[0026] ;
[0027] in, It is a multi-branch pooling block output. Indicates that the average pooling operation is repeated. Second-rate.
[0028] Preferably, the enhanced orthogonal spatial heterogeneous network with an embedded multidimensional generator in S3 comprises two parts: a hybrid generator and a regularized discriminator.
[0029] The hybrid generator uses a recurrent neural network model to capture the time-series features of the data, combines a variational autoencoder to map the data to the latent space and generate new data, and then optimizes the feature representation and network training through the SE module and identity mapping residual connection.
[0030] Initialization is performed using the hidden states of the recurrent neural network model:
[0031] ;
[0032] in, This represents the initial hidden state of the recurrent neural network (RNN) model. These are the input parameters for the recurrent neural network model. This indicates the initialization of the weight matrix;
[0033] The recurrent neural network model takes incomplete datasets, missing datasets, and the output of the data feature encoder as input, performs recursive processing, and utilizes the hidden state from the previous time step. and the latent variables at the current moment Calculate the hidden state at the current time step. :
[0034] ;
[0035] in, This represents the weight matrix of the hidden layer, responsible for performing weighted transformations on the hidden state from the previous time step. This represents the weight matrix of the input layer, used to weight the latent variables at the current time step. Weighting; It is the ReLU activation function. These are the parameters of the recurrent neural network model. For bias terms;
[0036] Through such recursive calculations, the recurrent neural network model gradually captures the time-series features of the data. Finally, by summing the time-series information, it obtains... As the output generated by the recurrent neural network model:
[0037] ;
[0038] in, For the integrated time series information, N represents the total number of time steps processed by the recurrent neural network;
[0039] The data processed by the recurrent neural network model is mapped to the latent space Z through a variational autoencoder, and new data is generated from the latent space using a decoder. The operation process of the variational autoencoder is as follows:
[0040] ;
[0041] in, This represents the set of variables and parameters of a variational autoencoder. These are variables in the latent space, representing the feature representation of the input data in the latent space after encoding. It is the parameter set in a variational autoencoder. This represents the weights of the encoder portion of the variational autoencoder. This indicates the bias of the encoder portion in a variational autoencoder; It is sampling noise, generated by a standard normal distribution. Represents a non-linear activation function. Indicates the weights of the decoder. The weights and biases of the decoder are used to decode the variables m in the latent space to generate new data.
[0042] By employing a variational autoencoder, the mapping from the original feature representation to the latent space and the generation of new data are achieved, further mining data features and increasing data diversity. The generated feature map is fed into the SE module, which uses a fully connected layer to adjust the importance of each channel using weighted coefficients, making the model more focused on key features. The formula is as follows:
[0043] ;
[0044] in, This indicates the input feature map The process involves calculating the feature map after channel weight adjustment. This represents the generator's output at the current moment. This represents the parameter set of the SE module. This represents the weight matrix of the fully connected layers in the SE module. It is a bias term. This indicates a global average pooling operation on the input features, where S is the Sigmoid activation function.
[0045] The features processed by the SE module are further optimized through identity mapping residual connections, as shown in the expression:
[0046]
[0047] in, Represents the residual of the identity mapping. It is the output of the generator from the previous time step, which is filtered by noise. and input features Generate data, These are the parameters of the corresponding network layer; the residual block combines the generator's output with its input. Adding the results to form residuals helps alleviate the vanishing gradient problem, accelerates network training and optimization, and ensures that the generator learns and generates data stably and efficiently in multiple rounds of computation.
[0048] In the regularized discriminator, the feature encoder encodes incomplete and missing datasets to obtain feature representations. The final feature encoding representation is obtained by weighted summation:
[0049] ;
[0050] in: This represents the feature representations generated by different feature encoders. These are the weight coefficients of each feature encoder, representing the importance of different features in the discrimination. It is the total number of feature encoders. This represents the final feature encoding representation. For the set of all relevant parameters;
[0051] The regularized discriminator performs gradient calculations on incomplete and missing datasets. The gradient calculation formula is as follows: ;
[0052] in, Indicates input data Find the gradient. It is the feature representation function in gradient calculation-related operations. These are the corresponding parameters;
[0053] The true / false probabilities are obtained by activating the input features using the sigmoid activation function.
[0054] :
[0055] in, It is the output of the regularization discriminator, representing the data. Let S be the probability of the real data, and S be the sigmoid activation function, ensuring that the output value is between 0 and 1. It is the regularization coefficient, used to control the normalization of the gradient. It is a hyperparameter used to control the size of the regularization term.
[0056] Preferably, S4 includes:
[0057] Computation of spatial location coding function Its expression is:
[0058] ;
[0059] ;
[0060] in, Represents a spatial location encoding function. It is a characteristic representation of the Fourier transform. The weight matrix represents the features. It is a sensor characteristic. It is the standard deviation of the Gaussian kernel. This refers to the number of position matrices considered. For fault type index, For dimensional indexing, This represents the total number of vector dimensions;
[0061] Based on the spatial location coding function results, further calculations are made of discriminative features between different sensors and between different fault types. Calculated using the following formula:
[0062] ;
[0063] in, Indicates position-encoded input. This represents a multilayer perceptron. Represents a linear unit with Gaussian error. Indicates the sensor state at the previous moment;
[0064] Different types of faults The calculation is related to intermediate variables. Related, calculate first :
[0065] ;
[0066] in, This represents the feature alignment vector between sensor features and the fault prototype. Represents the spatial characteristics of the sensor. Indicates the characteristics of the fault prototype;
[0067] Then through Calculate different fault types :
[0068]
[0069] Here, the Multilayer Perceptron (MLP) extracts and transforms features from the input data, while GELU introduces nonlinear characteristics into the computation results, making... and It can better distinguish the characteristics of different sensors and fault types.
[0070] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions, which has the following beneficial effects:
[0071] First, this invention designs an innovative hierarchical fusion multidimensional fault feature extraction network based on a cross-domain attention mechanism, which can effectively capture multidimensional features, including time-frequency domain features, spatial features, and features fused from multiple sensors. This network dynamically allocates feature weights through the combined action of channel and spatial attention mechanisms, enabling the model to extract key fault-related information more accurately. Simultaneously, by introducing asymmetric convolution and multi-branch convolution strategies, it not only improves the model's adaptability to complex fault modes but also significantly reduces computational resource consumption, overcoming the problem of feature loss that traditional feature extraction methods easily cause when processing multidimensional complex data, thereby improving the reliability and efficiency of diagnosis.
[0072] Secondly, this invention proposes an enhanced orthogonal spatial heterogeneous network with an embedded multidimensional generator for efficient data reconstruction addressing the problems of missing and incomplete data. The hybrid generator combines the advantages of recurrent neural networks and variational autoencoders, while introducing an identity mapping residual block to optimize gradient propagation, enhancing the model's training depth and stability. The regularized discriminator further improves the ability to discriminate the authenticity of generated data by integrating a cross-domain multi-scale feature extraction encoder and a gradient penalty mechanism. This generative adversarial network design overcomes the problem of unstable training caused by gradient vanishing or excessively deep network structures in traditional adversarial networks, and significantly improves the quality and feature completeness of generated data, providing more reliable data support for fault diagnosis.
[0073] Furthermore, this invention offers significant advantages in addressing label missing and errors. It proposes a label self-correction strategy based on spatially heterogeneous features, accurately capturing the spatial feature relationships between sensors and the correlation characteristics between fault types. Building upon this, a spatial information encoding system between sensors and fault types is established by aggregating multidimensional features through self-attention and cross-attention mechanisms. Further, by combining cosine similarity and a bidirectional soft assignment strategy, the network can accurately identify and correct erroneous labels while efficiently filling in missing labels. Moreover, the spatially differentiated feature frequency projection network improves the stability and accuracy of label correction by optimizing feature offsets, thus enabling fault diagnosis to maintain high performance even in complex data environments.
[0074] Finally, this invention designs a comprehensive loss optimization strategy. By introducing the synergistic effect of reconstruction loss, matching loss, and differentiation loss, it significantly improves the model's learning ability and fault diagnosis performance. The reconstruction loss optimizes the matching degree between generated and target data for frequency domain data, ensuring high fidelity in data reconstruction. The matching loss improves the matching efficiency between features through joint optimization of the soft allocation matrix and true matching pairs. The differentiation loss enhances the discriminative power of feature representation by correcting the offset of feature projection. These innovative design features enable the model to more efficiently recover complete data, improving the accuracy and stability of diagnosis.
[0075] In summary, the technical solution of this invention covers the entire process of fault data classification, preprocessing, feature extraction, data augmentation, label self-correction, and classification decision-making, providing an efficient and robust fault diagnosis method. Compared with existing technologies, this invention has significant technical advantages in handling data incompleteness, the accuracy of feature extraction, and label correction capabilities. Even in complex application scenarios such as missing data, incorrect labels, or abnormal sensor acquisition, this invention can still maintain high diagnostic accuracy and robustness. Furthermore, this invention has important practical significance for improving equipment reliability, reducing maintenance costs, and ensuring safety in industrial settings. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0077] Figure 1 The flowchart of the fault diagnosis method based on imperfect information under the orthogonal spatial heterogeneous mapping provided by the present invention is shown below.
[0078] Figure 2 The network flowchart of the cross-domain attention mechanism provided by this invention;
[0079] Figure 3 The multi-branch convolutional structure in the hierarchical fusion multidimensional fault feature extraction network provided by this invention;
[0080] Figure 4 The convolutional pooling structure in the hierarchical fusion multidimensional fault feature extraction network provided by this invention;
[0081] Figure 5 This invention provides a multi-branch pooling structure in a hierarchical fusion multidimensional fault feature extraction network.
[0082] Figure 6 A schematic diagram of an enhanced orthogonal spatial heterogeneous network with an embedded multidimensional generator provided by the present invention;
[0083] Figure 7 This invention provides a spatially differentiated label self-correction network.
[0084] Figure 8 The figure shows the experimental verification results of the ablation of the key model of the method of this invention.
[0085] Figure 9 This diagram illustrates a comparison of the learning performance of the method of this invention with existing incomplete fault information reconstruction and tag self-correction methods. Detailed Implementation
[0086] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0087] This invention discloses a fault diagnosis method based on imperfect information conditions using heterogeneous mapping in orthogonal spaces, such as... Figure 1 As shown, it includes:
[0088] S1: The integrity of the fault dataset is assessed, categorized into complete, incomplete, and missing datasets. An incomplete dataset refers to a situation where some sensor monitoring data is missing, or some labels are missing. This typically covers cases where some data or labels are available, but fragments are missing due to sensor limitations, i.e., incomplete data is characterized by partial missing sensor data or labels. Conversely, a missing dataset refers to a situation where the sensor failed to detect any data, or all labels for the monitoring data are lost. Compared to incomplete data, missing data encompasses all missing data. In the model of this invention, the simultaneous loss of data and labels is not considered. Preferably, preprocessing operations such as data standardization, pruning, splicing, and segmentation are performed on the complete input data for different fault types.
[0089] S2: The complete dataset is input into the cross-domain attention mechanism network, which assigns spatial domain weights and channel domain weights to data from different sensors. The output of the cross-domain attention mechanism network is input into the hierarchical fusion multidimensional fault feature extraction network, and after processing, it forms the output of the data feature encoder.
[0090] S3: The incomplete dataset, the missing dataset, and the output of the data feature encoder are input into an enhanced orthogonal spatial heterogeneous network embedded with a multidimensional generator to efficiently reconstruct the data for missing and incomplete problems, resulting in a reconstructed complete dataset. The reconstructed complete dataset and the original complete dataset together constitute the standard dataset. The standard dataset is then input again into a cross-domain attention mechanism network and a hierarchical fusion multidimensional fault feature extraction network to generate data features for different fault types and sensors.
[0091] S4 encodes the spatial information location of the features obtained after processing the standard dataset, forming a spatial relative location coding system for different fault types and sensors.
[0092] S5 proposes a label self-correction strategy based on spatial differentiation. By utilizing the spatial relative position coding system of different fault types and sensors, it first performs spatial information feature matching between different fault types and sensors; then, it projects spatial features onto the matched fault types and sensor information; finally, it obtains the fault classification result.
[0093] In this embodiment, the preprocessed complete dataset enters the cross-domain attention mechanism network, where weights are assigned in the channel and spatial dimensions to emphasize key features and suppress secondary information. For example... Figure 2 As shown, this structure processes the input data sequentially through a channel attention network and a spatial attention component network. The channel attention network comprises three parallel branches, which are concatenated and then configured with convolutional layers, activation functions, and another convolutional layer. The first branch includes a convolutional layer, an asymmetric convolutional block, batch normalization, and an activation function; the second branch includes a convolutional layer, two asymmetric convolutional blocks, batch normalization, and an activation function; and the third branch includes a convolutional layer and an average pooling layer. The spatial attention network includes an average pooling layer and a max pooling layer, which are concatenated and then configured with multiple asymmetric convolutional blocks and a sigmoid activation function.
[0094] The cross-domain attention mechanism network processing procedure is as follows:
[0095] ;
[0096] in, This represents the output of the cross-domain attention mechanism network. This represents the learnable weight parameters. Represents sensor data, This represents element-wise multiplication. Represents the ReLU activation function. This represents an asymmetric convolution block with kernel size k. This represents a 1×1 convolution operation. This indicates the average pooling operation. This indicates a max pooling operation. This represents the Sigmoid activation function. This represents the concatenation operation of features. This indicates a summation operation across all feature channels. This represents a non-linear activation function.
[0097] To extract features for each fault type in depth, the output of the cross-domain attention mechanism network is fed into a hierarchical fusion multidimensional fault feature extraction network, specifically including:
[0098] The hierarchical fusion multidimensional fault feature extraction network includes multi-branch convolution, convolutional pooling, and multi-branch pooling;
[0099] The output of the cross-domain attention module is divided into four sensor data streams, and each sensor data stream enters the hierarchical fusion multidimensional fault feature extraction network.
[0100] Multi-branch convolutional structure diagram as follows Figure 3 As shown, the processing procedure is as follows:
[0101] ;
[0102] in, This represents the output of a multi-branch convolutional block. BN and BP represent the data normalization and batch normalization operations of the multi-branch convolutional layer, respectively. This indicates a max pooling operation. f = 1, 3, represents the standard convolution of f × f. Represents convolution sum, This represents an asymmetric convolution operation with a kernel size of 3 and g repetitions.
[0103] Convolutional pooling structure diagram as shown Figure 4 After the above is shown, the output calculation is as follows:
[0104] ;
[0105] in, It is the output after convolutional pooling. This represents an asymmetric convolution operation with a kernel size of 1, repeated 3 times.
[0106] Multi-branch pooling structure diagram as follows Figure 5 As shown, the processing procedure is as follows:
[0107] ;
[0108] in, It is a multi-branch pooling block output. Indicates that the average pooling operation is repeated. Second-rate.
[0109] The data feature encoder formed above will be fed into an enhanced orthogonal spatial heterogeneous network with an embedded multidimensional generator. Identity mapping residual blocks will be introduced, and the data reconstruction process will be optimized by integrating fault feature frequencies and adversarial losses into a reconstruction loss function, thereby enhancing data integrity. Detailed reasoning processes are as follows: Figure 6 As shown.
[0110] The enhanced orthogonal spatial heterogeneous network with embedded multidimensional generator mainly consists of two parts: a hybrid generator and a regularized discriminator. The detailed steps of the two parts are as follows:
[0111] The core idea of the hybrid generator is to use recurrent neural networks to capture the time-series features of data, combine variational autoencoders to map the data to the latent space and generate new data, and then optimize the feature representation and network training through SE modules and identity mapping residual connections.
[0112] Initialization is performed using the hidden states of the recurrent neural network model:
[0113] ;
[0114] in, This represents the initial hidden state of the recurrent neural network (RNN) model. These are the input parameters for the recurrent neural network model. This indicates the initialization of the weight matrix;
[0115] The recurrent neural network model takes an incomplete dataset, a missing dataset, and a data feature encoder obtained through a hierarchical fusion multidimensional fault feature extraction network as input, performs recursive processing, and utilizes the hidden state from the previous time step. and the latent variables at the current moment Calculate the hidden state at the current time step. :
[0116] ;
[0117] in, This represents the weight matrix of the hidden layer, responsible for performing weighted transformations on the hidden state from the previous time step. This represents the weight matrix of the input layer, used to weight the latent variables at the current time step. Weighting; It is the ReLU activation function. These are the parameters of the recurrent neural network model. For bias terms;
[0118] Through such recursive calculations, the RNN gradually captures the time-series characteristics of the data. Finally, by summing the time-series information, it obtains... As the output of the recurrent neural network model, it provides a time-series-based feature representation for subsequent processing:
[0119] ;
[0120] in, The integrated time series information is the output generated by the recurrent neural network model, where N represents the total number of time steps processed by the recurrent neural network.
[0121] The data processed by the recurrent neural network is mapped to the latent space Z using a variational autoencoder, and new data is generated from the latent space using a decoder. The operation process of the variational autoencoder is as follows:
[0122] ;
[0123] in, This represents the set of variables and parameters of a variational autoencoder. These are variables in the latent space, representing the feature representation of the input data in the latent space after encoding. It is the parameter set in a variational autoencoder. This represents the weights of the encoder portion of the variational autoencoder. This indicates the bias of the encoder portion in a variational autoencoder; It is sampling noise, generated by a standard normal distribution. Represents a non-linear activation function. Indicates the weights of the decoder. This represents the weights and biases of the decoder, used to decode variables m in the latent space to generate new data.
[0124] By employing a variational autoencoder, the mapping from the original feature representation to the latent space and the generation of new data are achieved, further mining data features and increasing data diversity. The generated feature map is fed into the SE module, which uses a fully connected layer to adjust the importance of each channel using weighted coefficients, making the model more focused on key features. The formula is as follows:
[0125] ;
[0126] in, This indicates the input feature map The process involves calculating the feature map after channel weight adjustment. This represents the generator's output at the current moment. This represents the parameter set of the SE module, which includes parameters such as weights and biases of components such as fully connected layers within the module. This represents the weight matrix of the fully connected layers in the SE module. It is a bias term. This indicates a global average pooling operation on the input features, where S is the Sigmoid activation function.
[0127] The features processed by the SE module are further optimized through identity mapping residual connections, as shown in the expression:
[0128] ;
[0129] in, Represents the residual of the identity mapping. It is the output of the generator from the previous time step, which is filtered by noise. and input features Data generation; residual blocks add the generator's output to its input to form residuals, which helps alleviate the gradient vanishing problem, accelerates network training and optimization, and ensures that the generator learns and generates data stably and efficiently in multiple rounds of computation;
[0130] In the regularized discriminator, the feature encoder encodes incomplete and missing datasets to obtain feature representations. The final feature encoding representation is obtained by weighted summation:
[0131] ;
[0132] in: This represents the feature representations generated by different feature encoders. These are the weight coefficients of each feature encoder, representing the importance of different features in the discrimination. It is the total number of feature encoders. This represents the final feature encoding representation. For the set of all relevant parameters;
[0133] Next, the discriminator performs gradient calculations on the incomplete and missing datasets to normalize the features and ensure gradient stability. This process involves the following gradient calculation formula: .
[0134] in, Indicates input data Find the gradient. It is the feature representation function in gradient calculation-related operations. These are the corresponding parameters.
[0135] The true / false probabilities are obtained by activating the input features using the sigmoid activation function.
[0136]
[0137] in, It is the output of the regularization discriminator, representing the data. Let S be the probability of the real data, and S be the sigmoid activation function, ensuring that the output value is between 0 and 1. It is the regularization coefficient, used to control the normalization of the gradient. It is a hyperparameter used to control the size of the regularization term.
[0138] In the data reconstruction stage, to ensure that the frequency characteristics of the generator's output data match the target requirements, the overall reconstruction loss function is defined as follows, combining adversarial loss and mean squared error loss based on frequency characteristics:
[0139]
[0140] Where G represents the hybrid generator, which is responsible for generating simulated data; D is the regularization discriminator, which is used to distinguish between real and fake data; x represents real data, and z is the noise vector input to the generator; and These are respectively the expectation of the real data and the noise vector; and It is the log probability of the regularized discriminant's judgment of whether the data is true or false; F is the function that transforms the data into the frequency domain; This is the generator output data. Its frequency domain characteristics, It is the target frequency characteristic; Measuring the difference between the two is a hyperparameter that weighs the proportion of loss due to differences in frequency characteristics.
[0141] By working in tandem with a hybrid generator and a regularized discriminator, this enhanced orthogonal spatial heterogeneous network can effectively reconstruct missing data, significantly improve data integrity, provide high-quality data support for fault diagnosis, and ultimately enhance the performance and reliability of the diagnostic model.
[0142] Through the above data reconstruction, incomplete and missing information is input into the hybrid generator to form a reconstructed complete dataset, which together with the original complete dataset constitutes the standard dataset. In order to better achieve self-correction of mislabeled and missing labels, the standard dataset will be re-entered into the cross-domain attention mechanism network and the hierarchical fusion multidimensional fault feature extraction network to form a standard dataset feature encoder, which will serve as the input sequence for the spatial information location encoding architecture between different sensors and fault types.
[0143] The specific components of the spatial relative position coding system for different fault types and sensors include:
[0144] First, calculate the spatial location encoding function. Its expression is:
[0145] ;
[0146] ;
[0147] in, Represents a spatial location encoding function. It is a characteristic representation of the Fourier transform. The weight matrix represents the features. It is a sensor characteristic. It is the standard deviation of the Gaussian kernel. This refers to the number of position matrices considered. For fault type index, For dimensional indexing, This represents the total number of vector dimensions;
[0148] Based on the spatial location coding function results, discriminative features are further calculated between different sensors and between different fault types. Discriminative features Calculated using the following formula:
[0149] ;
[0150] in, Indicates position-encoded input. This represents a multilayer perceptron. Represents a linear unit with Gaussian error. Indicates the sensor state at the previous moment;
[0151] Fault type related characteristics The calculation is related to intermediate variables. Related, calculate first :
[0152]
[0153] Then through Calculate different fault types :
[0154]
[0155] The Multilayer Perceptron (MLP) here performs feature extraction and transformation on the input data, while the GELU activation function introduces nonlinear characteristics into the computation results, making... and It can better distinguish the characteristics of different sensors and fault types.
[0156] Furthermore, such as Figure 7 As shown, in order to complete and classify missing labels, spatial information feature matching is performed on different fault types and sensors based on different fault types and sensor location coding systems. Specifically, this includes:
[0157] Based on the discriminative features between different sensors and different fault types Calculate the similarity matrix:
[0158] ;
[0159] in, Represents the similarity matrix. Represents the dot product of vectors. Describes the Euclidean norm of a vector. , ;
[0160] Convert the similarity matrix into a soft-assignment matrix. :
[0161] ;
[0162] ;
[0163] ;
[0164] in, and These represent the similarity matrices after row and column softmax processing, respectively. The correlation index between different fault types and different sensors is represented by the soft allocation matrix, which measures and weights the relationship between different sensors and fault types by time-frequency domain feature similarity.
[0165] based on Computational spatial information feature matching:
[0166] ;
[0167] in, This represents the set of high-confidence pairings. This represents the threshold for each matching pair. ,if This indicates a high-confidence match;
[0168] The mutual information maximization strategy is adopted to further improve the quality and consistency of feature matching:
[0169] ;
[0170] in, This represents the set of pairs after further filtering, ensuring the quality and consistency of the pairs by maximizing the optimization of information.
[0171] Furthermore, to further improve label accuracy, spatial differential feature projection is performed on the matched fault types and sensor information to obtain fault classification results. The specific steps are as follows:
[0172] For the time-frequency domain characteristics F of sensor o and sensor p o and F p Perform a preliminary matching operation to obtain matching pairs. Based on this, the offset is calculated. And perform spatial feature projection:
[0173] ;
[0174] in, This means projecting the information from the 0th sensor onto the pth sensor;
[0175] Calculate the correction vector based on the results of spatial feature projection. :
[0176] ;
[0177] in, and These are weighting coefficients used to balance offset and feature differences;
[0178] Based on the calculated correction vector Update matching pairs: .
[0179] Furthermore, a matching loss is introduced during the label self-correction process. and differentiation loss To ensure the quality of data reconstruction and label self-correction, thereby improving the accuracy of fault diagnosis, the specific process is as follows:
[0180] Calculate the ground truth matching exponent and combine it with the soft allocation matrix. Calculate the matching loss, defined as the spatial information feature matching loss:
[0181] ;
[0182] in, This represents the set of indices for ground truth matching, i.e., the set of indices for correct matches. This represents the balancing factor, used to balance the weights of positive and negative samples. This represents a adjustment factor used to adjust the model's focus on difficult-to-classify samples. This indicates the total number of matching pairs;
[0183] Based on the differential offset and the corrected geospatial offset, the differential loss is... Defined as:
[0184] ;
[0185] in, It is a differential offset. It is the corrected geodetic offset, and 𝐾 is the number of differential projections.
[0186] To verify the practicality of the method of this invention, a case study of fault diagnosis of a key component of a power system was conducted. Multi-channel signal data was collected from a fault simulation test bench for a certain type of traction motor. This data covers various operating conditions of the motor under different health states, and four scenarios of information imperfection were designed to simulate common problems such as missing data and incomplete labels in actual fault diagnosis.
[0187] The experiment collected signals from nine channels at a sampling frequency of 25.6 kHz, corresponding to triaxial acceleration signals from the drive end (DE) and fan end (FE), as well as three-phase current signals. The dataset was extracted from the signals using a sliding window of length 1024, with no overlap between windows. A total of 1200 sample data points were extracted at each motor operating speed, with 80% used for model training and 20% for testing model performance. Furthermore, the experiment simulated eight different motor health states.
[0188] To verify the superiority of spatial differential feature projection technology and its self-correcting effect in synergy with spatial information feature matching, three comparative methods were designed for testing: Method A only introduces spatial information feature matching; Method B directly uses spatial differential feature projection but skips spatial information feature matching; the method of this invention combines the synergistic self-correcting mechanism of spatial information feature matching and spatial differential feature projection. These three methods were validated under the same network structure and four experimental conditions, and the quantitative results are as follows: Figure 8 As shown.
[0189] Experimental results show that Method A, which only introduces spatial information feature matching, has lower classification accuracy and various evaluation indicators than the method of this invention. This indicates that relying solely on spatial information feature matching is insufficient to handle the problems of missing and incorrect labels. In contrast, the method of this invention uses spatial differential feature projection technology to calculate the balance offset between fault features in the time and frequency domains, generating a correction vector to effectively correct missing and incorrect labels, thereby significantly improving classification accuracy. This verifies the effectiveness of spatial differential feature projection technology in addressing label errors and missing data.
[0190] On the other hand, compared with Method B, the method of this invention demonstrates superior performance in classification accuracy, macroscopic average F1 score, and ranking loss. Although Method B employs spatially differentiated feature projection, its classification accuracy is not stable enough because it skips spatial information feature matching. The method of this invention significantly improves the robustness and diagnostic accuracy of the model through a collaborative self-correction mechanism, further demonstrating the superiority of collaborative self-correction.
[0191] To further verify the performance of the method of this invention, an experiment was conducted to compare it with four existing fault diagnosis methods for incomplete information. These four methods include one based on a fuzzy clustering framework (Method C), one based on a dynamic Bayesian network (Method D), one based on a convolutional neural network (Method E), and one based on a three-way decision model (Method F). The experiment used a hierarchical fusion multidimensional fault feature extraction network as the basic architecture, the basic structure of which is shown in Table 1, and the training parameters are shown in Table 2. The experimental results are summarized in... Figure 9 .
[0192] Table 1. Summary of Basic Network Structure
[0193]
[0194] Table 2. Summary of Training-Related Parameters
[0195]
[0196] The comparative results show that the method of this invention outperforms the three existing methods in terms of classification accuracy, macro-average F1 score, and generalization ability. This indicates that the method of this invention can exhibit higher robustness and diagnostic efficiency under imperfect information conditions such as missing data and label errors. In particular, the self-correction mechanism combining spatial feature matching and differential feature projection enables the method of this invention to more effectively correct erroneous labels and recover high-quality data, thereby significantly improving diagnostic accuracy.
[0197] Through the above experiments and comparative analysis, the practicality and superiority of the method of this invention in dealing with complex fault diagnosis problems can be seen. Combining the collaborative self-correction technique of spatial information feature matching and differential feature projection, the method of this invention can effectively solve the common problems of missing information and incomplete labels in practical applications, and has high application value and technical advantages.
[0198] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0199] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fault diagnosis method based on imperfect information conditions under orthogonal space heterogeneous mapping, characterized in that, include: S1: Determine the integrity of the faulty dataset, classifying it into complete dataset, incomplete dataset, and missing dataset; S2: The complete dataset is input into the cross-domain attention mechanism network, and the output of the cross-domain attention mechanism network is input into the hierarchical fusion multidimensional fault feature extraction network. After processing, the data feature encoder output is formed. S3: Input the incomplete dataset, missing dataset, and the output of the data feature encoder into the enhanced orthogonal spatial heterogeneous network with embedded multidimensional generator to efficiently reconstruct the data missing and incomplete problems, and obtain the reconstructed complete dataset. The reconstructed complete dataset and the complete dataset together constitute the standard dataset. The standard dataset is then input into the hierarchical fusion multidimensional fault feature extraction network based on the cross-domain attention mechanism network to generate data features related to different fault types and sensors. S4: Spatial information location encoding is performed on the data features obtained after processing the standard dataset to form a spatial relative location encoding system for different fault types and sensors; S5: A label self-correction strategy based on spatial differentiation is proposed. By utilizing the spatial relative position coding system of different fault types and sensors, spatial information feature matching is performed on different fault types and sensors, and spatial feature projection is performed on the matched fault type and sensor information. The final fault classification results are obtained.
2. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, The cross-domain attention mechanism network consists of a channel attention network and a spatial attention network. The channel attention network has three parallel branches, which are concatenated and then set with convolutional layers, activation functions, and convolutional layers. The first branch includes a convolutional layer, an asymmetric convolutional block, batch normalization, and an activation function. The second branch includes a convolutional layer, two asymmetric convolutional blocks, batch normalization, and an activation function. The third branch includes a convolutional layer and an average pooling layer. The spatial attention network includes an average pooling layer and a max pooling layer, which are concatenated and then set with multiple asymmetric convolutional blocks and a sigmoid activation function. The cross-domain attention mechanism network processing procedure is as follows: ; in, This represents the output of the cross-domain attention mechanism network. This represents the learnable weight parameters. Represents sensor data, This represents element-wise multiplication. Represents the ReLU activation function. This represents an asymmetric convolution block with kernel size k. This represents a 1×1 convolution operation. This indicates the average pooling operation. This indicates a max pooling operation. This represents the Sigmoid activation function. This represents the concatenation operation of features. This indicates a summation operation across all feature channels. This represents a non-linear activation function.
3. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, The output of the cross-domain attention mechanism network is fed into a hierarchical fusion multidimensional fault feature extraction network, specifically including: The output of the cross-domain attention mechanism network is divided into four sensor data streams. Each sensor data stream enters the hierarchical fusion multidimensional fault feature extraction network, which includes multi-branch convolution, convolutional pooling, and multi-branch pooling. The multi-branch convolution process is as follows: ; in, This represents the output of a multi-branch convolutional block. BN and BP represent the data normalization and batch normalization operations of the multi-branch convolutional layer, respectively. This indicates a max pooling operation. f = 1, 3, represents the standard convolution of f × f. Represents convolution sum, This represents an asymmetric convolution operation with a kernel size of 3 and g repetitions. The output after convolutional pooling is calculated as follows: ; in, It is the output after convolutional pooling. This represents an asymmetric convolution operation with a kernel size of 1, repeated 3 times. The multi-branch pooling process is as follows: ; in, It is a multi-branch pooled block output. Indicates that the average pooling operation is repeated. Second-rate.
4. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, The enhanced orthogonal spatial heterogeneous network in S3, which embeds a multidimensional generator, consists of two parts: a hybrid generator and a regularized discriminator. The hybrid generator uses a recurrent neural network model to capture the time-series features of the data, combines a variational autoencoder to map the data to the latent space and generate new data, and then optimizes the feature representation and network training through the SE module and identity mapping residual connection. Initialization is performed using the hidden states of the recurrent neural network model: ; in, This represents the initial hidden state of the recurrent neural network (RNN) model. These are the input parameters for the recurrent neural network model. This indicates the initialization of the weight matrix; The recurrent neural network model takes incomplete datasets, missing datasets, and the output of the data feature encoder as input, performs recursive processing, and utilizes the hidden state from the previous time step. and the latent variables at the current moment Calculate the hidden state at the current time step. : ; in, This represents the weight matrix of the hidden layer, responsible for performing weighted transformations on the hidden state from the previous time step. This represents the weight matrix of the input layer, used to weight the latent variables at the current time step. Weighting; It is the ReLU activation function. These are the parameters of the recurrent neural network model. For bias terms; Through recursive calculation, the recurrent neural network model progressively captures the time-series characteristics of the data, and integrates the time-series information through a summation operation to obtain... As the output generated by the recurrent neural network model: ; in, For the integrated time series information, N represents the total number of time steps processed by the recurrent neural network; The data processed by the recurrent neural network model is mapped to the latent space Z through a variational autoencoder, and new data is generated from the latent space using a decoder. The operation process of the variational autoencoder is as follows: ; in, This represents the set of variables and parameters of a variational autoencoder. These are variables in the latent space, representing the feature representation of the input data in the latent space after encoding. It is the parameter set in a variational autoencoder. This represents the weights of the encoder portion of the variational autoencoder. This indicates the bias of the encoder portion in a variational autoencoder; It is sampling noise, generated by a standard normal distribution. Represents a non-linear activation function. Indicates the weights of the decoder. The weights and biases of the decoder are used to decode the variables m in the latent space to generate new data. The generated feature maps are fed into the SE module. The SE module uses a fully connected layer to adjust the importance of each channel using weighted coefficients, making the model more focused on key features. The formula is as follows: ; in, This indicates the input feature map The process involves calculating the feature map after channel weight adjustment. This represents the generator's output at the current moment. This represents the parameter set of the SE module. This represents the weight matrix of the fully connected layers in the SE module. It is a bias term. This indicates a global average pooling operation on the input features, where S is the Sigmoid activation function. The features processed by the SE module are further optimized through identity mapping residual connections, as shown in the expression: ; in, Represents the residual of the identity mapping. It is the output of the generator from the previous time step, which is filtered by noise. and input features Generate data, These are the parameters of the corresponding network layer; In the regularized discriminator, the feature encoder encodes incomplete and missing datasets to obtain feature representations. The final feature encoding representation is obtained by weighted summation: ; in: This represents the feature representations generated by different feature encoders. These are the weight coefficients of each feature encoder, representing the importance of different features in the discrimination. It is the total number of feature encoders. This represents the final feature encoding representation. For the set of all relevant parameters; The regularized discriminator performs gradient calculations on incomplete and missing datasets. The gradient calculation formula is as follows: ; in, Indicates input data Find the gradient. It is the feature representation function in gradient calculation-related operations. These are the corresponding parameters; The true / false probabilities are obtained by activating the input features using the sigmoid activation function. ; in, It is the output of the regularization discriminator, representing the data. Let S be the probability of the true data, and S be the sigmoid activation function, ensuring that the output value is between 0 and 1. It is the regularization coefficient, used to control the normalization of the gradient. It is a hyperparameter used to control the size of the regularization term.
5. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, S4 include: Computation of spatial location coding function Its expression is: ; ; in, Represents a spatial location encoding function. It is a characteristic representation of the Fourier transform. The weight matrix represents the features. It is a sensor characteristic. It is the standard deviation of the Gaussian kernel. This refers to the number of position matrices considered. For fault type index, For dimensional indexing, This represents the total number of vector dimensions; Calculate the discriminative features between different sensors and between different fault types. Calculated using the following formula: ; in, Indicates position-encoded input. This represents a multilayer perceptron. Represents a linear unit with Gaussian error. Indicates the sensor state at the previous moment; calculate : ; in, This represents the feature alignment vector between sensor features and the fault prototype. Represents the spatial characteristics of the sensor. Indicates the characteristics of the fault prototype; Then through Calculate different fault types : 。 6. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 5, characterized in that, By utilizing a spatial relative position coding system for different fault types and sensors, spatial information feature matching is performed on different fault types and sensors, specifically including: Based on the discriminative features between different sensors and different fault types Calculate the similarity matrix: ; in, Represents the similarity matrix. Represents the dot product of vectors. Describes the Euclidean norm of a vector. , ; Convert the similarity matrix into a soft-assignment matrix. : ; ; ; in, and These represent the similarity matrices after row and column softmax processing, respectively. The correlation index between different fault types and different sensors is represented by the soft allocation matrix, which measures and weights the relationship between different sensors and fault types by using time-frequency domain feature similarity. based on Computational spatial information feature matching: ; in, This represents the set of high-confidence pairings. This represents the threshold for each matching pair. ,if This indicates a high-confidence match; The mutual information maximization strategy is adopted to further improve the quality and consistency of feature matching: ; in, This represents the set of pairs after further filtering, ensuring the quality and consistency of the pairs by maximizing the optimization of information.
7. The fault diagnosis method based on orthogonal spatial heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, Spatial feature projection is performed on the matched fault types and sensor information to obtain the final fault classification result. The specific steps are as follows: For the time-frequency domain characteristics F of sensor o and sensor p o and F p Perform a preliminary matching operation to obtain matching pairs. Based on this, the offset is calculated. And perform spatial feature projection: ; in, This means projecting the information from the 0th sensor onto the pth sensor; Calculate the correction vector based on the results of spatial feature projection. : ; in, and These are weighting coefficients used to balance offset and feature differences; Based on the calculated correction vector Update matching pairs: .
8. The fault diagnosis method based on orthogonal space heterogeneous mapping under imperfect information conditions according to claim 1, characterized in that, Introducing reconstruction loss during data reconstruction process Introducing matching loss during the self-correction process of sticky notes and differentiation loss The specific process is as follows: In the data reconstruction stage, the overall reconstruction loss function is defined as follows: ; Where G represents the hybrid generator, which is responsible for generating simulated data; D is the regularization discriminator, which is used to distinguish between real and fake data; x represents real data, and z is the noise vector input to the generator; and These are respectively the expectation of the real data and the noise vector; and It is the log probability of the regularized discriminant's judgment of whether the data is true or false; F is the function that transforms the data into the frequency domain; This is the generator output data. Its frequency domain characteristics, It is a target frequency characteristic; Measuring the difference between the two is a hyperparameter that weighs the proportion of loss due to differences in frequency characteristics; Calculate the ground truth matching exponent and combine it with the soft allocation matrix. Calculate spatial information feature matching loss : ; in, This represents the set of indices for ground truth matching, i.e., the set of indices for correct matches. This represents the balancing factor, used to balance the weights of positive and negative samples. This represents a adjustment factor used to adjust the model's focus on difficult-to-classify samples. This indicates the total number of matching pairs; Based on the differential offset and the corrected geospatial offset, the differential loss is... Defined as: ; in, It is a differential offset. It is the corrected geodetic offset, and 𝐾 is the number of differential projections.
Citation Information
Patent Citations
Fingerprint feature recognition analysis method based on direction field guidance and space attention technology
CN119296143A
Process monitoring method for abnormal operation state in safe water treatment process
CN119513564A
Cited By
Full-life-cycle health diagnosis method of variable frequency water supply unit for water resource management
CN121412597A
Abnormal number AI risk diagnosis method and system based on multi-dimensional feature fusion
CN122226515A