A method for enhancing the diagnostic generalization of vertical takeoff and landing aircraft based on causal weights
By using causal weighting methods and deep learning technology in the fault diagnosis of vertical take-off and landing aircraft, a directed acyclic graph is established and causal strengthened, the problem of insufficient fault diagnosis capabilities in the existing technology is solved, and the generalization ability of the fault prediction model is significantly improved.
Patent Information
- Application Number
- CN202510352258.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing intelligent fault diagnosis methods have insufficient diagnostic capabilities in dealing with complex multi-factor systems, resulting in poor generalization capabilities of fault prediction models.
Using a method based on causal weight, a directed acyclic graph between flight characteristics and fault types is established through the information flow causal discovery method, mutual information between nodes is calculated and converted into causal intensity for causal enhancement, and a convolutional neural network and graph convolutional network are combined for feature extraction and fusion to build a fault prediction model.
It effectively improves the generalization ability of the fault prediction model, so that it can better adapt to different operating conditions and data changes, and accurately diagnose aircraft failures.
Smart Images

Figure CN119862504B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of aircraft fault diagnosis, and specifically to a method for enhancing the diagnostic generalization of a vertical takeoff and landing aircraft based on causal weights. Background Art
[0002] With the rapid development of artificial intelligence technology, vertical takeoff and landing aircraft have been widely used in multiple fields such as drones, military, and commercial aviation. These aircraft play a key role in performing tasks, ensuring the efficiency and safety of flight missions. However, with the increase in usage time, vertical takeoff and landing aircraft often face problems such as overload, fatigue, wear, and sensor failures under complex working conditions, especially when flying under different wind speeds and directions. These problems may cause the aircraft to malfunction or lead to serious safety accidents.
[0003] Currently, intelligent diagnostic methods based on machine learning and deep learning have emerged in the field of aircraft fault diagnosis. These methods are of great significance in improving the accuracy of fault diagnosis and reducing the dependence on manual feature extraction. However, most of the existing intelligent fault diagnosis frameworks focus on classifying and identifying the correlation between data, but ignore the causal relationship between data and labels. This makes the diagnostic ability of these methods insufficient when dealing with complex multi-factor systems, resulting in poor generalization ability of the obtained fault prediction models. Therefore, it is urgent to solve this problem. Summary of the Invention
[0004] In order to avoid and overcome the technical problems existing in the prior art, the present invention provides a method for enhancing the diagnostic generalization of a vertical takeoff and landing aircraft based on causal weights. The present invention can effectively improve the generalization ability of the fault prediction model.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for enhancing the diagnostic generalization of a vertical takeoff and landing aircraft based on causal weights, comprising the following enhancement steps:
[0007] S1. Obtain the fault flight data of the vertical takeoff and landing aircraft, and establish a directed acyclic graph between flight features and fault types through the information flow causal discovery method;
[0008] S2. Calculate the mutual information between each node in the directed acyclic graph, and convert the mutual information into the causal strength of each node to strengthen the causality of each node;
[0009] S3. Input the nodes after causal strengthening into a convolutional neural network for convolutional operation, and use causal weights for dynamic weighting during the convolutional operation to extract local features;
[0010] S4. Concatenate the local features and the causal features in the directed acyclic graph to obtain concatenated features; process the concatenated features using a multi-layer perceptron to obtain the adjacency matrix of the nodes.
[0011] S5. Use a graph convolutional network to fuse the concatenated features and the adjacency matrix to obtain high-order features.
[0012] S6. Input the high-order features into a fully connected network to complete the adversarial training with a convolutional neural network and a graph convolutional network as feature extractors and a fully connected network as a fault classifier; the feature extractor and the fault classifier constitute a fault prediction model.
[0013] As a further solution of the present invention: The specific steps of step S2 are as follows:
[0014] S21. Based on the time series data of each node in the directed acyclic graph, calculate the information entropy of each node; the calculation formula of the information entropy is as follows:
[0015]
[0016] In the formula, H(X i ) represents the information entropy of the i-th node X i in the directed acyclic graph; represents the value of node X i at time t; represents the probability distribution; log represents the logarithmic function with base 10.
[0017] S22. Use the information entropy to calculate the mutual information between each node, and the calculation formula of the mutual information is as follows:
[0018] I(X i , X j ) = H(X i ) + H(X j ) - H(X i , X j );
[0019] In the formula, I(X i , X j ) represents the mutual information between the i-th node X i and the j-th node X j in the directed acyclic graph; H(X j ) represents the information entropy of the j-th node X j in the directed acyclic graph; H(X i , X j ) represents the joint information entropy between node X i and node X j ;
[0020] S23. Arrange the calculated mutual information in sequence to form a square causal relationship matrix M, where each element represents the mutual information between the corresponding two nodes;
[0021] S24. Perform a normalization operation on the causal relationship matrix to obtain a causal strength matrix M″;
[0022] S25. Strengthen the causality of each node by combining the causal strengths in the causal strength matrix M″;
[0023] S26. Calculate the values of each node at each moment after causal strengthening according to the calculation formula of causal strengthening.
[0024] As a further solution of the present invention: The calculation formula of causal strengthening is expressed as follows:
[0025]
[0026] In the formula, represents the value of node X after causal strengthening i at time t; represents node X j at time t.
[0027] As a further solution of the present invention: The calculation formula of each element in the causal strength matrix M″ is as follows:
[0028]
[0029] In the formula, M″ ij represents the causal strength between the i-th node X i and the j-th node X j in the directed acyclic graph; w j represents the weight factor of node X j ; σ j represents the standard deviation of the time series data of node X j ; μ j represents the mean of the time series data of node X j ; M ij represents the mutual information between the j-th node X i and the j-th node X j in the directed acyclic graph, that is, I(X i , X j ); α represents a scaling factor; M min,j represents the minimum value of all mutual information related to node X j ; M max,j represents the maximum value of all mutual information related to node X j .
[0030] As a further solution of the present invention: The specific steps of step S3 are as follows:
[0031] S31. Organize the values of each node at each moment after causal strengthening into a data set corresponding to that moment;
[0032] S32. Input the data sets of each moment into a convolutional neural network in sequence for convolutional operations to extract the local features between each node through the convolutional neural network.
[0033] As a further solution of the present invention: The convolutional operation is expressed as follows:
[0034]
[0035] In the formula, represents the local feature of node X i at time t; represents the value of the k-th node X k at time t; represents and the similarity between; represents and the similarity between; represents and the calculated value of the softmax function between; n represents the total number of nodes; represents the value of node X j after causal strengthening at time t; w represents the convolutional kernel of the convolutional neural network; b represents the bias term of the convolutional neural network; F(·) represents the convolutional operation; exp(·) represents the exponential function with the natural constant e as the base.
[0036] As a further solution of the present invention: The specific steps of step S4 are as follows:
[0037] S41. Concatenate the local features and the causal features in the directed acyclic graph to obtain concatenated features; The concatenated features are expressed as follows:
[0038]
[0039] In the formula, represents the concatenated features formed after concatenation; represents the causal feature of node X i in the directed acyclic graph at time t; [·,·] represents the concatenation operation;
[0040] S42. Then perform a non-linear transformation on the concatenated features to obtain non-linear features; The non-linear features are expressed as follows:
[0041]
[0042] In the formula, represents the non-linear feature of merge represents the non-linear bias term; W merge represents the non-linear weight; σ(·) represents the activation function;
[0043] S43. Use a multi-layer perceptron to convert the concatenated features into high-dimensional features, and the high-dimensional features are represented as follows:
[0044]
[0045] In the formula, Z represents the feature matrix formed by combining various non-linear features; MLP represents the multi-layer perceptron; represents the high-dimensional feature corresponding to Z;
[0046] S44. Multiply the feature matrix by its transpose matrix to obtain the adjacency matrix; the adjacency matrix is represented as follows:
[0047]
[0048] In the formula, A represents the adjacency matrix; normal represents the matrix multiplication operation; represents the transpose matrix of
[0049] S45. Select the first r nearest neighbors of each node to form the sparse adjacency matrix A′;
[0050] A′ = Top-r(A);
[0051] In the formula, Top-r represents the operation of taking the first r nearest neighbors.
[0052] As a further solution of the present invention: The specific content of step S5 is as follows:
[0053] S51. Input the feature matrix and the sparse adjacency matrix into the graph convolutional network;
[0054] S52. Use two-layer receptive field convolution for convolution operation to obtain the corresponding high-order features;
[0055]
[0056] In the formula, H 0 represents the feature learned by the first-layer receptive field convolution; W 0 represents the weight matrix of the first-layer receptive field convolution; MRFConv represents the receptive field convolution; W 1 represents the weight matrix of the second-layer receptive field convolution; H 1Represents the features learned by the second-layer receptive field convolution, that is, high-order features.
[0057] As a further solution of the present invention: The objective function in the adversarial training process is as follows:
[0058] L Total = L C + λL DA ;
[0059] In the formula, L Total represents the objective function; L C represents the classification loss; L DA represents the domain alignment loss; λ represents a hyperparameter.
[0060] As a further solution of the present invention: The calculation formulas of each loss are specifically expressed as follows:
[0061]
[0062] In the formula, x s represents the source domain sample; x t represents the target domain sample; y represents the fault type, C(x s , y) represents the fault classifier model. Inputting the source domain sample, it outputs the probability that the source domain sample belongs to the fault type y. C(x t , y) represents the fault classifier model. Inputting the target domain sample, it outputs the probability that the target domain sample belongs to the fault type y; represents the output of the source domain classifier; represents the output of the target domain classifier;
[0063]
[0064] In the formula, D(·) represents the domain discriminator; represents the probability that the domain discriminator determines the generated source domain sample as false; represents the loss of the target domain sample.
[0065] Compared with the prior art, the beneficial effects of the present invention are:
[0066] 1. First, a directed acyclic graph between flight characteristics and fault types is established through the information flow causality discovery method, which can clearly present the causal relationship between the two and provide a basis for subsequent analysis. Then, the mutual information between nodes is calculated and transformed into causal strength for causal enhancement, strengthening the expression of causal relationships in the data. Causal weights are dynamically weighted in the convolutional neural network to extract local features, combined with the causal features of the directed acyclic graph, making full use of different types of feature information. High-order features are obtained by fusing features through the graph convolutional network, and then adversarial training is completed through the fully connected network. The constructed fault prediction model integrates various technical advantages, effectively improving the model generalization ability, enabling it to better adapt to different working conditions and data changes, and accurately diagnosing aircraft faults.
[0067] 2. Starting from calculating the node information entropy, the mutual information between nodes is calculated based on the information entropy. This information-theory-based method can accurately measure the degree of association between nodes. The mutual information is arranged into a causality matrix and normalized to obtain a causal strength matrix, providing a quantitative basis for causal enhancement. According to specific steps, the nodes are causally enhanced in combination with the causal strength matrix, enabling the nodes to more prominently reflect their causal connections with other nodes, further mining the causal information in the data, thus providing more valuable data features for subsequent model training and helping to improve the accuracy and generalization of the fault diagnosis model.
[0068] 3. The causal enhancement calculation formula clarifies how to adjust the values of nodes at each moment according to the causal strength matrix. Through this formula, the value of each node at a certain moment can comprehensively consider the values of other nodes at the same moment and their causal strengths, realizing the dynamic adjustment and fusion of node information. This makes the information transfer of nodes in the time series more reasonable, enhancing the coherence of causal relationships in the data, helping the model to more accurately capture the time-varying associations between aircraft faults and flight characteristics, and thus improving the reliability and generalization of fault prediction.
[0069] 4. The calculation formula for the elements of the causal strength matrix comprehensively considers multiple factors. It includes the node weight factor, the mean and standard deviation of time series data, the maximum and minimum values of mutual information, etc. Through this complex calculation method, the strength of the causal relationship between nodes can be more meticulously characterized. The weight factors of different nodes can be set according to their importance, and the statistical characteristics of time series data and the maximum and minimum values of mutual information further calibrate the causal strength, enabling the causal strength matrix to more accurately reflect the true causal relationship between nodes, providing more accurate parameters for subsequent causal enhancement and model training, and improving the performance of the fault diagnosis model.
[0070] 5. Organize the node values after causal enhancement into a dataset, providing an ordered and causally informative data input for the convolutional neural network. Sequentially inputting the dataset for convolutional operations enables the convolutional neural network to effectively extract the local features between each node. This data processing and feature extraction method not only retains the data advantages after causal enhancement but also utilizes the advantages of the convolutional neural network in local feature extraction, laying a good foundation for subsequent model fusion and fault diagnosis, and helping to improve the model's ability to identify aircraft fault features.
[0071] 6. The convolutional operation formula details how to perform dynamic weighting using causal weights in the convolutional neural network. The formula includes elements such as node similarity calculation, softmax function, and causal intensity matrix. Through this complex calculation method, the convolutional operation can more specifically weight the information of different nodes. When extracting local features, it highlights the node information closely related to the current node's causality and suppresses irrelevant or weakly related information, making the extracted local features better reflect the causal relationship between aircraft faults and flight characteristics, improving the quality of feature extraction, and thereby enhancing the model's fault diagnosis ability and generalization.
[0072] 7. By concatenating the local features and causal features to obtain concatenated features, different levels and types of feature information are fused, enriching the data representation. Performing operations such as non-linear transformation and multi-layer perceptron processing on the concatenated features gradually transforms the low-dimensional features into high-dimensional features, enabling better capture of complex patterns in the data. Finally, through matrix operations, the adjacency matrix and sparse adjacency matrix are obtained, providing a suitable input format for the graph convolutional network, which helps the graph convolutional network to more effectively mine the graph structure information in the data and improve the accuracy and generalization of the model for aircraft fault diagnosis.
[0073] 8. Input the feature matrix and sparse adjacency matrix into the graph convolutional network, and perform convolutional operations using two-layer receptive field convolution to obtain high-order features. This operation method can fully utilize the advantages of the graph convolutional network in processing graph-structured data, learning different levels of graph structure features through convolutions of different layers. The features learned by the first-layer receptive field convolution provide the basis for the second layer, and the second layer further mines more complex features. The obtained high-order features can more comprehensively reflect the complex relationship between aircraft flight characteristics and fault types, helping to improve the performance and generalization ability of the fault diagnosis model.
[0074] 9. The adversarial training objective function comprehensively considers the classification loss and the domain alignment loss. The classification loss is used to measure the accuracy of the model in classifying different samples, ensuring that the model can correctly identify the fault types of the aircraft. The domain alignment loss is committed to reducing the differences between the source domain and the target domain, enabling the model to maintain good performance under different working conditions or data distributions. By adjusting the hyperparameter λ to balance the two losses, the model can improve the classification accuracy while enhancing the generalization ability and better adapting to the aircraft fault diagnosis tasks in different environments.
[0075] 10. The classification loss formula can effectively guide the model to learn general fault features and improve the classification ability for different samples by comparing the classification situations of source domain and target domain samples. The domain alignment loss formula uses a domain discriminator to measure the probability of misjudging source domain samples and the loss of target domain samples respectively, prompting the model to align the data distributions of the source domain and the target domain during training, reducing the impact of inter-domain differences on the model performance, further enhancing the generalization of the model, and enabling the model to diagnose aircraft faults more stably and accurately when facing different working conditions and data changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 It is a flowchart of the enhancement steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0078] Please refer to Figure 1 , in the embodiments of the present invention, a method for enhancing the generalization of vertical takeoff and landing aircraft diagnosis based on causal weights includes the following content:
[0079] I. Construct a directed acyclic graph:
[0080] Obtain the fault flight data of the vertical takeoff and landing aircraft, and establish a directed acyclic graph between the flight features and the fault types through the information flow causal discovery method.
[0081] 1. Collect the fault flight data under different working conditions from the sensor system of the vertical takeoff and landing aircraft.
[0082] During the process of collecting fault flight data, the vertical takeoff and landing vehicle is operated under different working conditions. By setting different flight modes (such as vertical takeoff and landing, hovering, horizontal flight, etc.), and simulating different wind speed and wind direction conditions, the sensor data of the vehicle under these working conditions is collected. The time interval of data collection is usually uniform to ensure that each data point comes from the same time period, facilitating subsequent processing. During this process, the collected fault flight data includes:
[0083] Three-axis vibration (vibration in the X, Y, and Z axes): These vibration data reflect the dynamic response of the vehicle in different directions, recording the vibration amplitude and frequency in each direction.
[0084] Torque: Records the torque generated by the vehicle's engine or propulsion system, reflecting the load and power changes of the power system.
[0085] Rotation speed: Records the rotation speeds of key components such as rotors and motors, reflecting the operating state of the power system.
[0086] Each sensor continuously monitors the operating state of the vehicle throughout the experiment, generating regular numerical values. For example, the records of the three-axis vibration sensor are as follows:
[0087] A X =[x 1 ,x 2 ,x 3, ,...,x n ;
[0088] A Y =[y 1 ,y 2 ,y 3 ,...,y n ;
[0089] A Z =[z 1 ,z 2 ,z 3 ,...,z n ;
[0090] Among them, A X represents the vector composed of vibration data in the X-axis direction, x 1 represents the vibration data in the X-axis direction at the first moment, x 2 represents the vibration data in the X-axis direction at the second moment, x 3 represents the vibration data in the X-axis direction at the third moment, x n represents the vibration data in the X-axis direction at the nth moment. A Y represents the vector composed of vibration data in the Y-axis direction, y 1Represents the vibration data in the Y-axis direction at the first moment, y 2 Represents the vibration data in the Y-axis direction at the second moment, y 3 Represents the vibration data in the Y-axis direction at the third moment, y n Represents the vibration data in the Y-axis direction at the nth moment. A Z Represents the vector composed of vibration data in the Z-axis direction, z 1 Represents the vibration data in the Z-axis direction at the first moment, z 2 Represents the vibration data in the Z-axis direction at the second moment, z 3 Represents the vibration data in the Z-axis direction at the third moment, z n Represents the vibration data in the Z-axis direction at the nth moment.
[0091] Torque data: T = [T 1 , T 2 , T 3 ... T n ; Rotational speed data: V = [v 1 , v 2 , v 3 ... v n . Wherein, T represents the vector composed of vibration data, T 1 Represents the torque data at the first moment, T 2 Represents the torque data at the second moment, T 3 Represents the torque data at the third moment, T n Represents the torque data at the nth moment. V represents the vector composed of rotational speed data, v 1 Represents the rotational speed data at the first moment, v 2 Represents the rotational speed data at the second moment, v 3 Represents the rotational speed data at the third moment, v n Represents the rotational speed data at the nth moment.
[0092] To ensure data diversity, during the experiment, not only the operation data of normal aircraft are collected, but also various common aircraft fault scenarios are simulated. For example:
[0093] Fault state 1: The battery overheats, resulting in unstable power output.
[0094] Fault state 2: Engine failure, manifested as abnormal fluctuations in torque and rotational speed.
[0095] Fault state 3: Sensor failure, resulting in distorted or completely lost sensor data.
[0096] Through the above data acquisition process, sensor data of the vertical takeoff and landing aircraft under different working conditions and different fault modes can be obtained. Using these data, subsequent causal analysis, feature weight assignment, and fault diagnosis model training can be carried out to achieve efficient diagnosis and early warning of aircraft faults.
[0097] 2. Use the information flow causal discovery method to discover the causality of time series data.
[0098] 1) Data preprocessing: Filter the noise (Butterworth low-pass filter, median filter), interpolate the missing values (linear interpolation), and standardize (Z-score) the sensor time series data (such as vibration, torque, speed) to ensure data quality.
[0099] 2) Causal structure discovery:
[0100] Fast causal inference: Construct a partial ancestral graph (PAG) through conditional independence tests (such as kernel conditional independence test, KCIT) to identify the adjacency relationships (→, —, ) between variables, forming a Markov equivalence class.
[0101] Information flow algorithm: Based on the asymmetry of information transfer, calculate the information flow strength between variables, orient the causal edges, and generate a complete directed acyclic graph (DAG).
[0102] Output and application: Finally, a DAG reflecting the causal relationship of faults (such as "battery temperature → torque anomaly → vibration intensification") is obtained, providing a structural prior for subsequent feature weighting and graph convolutional networks, and guiding the model to focus on key causal features.
[0103] II. Causal strengthening:
[0104] Calculate the mutual information between each pair of nodes in the directed acyclic graph, and convert the mutual information into the causal strength of each node to strengthen the causality of each node.
[0105] In fault network diagnosis, it is a challenge to ensure that the model assigns weights according to the importance of faults in a multi-factor system. To optimize model learning, calculate the information dependence relationship, fuse the causal relationship, and apply a weighting matrix to process the input data.
[0106] Calculate the information entropy of each node, which is used to measure the randomness or uncertainty in the time series data of the node.
[0107] Assume that the value of node X i at time t is a discrete value and has a corresponding probability distribution Then the information entropy H(X i ) of node X i is defined as:
[0108]
[0109] Wherein, H(X i ) represents the information entropy of the i-th node X i in the directed acyclic graph; represents the value of node X i at time t; represents 's probability distribution; log represents the logarithmic function with base 10.
[0110] The mutual information between node X i and node X j quantifies the degree of shared information between them. The calculation formula of mutual information is as follows:
[0111] I(X i , X j ) = H(X i ) + H(X j ) - H(X i , X j );
[0112] Wherein, i(X i , X j ) represents the mutual information between the i-th node X i and the j-th node X j in the directed acyclic graph; H(X j ) represents the information entropy of the j-th node X j in the directed acyclic graph; H(X i , X j ) represents the joint information entropy between node X i and node X j .
[0113] By calculating the mutual information between all node pairs, the mutual information can be arranged in sequence and combined into a square causal relationship matrix M, and each element in it represents the mutual information between the corresponding two nodes.
[0114] After receiving the causal relationship matrix, first normalize it to scale the eigenvalues to the range of [0, 1]. This normalization ensures that each feature has a clear relative importance in subsequent analysis or modeling tasks and eliminates the scale difference. Apply standard normalization techniques, such as weighted min-max normalization and Z-score standardization, to achieve this. The causal intensity matrix M” obtained after the normalization operation is represented as follows:
[0115]
[0116] Wherein, M″ ijDenote the $i$-th node $X$ in the directed acyclic graph i and the $j$-th node $X$ j The causal strength between them; $w$ j Denote the node $X$ j The weight factor; $\sigma$ j Denote the node $X$ j The standard deviation of the time series data of the node $X$; $\mu$ j Denote the node $X$ j The mean of the time series data of the node $X$; $M$ ij Denote the $i$-th node $X$ in the directed acyclic graph i and the $j$-th node $X$ j The mutual information between them, i.e., $I(X$ i , $X$ j ); $\alpha$ represents a scaling factor used to control the impact of weighted min-max normalization on the final result; $M$ min,j Denote the minimum value of all mutual information related to the node $X$ j ; $M$ max,j Denote the maximum value of all mutual information related to the node $X$ j .
[0117] This process generates a normalized causal strength matrix $M''$, which is then used to adjust the input data of the convolutional neural network.
[0118] In the first convolutional layer of the convolutional neural network, the weighted input $X'$ of the node $X$ i is obtained by summing the weighted contributions of all other nodes, that is, strengthening the causality of each node by combining the respective causal strengths in the causal strength matrix $M''$. The calculation formula for causal strengthening is as follows: i
[0119]
[0120] In the formula, Denote the value of the node $X$ i after causal strengthening at time $t$; Denote the value of the node $X$ j at time $t$.
[0121] This input will be adjusted according to the causal relationship between nodes to generate a weighted input sequence for use by the convolutional neural network.
[0122] III. Extract local features:
[0123] To consider the influence of failure causes, an adaptive attention adjustment mechanism is adopted to enhance the flexibility of the model. This mechanism dynamically allocates weights, strengthens the contributions of nodes closely related to the output, and suppresses features unrelated to the failure. The values of each node at each moment after causal enhancement are organized into a dataset corresponding to that moment, and then the datasets of each moment are successively input into a convolutional neural network for convolutional operations with adaptive adjustment to extract local features between nodes through the convolutional neural network. The convolutional operation with this adaptive adjustment is given as follows:
[0124]
[0125] In the formula, represents the local feature of node X i at time t; represents the value of the k-th node X k at time t; represents and the similarity between; represents and the similarity between; represents and the calculated value of the softmax function between; n represents the total number of nodes; represents the value of node X j at time t after causal enhancement; w represents the convolutional kernel of the convolutional neural network; b represents the bias term of the convolutional neural network; F(·) represents the convolutional operation; exp(·) represents the exponential function with the natural constant e as the base.
[0126] This dynamic adjustment improves the model's ability to understand time series data, especially in capturing potential causal relationships. Normalization, weighting, and the attention mechanism enable the model to adjust its response according to the changing importance of failure causes, thereby improving training efficiency and prediction accuracy.
[0127] IV. Obtaining the adjacency matrix:
[0128] After processing time series sensor data through the convolutional neural network, the fully connected layer generates new feature representations, i.e., local features. Then, these local features are input into the graph convolutional network, where two different types of nodes need to be specially processed to mitigate the influence of different feature node types. To further enhance the training process by leveraging a directed acyclic graph, a node sharing mechanism is proposed. Initially, the causal features in the directed acyclic graph and the local features based on the convolutional neural network are concatenated as features in the following way:
[0129]
[0130] In the formula, represents the splicing feature formed after splicing; represents the causal feature of node X in the directed acyclic graph at time t i ; [·,·] represents the splicing operation.
[0131] This operation combines two feature vectors to generate a more informative and compact representation for each node. Then, the concatenated feature vector is processed through a non-linear transformation, which can be represented by the following equation:
[0132]
[0133] In the formula, represents the non-linear feature of merge ; b merge represents the non-linear bias term; W 1 2 represents the non-linear weight, usually represented in the form of a matrix, that is, it can be represented as a mapping matrix that transforms the concatenated feature vector into a new feature space; σ(·) represents the activation function, which introduces non-linearity in the transformation process, enabling the network to model complex relationships between features.
[0134]
[0135]
[0136] In the graph convolutional network, in order to construct the adjacency matrix A, the feature matrix formed by combining each non-linear feature is first processed through a multi-layer perceptron to obtain high-dimensional features, and then the adjacency matrix is obtained by multiplying the high-dimensional features by their transpose. Then, the first r nearest neighbors of each node are selected to form the sparse adjacency matrix A'. Specifically, it is represented as follows: In the formula, Z represents the feature matrix formed by combining each non-linear feature; MLP represents the multi-layer perceptron;
[0137] V. Obtaining high-order features:
[0138] Apply receptive field convolutions with different receptive fields (k 1 , k 2 , k 3 ) to the instance graph and learn through the receptive field convolutional layer to obtain the corresponding high-order feature representation:
[0139]
[0140] In the formula, H0 represents the features learned by the first-layer receptive field convolution; W 0 represents the weight matrix of the first-layer receptive field convolution; MRFConv represents the receptive field convolution; W 1 represents the weight matrix of the second-layer receptive field convolution; H 1 represents the features learned by the second-layer receptive field convolution, i.e., high-order features.
[0141] VI. Constructing a fault prediction model:
[0142] The goal of adversarial training is to make the model perform well in both the source domain (training data working conditions) and the target domain (new working conditions). Its core consists of the following two components:
[0143] 1. Feature extractor:
[0144] It is composed of a convolutional neural network and a graph convolutional network and is responsible for generating high-order feature representations.
[0145] The role of the graph convolutional network: Generate high-quality features containing causal relationships through graph structure modeling and feature fusion.
[0146] 2. Domain classifier:
[0147] An independent fully connected network used to determine whether the features come from the source domain or the target domain.
[0148] Adversarial goal: The feature extractor is optimized to "fool" the domain classifier so that it cannot distinguish between source domain and target domain features.
[0149] 3. Objective function:
[0150] During the adversarial training process, the objective function is as follows:
[0151] L Total = L C + λL DA ;
[0152] In the formula, L Total represents the objective function; L C represents the classification loss; L DA represents the domain alignment loss; λ represents a hyperparameter.
[0153]
[0154] In the formula, x s represents the source domain samples; x t represents the target domain samples; y represents the fault type, C(x s , y) represents the fault classifier model. Inputting the source domain samples, it outputs the probability that the source domain belongs to the fault type y, C(x t, y) represents the fault classifier model, which takes the target domain samples as input and outputs the probability that the target domain sample belongs to the fault type y; represents the output of the source domain classifier; represents the output of the target domain classifier.
[0155]
[0156] In the formula, D(·) represents the domain discriminator; represents the probability that the domain discriminator determines the generated source domain samples as false; represents the loss of the target domain samples.
[0157] By combining these two loss functions, the ultimate goal is to achieve accurate classification on the source domain while learning features that are equally effective on the target domain.
[0158] If the validation error of the target domain does not continuously decrease, terminate the training in advance. After the training is completed, the model used for fault prediction is the optimized "feature extractor + label classifier".
[0159] The prediction process using the fault prediction model is as follows:
[0160] The specific steps are as follows:
[0161] 1) Input new data:
[0162] Input the sensor data (such as vibration, rotational speed, torque, etc.) of the target domain (new working condition).
[0163] 2) Feature extraction:
[0164] The feature extractor (a hybrid structure composed of a convolutional neural network and a graph convolutional network) converts the original data into a high-dimensional feature representation. These features have been optimized through adversarial training and can ignore the working condition differences and focus on the fault-related patterns.
[0165] 3) Fault classification:
[0166] The label classifier (a fully connected network or a Softmax layer) outputs the probability of the fault type based on the extracted features, and thus the fault diagnosis is completed.
[0167] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weights, characterized in that: The enhancement steps include: S1. Obtain fault flight data of vertical take-off and landing aircraft, and establish a directed acyclic graph between flight characteristics and fault types through information flow causal discovery method; S2. Calculate the mutual information between nodes in the directed acyclic graph and convert the mutual information into the causal strength of each node to causally strengthen each node; S3, inputting the causally strengthened nodes into the convolutional neural network for convolution operation, and using the causal weights for dynamic weighting during the convolution operation to extract local features; S4, concatenating the local features with the causal features in the directed acyclic graph to obtain concatenated features; Use a multi-layer perceptron to process the concatenated features to obtain the adjacency matrix of the nodes; S5. Use graph convolutional network to fuse the concatenated features and adjacency matrix to obtain high-order features; S6. Input the high-order features into the fully connected network to complete adversarial training with the convolutional neural network and the graph convolutional network as feature extractors and the fully connected network as a fault classifier; the feature extractor and the fault classifier constitute a fault prediction model; The specific steps of step S2 are as follows: S21. Based on the time series data of each node in the directed acyclic graph, the information entropy of each node is calculated; the calculation formula of the information entropy is as follows: In the formula, H(X i ) represents the i-th node X in the directed acyclic graph i Information entropy of Represents node X i The value at time t; express The probability distribution of ; log represents the logarithmic function with base 10; S22. Use information entropy to calculate the mutual information between nodes. The calculation formula of mutual information is as follows: I(X i ,X j )=H(X i )+H(X j )-H(X i ,X j ); In the formula, I(X i ,X j ) represents the i-th node X in the directed acyclic graph i and the jth node X j The mutual information between j ) represents the jth node X in the directed acyclic graph j The information entropy of H(X i ,X j ) represents node X i and node X j The joint information entropy between S23, arranging the calculated mutual information in order to form a causal relationship matrix M in the form of a square matrix, wherein each element represents the mutual information between two corresponding nodes; S24, performing a normalization operation on the causal relationship matrix M to obtain a causal strength matrix M'; S25, combining the causal strengths in the causal strength matrix M' to causally strengthen each node; S26, calculating the value of each node after causal reinforcement at each time through the calculation formula of causal reinforcement; The calculation formula of causal reinforcement is as follows: In the formula, Represents node X after causal reinforcement i The value at time t; M″ ij Represents the i-th node X in the directed acyclic graph i and the jth node X j the causal strength between them; Represents node X j The value at time t.
2. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 1, characterized in that: The calculation formula for each element in the causal intensity matrix M" is as follows: In the formula, w j Represents node X j The weight factor of j Represents node X j The standard deviation of the time series data; μ j Represents node X j The mean of the time series data; M ij Represents the i-th node X in the directed acyclic graph i and the jth node X j The mutual information between them, that is, I(X i ,X j ); α represents the scaling factor; M min,j Represents the node X j The minimum value of all relevant mutual information; M max,j Represents the node X j The maximum value of all relevant mutual information.
3. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 2, characterized in that: The specific steps of step S3 are as follows: S31, arranging the values of each node at each time after causal reinforcement into a data set at the corresponding time; S32, inputting the data sets at each moment into the convolutional neural network in sequence for convolution operation, so as to extract the local features between each node through the convolutional neural network.
4. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 3 is characterized in that: The convolution operation in a convolutional neural network is represented as follows: In the formula, Represents node X i Local features at time t; represents the kth node X k The value at time t; express and The similarity between express and The similarity between express and The softmax function calculation value between ; n represents the total number of nodes; Represents node X after causal reinforcement j The value at time t; w represents the convolution kernel of the convolutional neural network; b represents the bias term of the convolutional neural network; F(·) represents the convolution operation; exp(·) represents the exponential function with the natural constant e as the base.
5. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 4, characterized in that: The specific steps of step S4 are as follows: S41. The local features and the causal features in the directed acyclic graph are spliced together to obtain spliced features. The spliced features are expressed as follows: In the formula, express The splicing features formed after splicing; Represents the node X in the directed acyclic graph at time t i The causal characteristics of ; [·,·] represents the concatenation operation; S42, then performing nonlinear transformation on the splicing features to obtain nonlinear features; The nonlinear characteristics are expressed as follows: In the formula, express The nonlinear characteristics of merge represents the nonlinear bias term; W merge represents nonlinear weight; σ(·) represents activation function; S43. Use a multi-layer perceptron to transform the concatenated features into high-dimensional features. The high-dimensional features are expressed as follows: Where Z represents the feature matrix formed by the combination of various nonlinear features; MLP represents multi-layer perceptron; Represents the high-dimensional features corresponding to Z; S44. Multiply the feature matrix by its transposed matrix to obtain an adjacency matrix; the adjacency matrix is expressed as follows: In the formula, A represents the adjacency matrix; normal represents the matrix multiplication operation; express The transposed matrix of S45, selecting the first r nearest neighbors of each node to form a sparse adjacency matrix A′; A′=Top-r(A); In the formula, Top-r means taking the first r nearest neighbors.
6. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 5, characterized in that: The specific content of step S5 is as follows: S51, input the feature matrix and the sparse adjacency matrix into the graph convolutional network; S52, using two layers of receptive field convolution to perform convolution operation to obtain corresponding high-order features; In the formula, H0 represents the features learned by the first layer of receptive field convolution; W0 represents the weight matrix of the first layer of receptive field convolution; MRFConv represents the receptive field convolution; W1 represents the weight matrix of the second layer of receptive field convolution; H1 represents the features learned by the second layer of receptive field convolution, that is, the high-order features.
7. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 6, characterized in that: The objective function in the adversarial training process is as follows: THE Total =L C +λL DA ; Where, L Total represents the objective function; L C represents the classification loss; L DA represents the domain alignment loss; λ represents a hyperparameter.
8. The method for enhancing the generalization of vertical take-off and landing aircraft diagnosis based on causal weight according to claim 7, characterized in that: The calculation formulas for each loss are specifically expressed as follows: In the formula, x s represents the source domain sample; x t represents the target domain sample; y represents the fault type, C(x s ,y) represents the fault classifier model, inputs the source domain sample, and outputs the probability that the source domain sample belongs to fault type y, C(x t ,y) represents the fault classifier model, which inputs the target domain sample and outputs the probability that the target domain sample belongs to fault type y; represents the source domain classifier output; represents the target domain classifier output; Where D(·) represents the domain discriminator; represents the probability that the domain discriminator judges the generated source domain sample as fake; represents the loss of target domain samples.
Citation Information
Patent Citations
Complex working condition gear fault diagnosis method based on matching element learning under few samples
CN116593157A
Power equipment fault diagnosis method and system based on causal knowledge guidance
CN118964900A