A method for fault detection of aviation inverter based on information fusion algorithm

By using information fusion algorithm and Bayesian deep learning model in aircraft converter fault detection, combined with variational autoencoder and Bayesian theorem, the accuracy and reliability problems of fault detection in the existing technology are solved, and accurate diagnosis and early warning of converter faults are achieved.

CN119622648BActive Publication Date: 2025-05-09NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510156813.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-09
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and timely detect potential failures of aircraft converters under complex operating conditions, and deep learning models have problems with insufficient overfitting and generalization capabilities in fault detection.

Method used

Using an information fusion algorithm method, a Bayesian deep learning model is constructed through a variational autoencoder, and Bayesian theorem and regularization terms are introduced during model training and optimization to improve the generalization ability of the model and its robustness to uncertainty.

Benefits of technology

It realizes accurate diagnosis and early warning of aviation converter faults, improves the accuracy and reliability of fault detection, and enhances the operation and maintenance guarantee level and flight safety of aviation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_158
    Figure SMS_158
  • Figure QLYQS_41
    Figure QLYQS_41
  • Figure QLYQS_44
    Figure QLYQS_44
Patent Text Reader

Abstract

The present invention discloses an aviation inverter fault detection method based on an information fusion algorithm, comprising collecting operating data of the aviation inverter and constructing a data set; extracting time information and space information from the data in the data set, and fusing time features and space features based on a variational autoencoder to obtain spatiotemporal data; constructing a Bayesian deep learning model, and training and optimizing the Bayesian deep learning model; inputting the spatiotemporal data into the trained and optimized Bayesian deep learning model for processing, outputting a numerical value representing the probability of a fault, and a variance based on the probability numerical value as an uncertainty estimate; presetting a probability threshold, and when the output probability is greater than the threshold, determining that the inverter has a fault; otherwise, determining that it is in a normal operating state; providing more accurate and reliable diagnostic results for fault detection, and having strong adaptability and robustness in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft equipment fault detection, and in particular to an aviation inverter fault detection method based on an information fusion algorithm. Background Art

[0002] With the rapid development of aviation technology, aircraft inverters, as key power conversion components, are crucial to the safety and reliability of the entire aviation system. In actual flight, inverters may fail due to various complex factors, such as long-term wear and tear, electrical stress, and harsh environmental influences. Traditional inverter fault detection methods are mainly based on simple electrical parameter monitoring, such as threshold judgment of single signals such as voltage and current. This method is difficult to accurately and timely detect potential inverter faults under complex working conditions.

[0003] In recent years, deep learning technology has shown certain potential in the field of fault detection. However, when dealing with aircraft inverter fault detection problems, existing deep learning algorithms often ignore the rich information correlation of fault signals in time and space dimensions. The fault characteristics of the inverter not only show dynamic changes in the time series, but also have mutual influence and correlation in different circuit structures and component spatial positions. Simple time series deep learning models or spatial feature extraction methods cannot fully and effectively capture these complex fault information.

[0004] The Bayesian method has unique advantages in uncertainty reasoning and probabilistic modeling, and can quantitatively evaluate the uncertainty of the model, providing a new idea for solving the problems of overfitting and insufficient generalization ability of deep learning models in fault detection. However, the research on combining Bayesian theory with deep learning and applying it to the time-space information fusion of aircraft inverter fault detection is still in its infancy. There is a lack of mature and efficient algorithms to achieve accurate diagnosis and early warning of inverter faults, which seriously restricts the operation and maintenance level of aviation equipment and flight safety. Therefore, an aviation inverter fault detection method based on information fusion algorithm is urgently needed to solve the above problems. Summary of the invention

[0005] The purpose of the present invention is to provide an aviation inverter fault detection method based on an information fusion algorithm, which can effectively solve the problems existing in the above-mentioned prior art.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solution: an aviation inverter fault detection method based on information fusion algorithm comprises the following steps:

[0007] Collect the operating data of aviation inverters and build data sets;

[0008] Extract time information and space information from the data set, and fuse time features and space features based on variational autoencoders to obtain spatiotemporal data.

[0009] Build, train and optimize Bayesian deep learning models;

[0010] The spatiotemporal data is input into the trained and optimized Bayesian deep learning model for processing, and the output is a numerical value representing the probability of failure, as well as the variance based on the probability value as uncertainty estimation;

[0011] A probability threshold is preset, and when the output probability is greater than the threshold, it is determined that the converter is faulty; otherwise, it is determined to be in normal operation.

[0012] Preferably, the data in the data set are preprocessed, including:

[0013] Data cleaning: Assume that a variable time series is , calculate the median of the sequence :

[0014] ;

[0015] For each data point , calculate its absolute deviation , and then calculate the median absolute deviation ;

[0016] like , then Determined as an outlier, where i is the data point number, and corrected in the following way:

[0017] ;

[0018] in is a symbolic function, which means taking The symbol, is a constant;

[0019] Data filtering, including:

[0020] Strain step filtering:

[0021] For the input sensor data sequence , the output of the filter for:

[0022] ;

[0023] in, is the filter coefficient at time n, and N is the filter order;

[0024] The update formula of the filter coefficient is:

[0025]

[0026] in, is the estimation error, is the expected response (the original pure signal in the absence of noise), is the step size factor, and its update formula is:

[0027] ;

[0028] in, and It is a constant adjusted according to the actual situation. The step size can be dynamically adjusted according to the energy of the input signal, so that the filter can have better performance under different signal conditions, converge quickly and maintain a low steady-state error;

[0029] Data normalization processing:

[0030] A normalization algorithm based on probability distribution is used; first, the sensor data is calculated The mean and standard deviation :

[0031] ;

[0032] ;

[0033] Then, normalize the data:

[0034] ;

[0035] in, is the inverse cumulative distribution function of the standard normal distribution, yes The cumulative distribution function of is a small positive number (here =0.001), which is used to avoid the influence of extreme values; this normalization method takes into account the probability distribution characteristics of the data. Compared with traditional linear normalization, it can better retain the characteristic information of some non-uniformly distributed data, and can improve the stability and convergence speed of the model in subsequent model training;

[0036] Preferably, the time information extraction includes:

[0037] Skewness calculation based on the median:

[0038] ;

[0039] in, Express expectations, is the median, is the median absolute deviation;

[0040] Kurtosis calculation based on interquartile range:

[0041] ;

[0042] in is the second quartile, i.e. the median, and are the first and third quartiles, respectively;

[0043] DTW distance calculation:

[0044] For two lengths respectively and Time Series and , build a The distance matrix ,in, ;

[0045] Define a cumulative distance matrix ,initialization , calculated by the following recursive formula:

[0046] ;

[0047] The DTW distance is ;

[0048] Wavelet energy entropy calculation:

[0049] For time series Perform wavelet transform to obtain wavelet coefficients ,in is the scale parameter, is the translation parameter;

[0050] For each scale , calculate the wavelet energy at this scale :

[0051] ;

[0052] Calculate wavelet energy entropy :

[0053] ;

[0054] in is the total number of scales.

[0055] Preferably, the spatial information extraction is specifically:

[0056] Spatial feature construction based on physical structure: including heat conduction and temperature gradient features and electrical parameter difference features. The heat conduction and temperature gradient features are obtained by analyzing the temperature distribution inside the converter and the temperature difference between adjacent positions. The electrical parameter difference features are obtained by calculating the electrical parameter difference between adjacent components in the converter.

[0057] Spatial information extraction based on graph theory:

[0058] The converter graph model is constructed by taking the components or sensor nodes of the converter as vertices and the electrical connections, thermal conduction relations or physical adjacent relations between the components as edges;

[0059] Calculate the Laplacian matrix of the graphical model and extract the eigenvalues ​​and eigenvectors of the Laplacian matrix as spatial features; and

[0060] Aggregate and update the features of sensor nodes based on graph convolutional neural networks to extract representative spatial features;

[0061] Preferably, the basic model of the Bayesian deep learning model is a hybrid architecture constructed by combining a convolutional neural network and a long short-term memory network; and

[0062] In the convolutional neural network, a dilated convolution is used to expand the receptive field of the convolution kernel;

[0063] In the long short-term memory network, the attention mechanism is introduced:

[0064] Calculate attention weight: Assume is the hidden state sequence output by the long short-term memory network, and the attention weight vector is calculated through a fully connected layer and a softmax function :

[0065] ;

[0066] ;

[0067] in, and are the weights and biases of the fully connected layer, and the softmax function is used to normalize the attention scores;

[0068] The hidden states are weighted summed according to the attention weights to obtain the context vector :

[0069] ;

[0070] Finally, the vector With the last hidden state of LSTM Splicing, as the final output feature of the model;

[0071] Enable Bayesian deep learning models to dynamically assign weights based on the importance of input data;

[0072] A regularization term is added to the Bayesian deep learning model.

[0073] Preferably, the Bayesian deep learning model introduces Bayesian theorem, regards model parameters as random variables, and calculates uncertainty estimates of model parameters at each forward propagation calculation of the model output based on the Bayesian back-propagation algorithm according to the observed data. During the back-propagation process, the posterior distribution of the parameters is updated according to the uncertainty of the parameters and the gradient of the loss function.

[0074] Preferably, the training of the Bayesian deep learning model includes:

[0075] Dividing the data into a training set, a validation set, and a test set, and dividing the proportions of the training set, the validation set, and the test set in turn based on a random division method;

[0076] Perform data enhancement operations on the data in the dataset, including adding noise, shifting the time series, and scaling the time series.

[0077] Preferably, the optimization of the Bayesian deep learning model is specifically as follows: updating the model parameters according to the gradient information of the loss function to minimize the loss function so that the model gradually converges to a preset solution, and the updated model parameters are:

[0078] ;

[0079] in, are model parameters, is the learning rate, , , is the second-order moment estimate, The first moment estimate, and is the decay rate hyperparameter, Is a positive number.

[0080] Preferably, the differences in the features of the middle layer of the Bayesian deep learning model under fault and normal conditions are analyzed, characteristic patterns related to the fault are extracted, and the fault type or the location where the fault occurs is determined.

[0081] Preferably, the Bayesian deep learning model is updated in real time, including:

[0082] Monte Carlo dropout is used to calculate the uncertainty index for new data, and the new data samples are sorted according to the uncertainty index, and the first several samples with the highest uncertainty are selected to form a selected sample set;

[0083] Based on a selected sample set, an incremental learning approach is adopted. When new data samples arrive, the model adjusts its own parameters by learning the new data based on the existing parameters to update the model.

[0084] Beneficial effects: In the present invention, by fusing temporal features and spatial features based on a variational autoencoder, the complex relationship between temporal and spatial features can be effectively captured, and the features are optimized and fused in the encoding-decoding process, and the generated fusion features have stronger expression ability and discrimination power, providing an effective feature fusion strategy for aviation inverter fault detection; wherein the Bayesian deep learning model is constructed, trained and optimized, which can effectively process the spatiotemporal data of aviation inverters, and at the same time, taking into account the uncertainty of parameters, provide more accurate and reliable diagnostic results for fault detection, and have strong adaptability and robustness in practical applications;

[0085] In addition, real-time updating of the model can enable the fault detection model to continuously adapt to new data and operating condition changes during the long-term operation of the aviation inverter, maintain a good performance state, improve the accuracy and reliability of fault detection, and provide strong protection for the safe and stable operation of the aircraft. DETAILED DESCRIPTION

[0086] The embodiments of the present invention are described below in conjunction with the embodiments of the present invention. The terms used in the embodiments of the present invention are only used to explain the specific embodiments of the present invention, and are not intended to limit the present invention. The embodiments of the present application are described below.

[0087] Embodiment: A method for detecting faults of an aviation inverter based on an information fusion algorithm comprises the following steps:

[0088] Collect the operating data of aviation inverters and build data sets;

[0089] In one case, data is acquired by placing sensors on the converter, where:

[0090] The voltage, current, vibration and temperature data of the power module, capacitor, inductor and control circuit of the aviation inverter are obtained through sensors. The fast-changing parameters such as voltage, current and vibration data are sampled at a high frequency to ensure that the details of the changes in the inverter's operating status can be captured. As a relatively slow-changing parameter, temperature uses a lower sampling frequency to avoid excessive data redundancy and reduce the amount of calculation.

[0091] Preprocess the data in the dataset, including:

[0092] Data cleaning: Assume that a variable time series is , calculate the median of the sequence :

[0093] ;

[0094] For each data point , calculate its absolute deviation , and then calculate the median absolute deviation ;

[0095] like , then Determined as an outlier, where i is the data point number, when the data point When it is judged as an outlier, the formula will be The relationship with the median M is corrected to a reasonable value and corrected in the following way:

[0096] ;

[0097] in:

[0098] is the median of the time series;

[0099] is the Median Absolute Deviation, which is the median of the absolute deviations of the data points from the median;

[0100] is a constant used to control the range of outliers;

[0101] is a symbolic function, which means taking the sign of (positive or negative);

[0102] (1) If Greater than ,illustrate is an outlier that is too large. The corrected value will be set to ;

[0103] (2) If Less than ,illustrate is an outlier that is too small. The corrected value will be set to ;

[0104] In this way, outliers will be corrected to a reasonable range to avoid adverse effects on subsequent data analysis and model training;

[0105] Data filtering, including:

[0106] Strain step filtering: For the input sensor data sequence , the output of the filter for:

[0107] ;

[0108] in, is the filter coefficient at time n, and N is the filter order;

[0109] The update formula of the filter coefficient is:

[0110]

[0111] in, is the estimation error, is the expected response (the original pure signal in the absence of noise), is the step size factor, and its update formula is:

[0112] ;

[0113] in, and It is a constant adjusted according to the actual situation. The step size can be dynamically adjusted according to the energy of the input signal, so that the filter can have better performance under different signal conditions, converge quickly and maintain a low steady-state error;

[0114] Data normalization: A normalization algorithm based on probability distribution is used; first, the sensor data is calculated The mean and standard deviation :

[0115] ;

[0116] ;

[0117] Then, normalize the data:

[0118] ;

[0119] in, is the inverse cumulative distribution function of the standard normal distribution, yes The cumulative distribution function of is a small positive number (here =0.001), which is used to avoid the influence of extreme values; this normalization method takes into account the probability distribution characteristics of the data. Compared with traditional linear normalization, it can better retain the characteristic information of some non-uniformly distributed data, and can improve the stability and convergence speed of the model in subsequent model training;

[0120] Through the above data collection and preprocessing, noise and outliers can be effectively removed, the quality and stability of data can be improved, and more reliable basic data can be provided for subsequent spatiotemporal information extraction and fault detection, thereby improving the performance and accuracy of the entire algorithm in aviation inverter fault detection;

[0121] Extract temporal information and spatial information from the data in the dataset;

[0122] Time information extraction includes:

[0123] The skewness calculation based on the median can reduce the impact of extreme values ​​on the skewness calculation and more accurately reflect the asymmetry of data distribution. Specifically:

[0124] ;

[0125] in, Express expectations, is the median, is the median absolute deviation;

[0126] Kurtosis calculation based on interquartile range:

[0127] ;

[0128] in is the second quartile, i.e. the median, and are the first and third quartiles, respectively;

[0129] Compared with the traditional kurtosis formula:

[0130] ;

[0131] The traditional kurtosis formula is not sensitive enough to the data with heavy-tail distribution. The kurtosis calculation based on the interquartile range in this embodiment can better capture the distribution characteristics of the data near the peak value, and is more effective in detecting abnormal spike signals in the converter.

[0132] DTW distance calculation:

[0133] For two lengths respectively and Time Series and , construct a The distance matrix ,in, ;

[0134] Define a cumulative distance matrix ,initialization , calculated by the following recursive formula:

[0135]

[0136] The DTW distance is ; On this basis, the mean, standard deviation and other statistics of DTW features are calculated as part of the time domain features; for example, for a set of time series , its DTW mean feature is:

[0137]

[0138] in is the average sequence of the group of time series; this DTW feature can effectively measure the similarity and change trend between time series, and plays an important role in detecting the operating state changes of the converter under different working conditions;

[0139] Wavelet energy entropy calculation:

[0140] For time series Perform wavelet transform to obtain wavelet coefficients ,in is the scale parameter, is the translation parameter;

[0141] For each scale , calculate the wavelet energy at this scale :

[0142] ;

[0143] Calculate wavelet energy entropy :

[0144] ;

[0145] in is the total number of scales; wavelet energy entropy can reflect the energy distribution of the signal at different scales, and has high sensitivity for detecting transient faults in the converter and complexity changes of the signal. As a time domain feature, it can provide more information about the signal characteristics, which helps to improve the accuracy and reliability of fault detection.

[0146] For spatial information extraction, this embodiment includes spatial feature construction based on physical structure and spatial information extraction based on graph theory;

[0147] For the construction of spatial features based on physical structures, including:

[0148] Heat conduction and temperature gradient characteristics: During the operation of aviation inverters, the generation and transfer of heat follows the principle of heat conduction. By analyzing the temperature distribution inside the inverter and the temperature difference between adjacent positions, information about the inverter's heat dissipation status and potential faults can be obtained. For example, local overheating may indicate a component failure or a blockage in the heat dissipation system.

[0149] In one case, considering the heat dissipation structure and heat conduction path of the converter, it is assumed that the temperature distribution of the converter satisfies the heat conduction equation:

[0150] ;

[0151] in, is the material density, is the specific heat capacity, is the temperature field, is the thermal conductivity, is the heat source term; in the discrete case, for the temperature sensor nodes on the two-dimensional converter plane , the temperature variation with time can be approximately expressed as:

[0152] ;

[0153] in, is the time step, and is the spatial step length, Indicates the number of time steps.

[0154] Calculate the temperature gradient characteristics, Temperature gradient in direction for:

[0155] ;

[0156] exist Temperature gradient in direction for:

[0157] ;

[0158] These temperature gradient characteristics can reflect the direction and rate of heat transfer inside the converter, and play an important role in detecting heat dissipation system failures or local overheating problems;

[0159] Electrical parameter difference characteristics: When the electrical components in the converter are operating normally, their electrical parameters remain relatively stable; when a component fails, its electrical parameters will change, and this change may be reflected in the electrical parameter differences of adjacent components; the failure of the converter component module may cause its output voltage or current to be inconsistent with that of the adjacent modules; for the electrical components in the converter, such as power modules, capacitors, inductors, etc., calculate the electrical parameter differences between adjacent components;

[0160] In one case, for two adjacent power modules

[0161] A and B, whose output voltages are and , the currents are and , then the voltage difference , current difference ; Calculate the mean and variance of these electrical parameter differences as part of the spatial features; For example, for a set of adjacent power modules, the voltage difference sequence , its mean ,variance ; These characteristics can reflect the consistency and stability of the internal electrical performance of the converter, and play an important role in detecting the failure or performance degradation of electrical components;

[0162] For spatial information extraction based on graph theory, including:

[0163] Constructing a graph model of the converter: Abstract each component or sensor node of the converter as a graph vertex. The electrical connection, heat conduction relationship or physical proximity relationship between the components are represented by the edges of the graph. Graph theory is used to analyze the spatial structure and information propagation path inside the converter. By constructing a graph model, we can more intuitively understand the relationship between the various parts inside the converter, providing a basis for subsequent feature extraction and fault diagnosis.

[0164] In one case, each component or sensor node of the converter is considered as a vertex of the graph , electrical connections, thermal conduction relationships, or physical proximity relationships between components are considered as edges of the graph. and Connected , construct an undirected graph ;

[0165] Defining the adjacency matrix , if the vertex and If connected, ,otherwise ; At the same time, define the degree matrix ,in , that is, the vertex degree;

[0166] Graph Laplace feature extraction: The Laplace matrix of a graph contains the structural information of the graph. Its eigenvalues ​​and eigenvectors can reflect the topological characteristics of the graph and the degree of connection between nodes. In the fault detection of aviation inverters, by analyzing the Laplace features, it can be found that when a fault occurs, the structure of the graph will change, resulting in changes in the eigenvalues ​​and eigenvectors. These changes can be used as a basis for fault diagnosis and locate the area where the fault occurs.

[0167] Computational graph Laplacian matrix , perform eigendecomposition on the Laplacian matrix ,in is the eigenvector matrix, is a diagonal matrix whose diagonal elements are eigenvalues:

[0168] Extract the eigenvalues ​​and eigenvectors of the Laplacian matrix as spatial features; use the previous The smallest non-zero eigenvalue and its corresponding eigenvector ,These features can reflect the structural characteristics of the graph and the ,closeness of the connections between the nodes; by analyzing the size and ,distribution of each element in the eigenvector, the nodes that play a key role in the ,structure of the graph are determined. When a fault occurs, the eigenvalues ​​and eigenvectors corresponding to these key ,nodes may change significantly, which helps locate the source of the fault.

[0169] Graph convolutional neural network (GCN) feature extraction: Graph convolutional neural network can perform convolution operations on graph structured data and automatically learn the spatial dependencies and feature representations between nodes. In aviation inverter fault detection, GCN can aggregate and update the features of sensor nodes based on the inverter graph model to extract representative spatial features. These features can comprehensively reflect the mutual influence between the components inside the inverter and improve the accuracy and robustness of fault detection.

[0170] In one case, for the graph signal (in is the number of nodes, is the signal dimension), the operation of the graph convolutional layer is defined as follows:

[0171] ;

[0172] in, is the identity matrix It is The feature map of the layer, It is The weight matrix of the layer, is the activation function (ReLU).

[0173] Through multi-layer graph convolution operations, spatial features at different levels can be extracted. These features can obtain the complex spatial dependencies and topological structure information inside the converter, and have high accuracy and adaptability for fault detection. For example, after several layers of graph convolution, the final feature map obtained is It can be connected as input to the subsequent fault detection model, where the feature vector of each node contains the information fusion of itself and adjacent nodes, which can more comprehensively reflect the spatial state information of the converter.

[0174] The above spatial information extraction method is based on the physical structure and graph theory of the converter, and can provide rich and accurate spatial information for aviation converter fault detection, thereby improving the accuracy and reliability of fault detection.

[0175] Based on the variational autoencoder, the temporal and spatial features are fused to obtain spatiotemporal data;

[0176] The temporal and spatial features are encoded into a low-dimensional latent space through a variational autoencoder (VAE). The feature fusion and optimization are achieved through the operation of the latent space. The fused feature representation is then decoded and restored to mine the complex relationship between temporal and spatial features, and a more discriminative feature vector is generated for aviation inverter fault detection.

[0177] In one case: Let the time feature vector be , the spatial eigenvector is , the input vector concatenated into them is , for a containing A dataset of samples , the working process of VAE is as follows:

[0178] Encoder part: The encoder takes the input vector The mean of the latent space is mapped to the latent space through two fully connected layers. and variance ; The specific formula is:

[0179] Mean , is the weight matrix, is the bias vector, is the dimension of the latent space;

[0180] variance , and are the corresponding weights and biases;

[0181] In order to ensure that the variance and If it is non-negative, perform softplus transformation on it: ,in ;

[0182] Sampling from the latent space ,in ,and is the random noise of standard normal distribution, Represents element-wise multiplication; makes the encoding process random, which helps to obtain richer potential representations;

[0183] Decoder part: The decoder converts the latent variables Decoded back into the reconstructed feature vector through one or more fully connected layers ,in ,and is the weight matrix, is the bias vector.

[0184] Loss function and training: The goal of VAE is to maximize the evidence lower bound (ELBO), which is formulated as:

[0185] ;

[0186] in is the KL divergence, is the prior distribution (set to standard normal distribution ); the first item represents the reconstruction loss, that is, the reconstructed data generated by the expected decoder As close to the original input as possible , use the mean square error (MSE) to approximate the calculation; the mean square error calculation formula here is:

[0187] ;

[0188] Item 2 is the KL divergence term, which is used to measure the potential distribution of the encoder output With the prior distribution The difference is calculated as:

[0189] ;

[0190] The back propagation algorithm is used to maximize ELBO to train the parameters of VAE. and ;

[0191] Feature fusion and output: trained VAE, fused spatiotemporal feature vector By reconstructing the feature vector output by the decoder After transformation, a fully connected layer is used for feature extraction:

[0192] ,in is the feature dimension after fusion.

[0193] The VAE-based spatiotemporal feature fusion method can effectively capture the complex relationship between temporal and spatial features by learning the distribution of the latent space, and optimizes and fuses the features in the encoding-decoding process. The generated fused features have stronger expressiveness and discriminative power, providing an effective feature fusion strategy for aviation inverter fault detection.

[0194] A Bayesian deep learning model is constructed, whose basic model is a hybrid architecture combining convolutional neural network (CNN) and long short-term memory network (LSTM); CNN is good at automatically extracting local spatial features of data, and performs convolution operations by sliding convolution kernels on data to capture the spatial correlation between sensor data at different locations of the inverter; LSTM is specifically used to process long-term dependencies in time series data, and its internal gating mechanism can selectively remember or forget historical information, thereby effectively modeling and analyzing features in the time dimension. Combining the two can give full play to their advantages in spatiotemporal feature extraction and better deal with complex data patterns in aviation inverter fault detection;

[0195] In one case, suppose the input spatiotemporal data is ,in is the batch size, C is the number of channels, and are the height and width of the spatial dimension respectively.

[0196] CNN part: Convolutional layer and pooling layer are used to extract spatial features; the convolution operation formula is:

[0197] ;

[0198] in It is Tier The output of the feature map is It is Tier The convolution kernel and The convolution kernel weights of the input feature map, It is Tier Input feature maps, is the bias term, Represents the convolution operation;

[0199] The pooling layer uses maximum pooling to reduce the resolution of the feature map. The maximum pooling formula is:

[0200] ;

[0201] in is the output after pooling, is the pooling step size, It is the pooling area;

[0202] In this embodiment, in the CNN part, dilated convolution is used to expand the receptive field of the convolution kernel, so as to obtain more extensive spatial information without increasing the number of parameters, thereby better capturing the long-distance dependencies between different components in the converter;

[0203] For dilated convolution: the sampling interval of its convolution kernel is (void ratio), the convolution operation formula becomes: ;in Represents the input feature map in the spatial dimension Sampling for intervals;

[0204] LSTM part: The spatial feature map extracted by CNN is flattened and used as the input sequence of LSTM; for the LSTM unit at time The calculation formula is as follows:

[0205] Input Gate:

[0206] ;

[0207] Forget Gate:

[0208] ;

[0209] Output Gate:

[0210] ;

[0211] Memory unit:

[0212] ;

[0213] Hidden state:

[0214] ;

[0215] in, is the input at the current moment, is the hidden state at the previous moment, and are the corresponding weight matrices and bias vectors, is the sigmoid function, Indicates that the above formula is element-wise multiplication, and tanh is the hyperbolic tangent function;

[0216] Dynamically allocate weights to improve the diagnostic accuracy and efficiency of the model; at the same time, in order to prevent overfitting, regularization terms are added to the model to enhance the generalization ability of the model;

[0217] LSTM with attention mechanism:

[0218] First, calculate the attention weight, assuming is the hidden state sequence output by LSTM, and the attention weight vector is calculated through a fully connected layer and softmax function :

[0219] ;

[0220] ;

[0221] in, and are the weights and biases of the fully connected layer, and the softmax function is used to normalize the attention scores;

[0222] Then, the hidden states are weighted summed according to the attention weights to obtain the context vector :

[0223] ;

[0224] Finally, the vector With the last hidden state of LSTM Splicing, as the final output feature of the model;

[0225] The Bayesian theorem is introduced into the Bayesian deep learning model, and the model parameters are regarded as random variables. Based on the observed data, the uncertainty estimate of the model parameters is calculated at each forward propagation when the model output is calculated based on the Bayesian back-propagation algorithm. During the back-propagation process, the posterior distribution of the parameters is updated according to the uncertainty of the parameters and the gradient of the loss function.

[0226] In one case: Let the parameters of the model be , the observed data is , the prior distribution is , the likelihood function is , then according to Bayes' theorem, the posterior distribution of the parameter is:

[0227]

[0228] in, is the marginal distribution of the data;

[0229] During the model training process, the Bayesian back propagation (BBP) algorithm is used to update the posterior distribution of the model parameters. The traditional gradient calculation is modified by considering the uncertainty of the parameters in the back propagation process. Specifically, when the model output is calculated in each forward propagation, the uncertainty estimate of the model parameters (usually expressed by variance or covariance) is calculated at the same time. During the back propagation process, the posterior distribution of the parameters is updated according to the uncertainty of the parameters and the gradient of the loss function. This method enables the model to gradually adjust the distribution of parameters during the training process, thereby improving the generalization ability of the model and its robustness to uncertainty.

[0230] Assume the loss function of the model is , for a weight matrix and the bias vector The neural network layer, in the forward propagation, calculates the output , and calculate the variance of weights and biases at the same time and ;

[0231] In back propagation, according to the loss function Output Gradient And the variance of the parameters, calculate the gradient update of weights and biases:

[0232] For weight : ;

[0233] For bias b: ;in is the learning rate, is a regularization parameter. In this way, the model can gradually adjust the distribution of parameters during training, thereby improving the generalization ability of the model and its robustness to uncertainty.

[0234] The Bayesian deep learning model constructed through the above content can effectively process the spatiotemporal data of aviation inverters, while taking into account the uncertainty of parameters, providing more accurate and reliable diagnostic results for fault detection, and has strong adaptability and robustness in practical applications.

[0235] Train and optimize Bayesian deep learning models;

[0236] Training of the Bayesian deep learning model involves:

[0237] Dividing the data into a training set, a validation set, and a test set, and dividing the proportions of the training set, the validation set, and the test set in turn based on a random division method;

[0238] In one case, suppose the original data set contains samples, marked as ,in represents the input spatiotemporal data samples, Indicates the corresponding fault label;

[0239] The random division method is used to set the training set ratio to , the validation set accounts for , the test set accounts for ;

[0240] The sample index collection of the training set It is obtained by sampling, and its size is ; Use evenly distributed A random number generator in the range, selected non-repeated indexes, and the samples corresponding to these indexes constitute the training set and

[0241] Sample index collection of the validation set Select from the remaining samples, the size is , the validation set is and

[0242] The test set consists of the remaining samples, i.e. and .

[0243] Perform data enhancement operations on the data in the dataset, including:

[0244] Add noise: For the input sample , the operation of adding Gaussian noise can be expressed as ,in Indicates that the mean is 0 and the covariance matrix is Gaussian noise, is the identity matrix; for example, for a two-dimensional spatiotemporal data sample (in is the height, is the width, is the number of channels), noise Each element of ( Independently sample from a Gaussian distribution and then add noise to the original sample to get the enhanced sample ;

[0245] Shift the time series: If Step translation ( is an integer, which can be positive or negative), and the enhanced sequence is obtained . Here we use the modulo operation Ensure that the index after translation is within the valid range, so as to achieve circular translation and keep the sequence length unchanged;

[0246] Scaling a time series: Given a scaling factor , the scaled time series is ; If the original time series represents the voltage or current value of the converter over a period of time, the scaling operation can simulate the gain change of the sensor or the change of the voltage or current amplitude under different working conditions.

[0247] Optimization of Bayesian deep learning models includes:

[0248] Update the model parameters according to the gradient information of the loss function to minimize the loss function and make the model gradually converge to the preset solution;

[0249] In one case, let the model be Fault type The predicted probability is is the number of fault types , the real label is encoded using one-hot encoding , and for each sample , there is only one , the rest are 0 , then the cross entropy loss function is:

[0250] ;

[0251] This loss function measures the difference between the probability distribution predicted by the model and the probability distribution of the true label; for each sample, only the category corresponding to the true label is is 1, and the others are 0, so the loss function mainly focuses on the model's prediction probability of the true fault type ; When the model predicts a higher probability of the true fault type, the loss value is smaller, thereby guiding the model to learn feature representations that can accurately classify different fault types;

[0252] Add a regularization term:

[0253] For the parameters of the model (including the weight matrix and bias vector in the neural network, etc.), the regularization term is:

[0254] ;

[0255] The summation here is performed on all parameters of the model. is the regularization parameter, which controls the strength of the regularization term. The value will impose stronger constraints on the parameters, preventing them from being too large and thus avoiding overfitting of the model.

[0256] The final loss function is: ,in For cross picking loss;

[0257] In this embodiment, the optimization process involves the first-order moment estimation of two parameter vectors and the second moment estimate , and a time step ; In each iteration For the model parameters , whose gradient is , is the value of the loss function under the current parameter value; update the first-order moment estimate:

[0258] ;

[0259] Update the second moment estimate:

[0260] ;

[0261] in and is the decay rate hyperparameter, for example, ; To correct the bias, calculate the bias-corrected first-order moment estimate and the second moment estimate ;

[0262] Finally, update the model parameters:

[0263] ;

[0264] in is the learning rate, is a small positive number (such as 1e-8) to prevent the denominator from being zero. This adaptive adjustment of the learning rate enables the optimization algorithm to use different learning rates for different parameters according to the gradient history information of the parameters during training, so that it can converge to a better parameter value faster and is relatively insensitive to the choice of hyperparameters, showing better performance and stability.

[0265] Through the above, the Bayesian deep learning model can be effectively trained and optimized, so that it has high accuracy, generalization ability and stability in aviation inverter fault detection tasks, thereby improving the reliability of fault detection.

[0266] The spatiotemporal data is input into the trained and optimized Bayesian deep learning model for processing, and a numerical value representing the probability of a fault is output, as well as the variance based on the probability value as an uncertainty estimate; a probability threshold is preset, and when the output probability is greater than the threshold, the converter is judged to have a fault; otherwise, it is judged to be in normal operation; the uncertainty of the occurrence of faults is taken into account, which can more accurately reflect the actual situation and reduce the possibility of misjudgment;

[0267] In one case, suppose the model is The failure prediction probability is (in Indicates a fault condition. indicates a normal state), the pre-set probability threshold is ; Then the fault detection decision rule is:

[0268] Fault determination ;

[0269] Uncertainty estimation: During the inference phase of the model, a model with a dropout layer is applied (assuming forward propagation), each time obtaining a fault prediction probability ; then the mean of the predicted probabilities is:

[0270] ;

[0271] The variance of the predicted probabilities is used as an estimate of uncertainty:

[0272] ;

[0273] Analyze the differences in the features of the middle layer of the Bayesian deep learning model under fault and normal conditions, extract the feature patterns related to the fault, and determine the fault type or location where the fault occurs;

[0274] In one case, let the output feature of the intermediate layer of the model be ( is the feature dimension), for different fault types ), by clustering and extracting the intermediate layer features of known fault samples in the training set, the feature center corresponding to the fault type is obtained and the covariance matrix ;

[0275] In the diagnosis stage, for a sample to be diagnosed, the Mahalanobis distance between its middle layer feature and the feature center of each fault type is calculated:

[0276] ;

[0277] Then the sample is judged as the fault type Gj with the smallest distance, that is:

[0278] ;

[0279] As the aviation inverter continues to operate, new operating data will be continuously generated. In order to enable the fault detection model to adapt to the performance changes and operating condition evolution of the inverter, the Bayesian deep learning model is updated in real time, including:

[0280] Monte Carlo dropout is used to calculate the uncertainty index for new data, and the new data samples are sorted according to the uncertainty index, and the first several samples with the highest uncertainty are selected to form a selected sample set;

[0281] In one case, for a new data sample , using Monte Carlo dropout to calculate its uncertainty index ; The model forward propagation (the dropout layer randomly discards neuron connections with a certain probability during each forward propagation), and we get Prediction results ; Use the variance of the predictions as an indicator of uncertainty:

[0282] ;

[0283] in is the mean of the predicted results;

[0284] Sort the new data samples according to the uncertainty index and select the top The samples with the highest uncertainty form a selected sample set ;

[0285] Based on the selected sample set, the incremental learning method is adopted. When new data samples arrive, the model adjusts its own parameters by learning the new data based on the existing parameters to update the model.

[0286] In one case, let the parameters of the existing model be , the new data sample set, that is, the selected sample set is ,in Determine the new input spatiotemporal data, is the corresponding fault label;

[0287] Use gradient-based optimization methods (such as stochastic gradient descent SGD or its variants) to update model parameters; for the loss function , the update formula of model parameters is:

[0288] ;

[0289] in Fixed learning rate, Fixed parameters Gradient operator of the new data; After reaching the next generation, the final model parameters are:

[0290] ;

[0291] The generation process here can be done using mini-batch stochastic gradient descent (Mini-Batch SGD), which divides the new data into 3 mini-batches ( , each small batch of stars contains samples, in each iteration In this case, a random number is selected from these small groups of stars. , and calculate the average gradient on this small batch of stars to update the parameters:

[0292] ;

[0293] Through the above-mentioned model updating mechanism, the fault detection model can continuously adapt to new data and operating condition changes during the long-term operation of the aviation inverter, maintain a good performance state, improve the accuracy and reliability of fault detection, and provide strong guarantee for the safe and stable operation of the aircraft.

[0294] The above describes in detail the implementation modes of the present invention, but the present invention is not limited to the above implementation modes. For ordinary technicians in this technical field, after knowing the contents recorded in the present invention, they can make several equivalent changes and substitutions thereto without departing from the principle of the present invention. These equivalent changes and substitutions should also be regarded as belonging to the protection scope of the present invention.

Claims

1. An aviation inverter fault detection method based on information fusion algorithm, characterized in that: The steps include: Collect the operating data of aviation inverters and build data sets; Extract time information and space information from the data set, and fuse time features and space features based on variational autoencoders to obtain spatiotemporal data. The time information extraction includes: Skewness calculation based on the median: in, Express expectations, is the median, is the median absolute deviation; Kurtosis calculation based on interquartile range: in, is the second quartile, i.e. the median, are the first and third quartiles, respectively; DTW distance calculation: For two lengths respectively and Time Series and , construct a The distance matrix ,in, ; Define a cumulative distance matrix ,initialization , calculated by the following recursive formula: The DTW distance is ; Wavelet energy entropy calculation: For time series Perform wavelet transform to obtain wavelet coefficients ,in is the scale parameter, is the translation parameter; For each scale , calculate the wavelet energy at this scale : Calculate wavelet energy entropy , in is the total number of scales; The spatial information extraction is specifically as follows: Spatial feature construction based on physical structure: including heat conduction and temperature gradient features and electrical parameter difference features. The heat conduction and temperature gradient features are obtained by analyzing the temperature distribution inside the converter and the temperature difference between adjacent positions. The electrical parameter difference features are obtained by calculating the electrical parameter difference between adjacent components in the converter. Spatial information extraction based on graph theory: The converter graph model is constructed by taking the components or sensor nodes of the converter as vertices and the electrical connections, thermal conduction relations or physical adjacent relations between the components as edges; Calculate the Laplacian matrix of the graphical model and extract the eigenvalues ​​and eigenvectors of the Laplacian matrix as spatial features; and Aggregate and update the features of sensor nodes based on graph convolutional neural networks to extract representative spatial features; Build, train and optimize Bayesian deep learning models; The spatiotemporal data is input into the trained and optimized Bayesian deep learning model for processing, and the output is a numerical value representing the probability of failure, as well as the variance based on the probability value as uncertainty estimation; A probability threshold is preset, and when the output probability is greater than the threshold, it is determined that the converter is faulty; otherwise, it is determined to be in normal operation.

2. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 1, characterized in that: Preprocess the data in the dataset, including: Data cleaning: Assume that a variable time series is , calculate the median of the sequence : For each data point , calculate its absolute deviation , and then calculate the median absolute deviation ; like , then It is determined to be an outlier and corrected in the following way: in is a symbolic function; Data filtering includes strain step filtering processing and data normalization processing, wherein the data normalization processing adopts a normalization algorithm based on probability distribution.

3. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 2, characterized in that: The basic model of the Bayesian deep learning model is a hybrid architecture built by combining convolutional neural networks and long short-term memory networks; as well as In the convolutional neural network, a dilated convolution is used to expand the receptive field of the convolution kernel; In the long short-term memory network, the attention mechanism is introduced: Calculate attention weight: Assume is the hidden state sequence output by the long short-term memory network, and the attention weight vector is calculated through a fully connected layer and a softmax function : in, are the weights and biases of the fully connected layer, and the softmax function is used to normalize the attention scores; The hidden states are weighted summed according to the attention weights to obtain the context vector : ; Finally, the vector With the last hidden state of LSTM Splicing, as the final output feature of the model; Enable Bayesian deep learning models to dynamically assign weights based on the importance of input data; A regularization term is added to the Bayesian deep learning model.

4. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 3, characterized in that: The Bayesian deep learning model introduces the Bayesian theorem, regards the model parameters as random variables, and calculates the uncertainty estimate of the model parameters at each forward propagation calculation of the model output based on the Bayesian back-propagation algorithm according to the observed data. During the back-propagation process, the posterior distribution of the parameters is updated according to the uncertainty of the parameters and the gradient of the loss function.

5. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 3, characterized in that: The training of the Bayesian deep learning model includes: Dividing the data into a training set, a validation set, and a test set, and dividing the proportions of the training set, the validation set, and the test set in turn based on a random division method; Perform data enhancement operations on the data in the dataset, including adding noise, shifting the time series, and scaling the time series.

6. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 1, characterized in that: The optimization of the Bayesian deep learning model is specifically as follows: updating the model parameters according to the gradient information of the loss function to minimize the loss function so that the model gradually converges to a preset solution. The updated model parameters are: ; in, are model parameters, is the learning rate, , is the second-order moment estimate, The first moment estimate, is the decay rate hyperparameter, Is a positive number.

7. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 5, characterized in that: Analyze the differences in the features of the middle layer of the Bayesian deep learning model under fault and normal conditions, extract the characteristic patterns related to the fault, and determine the fault type or location where the fault occurs.

8. The method for detecting faults of an aviation converter based on an information fusion algorithm according to claim 1, characterized in that: Updates to Bayesian deep learning models in real time, including: Monte Carlo dropout is used to calculate the uncertainty index for new data, and the new data samples are sorted according to the uncertainty index, and the first several samples with the highest uncertainty are selected to form a selected sample set; Based on a selected sample set, an incremental learning approach is adopted. When new data samples arrive, the model adjusts its own parameters by learning the new data based on the existing parameters to update the model.

Citation Information

Patent Citations

  • Partial discharge fault state identification method based on ensemble learning

    CN111626153A

  • Behavior recognition feature selection method based on improved feature subset discrimination

    CN111709441A