An unmanned aerial vehicle intelligent fusion diagnosis and analysis method based on multi-modal fusion

By employing a multimodal fusion diagnostic analysis method, unified temporal alignment and quality perception fusion of multi-source UAV data were achieved, improving the accuracy and robustness of fault identification and health monitoring, and providing reliable fault warnings and operation and maintenance decision support for UAVs throughout the entire flight phase.

CN122112972APending Publication Date: 2026-05-29YUNNAN DIANENG SMART ENERGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN DIANENG SMART ENERGY CO LTD
Filing Date
2026-02-24
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing methods for monitoring the health of unmanned aerial vehicles (UAVs) and diagnosing faults are difficult to fully utilize the correlation between multi-source data. They are prone to misjudgment and omission, especially under complex operating conditions and multi-factor coupling. Furthermore, they lack online intelligent diagnostic processes and cross-mission knowledge transfer mechanisms throughout the entire flight phase.

Method used

A multimodal fusion diagnostic analysis method is adopted. Through multi-source data preprocessing, multimodal fusion diagnostic pre-module and improved PatchTST network technology, unified temporal alignment, modal quality perception and weighted fusion of multi-source heterogeneous data are achieved. A hierarchical health factor decoupling layer, an anomaly sensitive dual-stream layer and a cross-task prototype memory enhancement layer are constructed to perform temporal modeling and fusion diagnostic feature extraction.

Benefits of technology

It significantly improves the accuracy and robustness of health monitoring and fault diagnosis under all flight conditions of UAVs, and can provide reliable real-time fault warning and operation and maintenance decision support in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122112972A_ABST
    Figure CN122112972A_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle intelligent fusion diagnosis analysis methods based on multi-modal fusion, comprising: unmanned aerial vehicle multi-source data is collected, and multiple-source heterogeneous data is obtained by preprocessing;Multi-modal fusion diagnosis front-end module is constructed, and multiple-channel time series input tensor is obtained;Multiple-channel time series input is divided into time slice, and linear mapping and position coding are carried out;Improved PatchTST network is constructed, and fusion diagnosis feature is generated;Based on fusion diagnosis feature, fault and health index are obtained;Health index and operation and maintenance rule comparison generate diagnosis decision result;According to diagnosis decision result, update cross-task prototype memory library.The application realizes the intelligent fault diagnosis and operation and maintenance decision support of unmanned aerial vehicle full working condition by multi-modal fusion diagnosis front-end module and improved PatchTST network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) diagnostic technology, and in particular to an intelligent fusion diagnostic analysis method for UAVs based on multimodal fusion. Background Technology

[0002] With the widespread application of drones in logistics, environmental monitoring, power line inspection, and emergency rescue, the demand for drones to perform long-duration, all-weather missions in complex environments is increasing, placing higher demands on operational safety and health status monitoring. Current drone health monitoring and fault diagnosis largely rely on flight control logs, a limited number of key sensor signals, and human experience. They often employ threshold judgments, rule-based comparisons, or statistical analysis methods based on single sensor data. Some solutions introduce traditional machine learning models to identify faults in single vibration signals, motor current signals, or flight control parameters. Current methods typically analyze only a single type of signal or a limited number of features, making it difficult to fully utilize the correlations between multi-source data from flight control, power and energy sources, structure and vibration, environment, and loads. They lack sensitivity to faults caused by complex operating conditions and multi-factor coupling, and are prone to misjudgments and missed diagnoses in real-world flight scenarios with noise interference, sensor anomalies, and frequent operating condition switching. The diagnostic accuracy and robustness are insufficient to meet the requirements for highly reliable drone operation.

[0003] With the development of deep learning and time series modeling techniques, networks based on convolutional neural networks, recurrent neural networks, and self-attention mechanisms have been introduced into the field of UAV state recognition and fault diagnosis. However, existing methods mostly model single-modal time series or local features, lacking a multimodal fusion mechanism for multi-source heterogeneous data. They often simply concatenate multiple signals and input them into the network without addressing the data quality, temporal alignment, and diagnostic contribution of different modalities. Existing methods also focus on offline analysis or diagnostic model design for specific operating conditions, and have not yet formed an online intelligent diagnostic process covering the entire flight phase, including takeoff, climb, cruise, maneuvering, and landing. They lack a mechanism for cross-task knowledge transfer and memory enhancement using historical flight mission diagnostic results and feature patterns, making it difficult to provide stable and reliable technical support for real-time fault warning and operation and maintenance decisions during UAV operation.

[0004] Therefore, how to provide a method for intelligent fusion diagnostic analysis of UAVs based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] One objective of this invention is to propose an intelligent fusion diagnostic analysis method for unmanned aerial vehicles (UAVs) based on multimodal fusion. This invention comprehensively utilizes multi-source data preprocessing, a multimodal fusion diagnostic pre-module, and improved PatchTST network technology. It details the process of collecting multi-source heterogeneous data on flight control, power and energy, structure and vibration, environment, and loads during UAV flight. After unified temporal alignment, sample construction, modal quality perception, and adaptive weighted fusion, a multi-channel temporal input tensor is generated. Then, through the introduction of an improved PatchTST network with a hierarchical health factor decoupling layer, an anomaly-sensitive dual-flow layer, and a cross-task prototype memory enhancement layer, temporal modeling and fusion diagnostic feature extraction are performed. Finally, the entire process of fault type identification, fault severity assessment, health index calculation, diagnostic decision generation, and updating of the cross-task prototype memory bank is completed. This invention innovatively achieves quality-aware fusion of multi-source heterogeneous data, hierarchical decoupling of health factors, control factors, and environmental factors, dual-stream sensitive modeling of normal trends and abnormal patterns, and cross-task prototype memory-enhanced diagnosis in terms of method and network structure. It can significantly improve the accuracy, robustness, and online operation and maintenance decision-making capabilities of health monitoring and fault diagnosis under all flight conditions of UAVs.

[0006] According to an embodiment of the present invention, a method for intelligent fusion diagnostic analysis of unmanned aerial vehicles based on multimodal fusion includes: Multi-source data is collected during the flight of the drone and preprocessed to obtain multi-source heterogeneous data; A multimodal fusion diagnostic pre-module is constructed. The module is divided into stages by aligning sample units, and the multi-source heterogeneous data is time-aligned and resampled. The quality fusion unit calculates the quality index and obtains the multi-channel time-series input tensor. The multi-channel temporal input tensor is divided into time segments along the time dimension, which are used as patch sequences. Linear mapping and positional encoding are then performed to obtain the initial patch feature sequence. An improved PatchTST network is constructed, and a hierarchical health factor decoupling structure is introduced to map the initial patch feature sequence to multi-factor features. It is decomposed into normal flow features and abnormal flow features through an anomaly-sensitive dual-stream structure. A cross-task prototype memory enhancement structure is used to set up a cross-task prototype memory library, and similarity retrieval and feature fusion are performed to generate fused diagnostic features. The fused diagnostic features are classified and regressed to obtain the fault type, fault severity, and health index of the UAV. The fault type, fault severity, and health index are compared with preset thresholds and operation and maintenance rules to generate diagnostic decision results; The cross-mission prototype memory is updated based on the diagnostic decision results, providing historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions.

[0007] Optionally, the multi-source data includes flight control data, power and energy data, structural and vibration data, environmental data, and load data.

[0008] Optionally, obtaining multi-source heterogeneous data includes: During the flight of the drone, multi-source data is collected in real time through airborne sensors and flight control platform to form a multi-source raw data stream; A timestamp based on a unified time base is appended to the multi-source raw data streams, and the data is written into the data buffer in the order of collection to form a multi-source data sequence with time stamps; The time-stamped multi-source data sequences are sequentially subjected to denoising filtering, outlier detection and removal, coordinate transformation and dimensional normalization to obtain multi-source heterogeneous data for multimodal fusion diagnostic analysis.

[0009] Optionally, obtaining the multi-channel temporal input tensor includes: The flight start and end times are determined based on flight parameters such as altitude, speed, attitude, and thrust commands. The flight phase is divided into takeoff, climb, cruise, maneuver, and landing phases according to preset rules, resulting in multiple flight phase time intervals. For each flight phase time interval, the data is slidably extracted on the multi-source heterogeneous data according to the time window length and time step, and a time window sequence containing continuous time step data is constructed. Each time window is used as a time axis of a structured multimodal diagnostic sample, and the corresponding flight phase marker is recorded in the sample. Based on a unified time reference, a multimodal fusion diagnostic pre-module is constructed, consisting of an alignment sample unit and a quality fusion unit. The alignment sample unit performs interpolation and resampling processing on various types of sensor data in each structured multimodal diagnostic sample, so that various types of sensor data have the same time step interval and the same number of time steps on the time axis, forming a multi-channel time series data tensor on a unified time axis. The quality fusion unit calculates the missing rate, outlier count, signal energy, and out-of-bounds count of the time-series data tensor for each channel. Normalization is then used to convert each quality indicator into a modal quality score with a value range between zero and one, thus forming the corresponding quality score vector. The quality fusion unit determines the channel weighting coefficients based on the modal quality scores of each channel in the quality scoring vector. Weighted summation is performed on the multi-channel time series data in the channel dimension to generate weighted fused time series features. The weighted fused time series features are then combined with the multi-channel time series data to form a multi-channel time series input tensor.

[0010] Optionally, obtaining the initial patch feature sequence includes: Based on the time segment length and time step, the multi-channel time series input tensor is truncated along the time dimension using a sliding window method to obtain a patch sequence composed of multiple time segments, each time segment containing multi-channel time series data of multiple consecutive time steps; The multi-channel time-series data of each time segment in the patch sequence is unfolded into a one-dimensional vector according to the fixed order of time steps and channels. Each one-dimensional vector is linearly mapped to a preset feature dimension to generate a patch embedding vector sequence corresponding to each time segment. Assign a unique position index to each patch embedding vector in the patch embedding vector sequence, generate a corresponding position encoding vector based on the position index, and perform a vector addition operation between the position encoding vector and the patch embedding vector to obtain an initial patch feature sequence with time and position information.

[0011] Optionally, the generation of fusion diagnostic features includes: An improved PatchTST network was constructed, which integrates a hierarchical health factor decoupling layer, an anomaly-sensitive dual-flow layer, and a cross-task prototype memory enhancement layer into the same temporal coding framework. The initial patch feature sequence is used as input and the fused diagnostic features are used as output. The hierarchical health factor decoupling layer inputs the initial patch feature sequence into the encoding unit, and generates control factor subspace features, environmental factor subspace features and health factor subspace features for each patch feature through three sets of linear mappings. The abnormality-sensitive dual-flow layer calculates the abnormality tendency score of each patch based on the health factor branch features. According to the abnormality tendency score and the gating function, the patch features are decomposed into normal flow features and abnormal flow features. The normal flow features are input into the normal flow encoder branch, and the abnormal flow features are input into the abnormal flow encoder branch. Self-attention operation and feedforward operation are performed respectively to obtain normal flow coding features and abnormal flow coding features. The anomaly-sensitive dual-stream layer inputs normal stream coding features and anomaly stream coding features into the cross-attention unit to perform cross-attention operations, and concatenates the normal stream coding features and anomaly stream coding features to generate anomaly-sensitive dual-stream fusion features; The cross-task prototype memory enhancement layer constructs a cross-task prototype memory bank based on health factor branch features and abnormality-sensitive dual-stream fusion features. The abnormality-sensitive dual-stream fusion features of the current sample are used as query features. Prototype features similar to the query features are retrieved from the cross-task prototype memory bank. A weighted fusion operation is performed on the query features and the retrieved prototype features, and the weighted fusion result is used as the fusion diagnostic feature.

[0012] Optionally, obtaining the drone's fault type, fault severity, and health index includes: Fault type identification is performed on the fused diagnostic features. Multiple output channels are set up, each corresponding to a fault type or normal state. Multi-class classification calculation is performed on the fused diagnostic features to obtain the probability value of each fault type. The fault type of each UAV is determined according to the category corresponding to the highest probability. The severity of the fault is assessed by fusing diagnostic features and regression calculation is performed on the fusing diagnostic features to obtain a fault severity score ranging from zero to one. Zero is taken as the fault-free state and one is taken as the state of maximum fault severity. The fault severity score is taken as the fault severity of each UAV. The system integrates diagnostic features and fault severity scores as input, performs weighted calculations according to weighting rules, and generates a health index with a value range of zero to one. Zero is used as a complete failure state, and one is used as a complete health state. The system outputs the health index corresponding to each UAV.

[0013] Optionally, generating diagnostic decision results includes: Read the fault type, fault severity and health index corresponding to each monitored object, retrieve the health index threshold, fault severity threshold and operation and maintenance rule parameters corresponding to each monitored object from the storage medium, and set the health index threshold and fault severity threshold for each monitored object. For each monitored object, the health index is compared with the corresponding health index threshold, and the fault severity is compared with the corresponding fault severity threshold. Based on the comparison results, the operating status of the monitored object is divided into normal state, warning state and dangerous state, and the diagnostic level of the monitored object is determined in combination with the fault type. Logical judgments are made on the overall operating status based on the diagnostic level and operation and maintenance rule parameters of each monitored object. When any monitored object is in a dangerous state, a return command or emergency shutdown command is generated. When there is a monitored object in a warning state but not in a dangerous state, a load adjustment command and maintenance suggestion are generated. The generated commands and maintenance suggestions are combined to form a diagnostic decision result.

[0014] Optionally, updating the cross-mission prototype memory based on diagnostic decision results to provide historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions includes: Obtain the fusion diagnostic features, fault type, fault severity, health index and diagnostic decision results corresponding to the current flight mission, associate the flight mission identifier, flight time, flight phase information with fault type, fault severity and health index and store them in the health record, and update the corresponding prototype features or add new prototype features in the cross-mission prototype memory using fusion diagnostic features, fault type, health index range and flight phase as conditions. In the cross-task prototype memory, when the number of prototype features exceeds the capacity limit, the prototype features are filtered according to the generation time order of the prototype features, and the prototype feature record at the bottom is deleted to keep the number of prototype features in the cross-task prototype memory within the capacity range.

[0015] The beneficial effects of this invention are: This invention introduces a multimodal fusion diagnostic pre-module, achieving unified temporal alignment, phased sample construction, and modal quality perception weighted fusion of multi-source heterogeneous data from flight control, power and energy, structure and vibration, environment, and loads. This significantly improves the utilization efficiency of UAV operational status information. Compared with existing diagnostic methods based on single signals or simple feature splicing, this invention can comprehensively utilize data from different sources and of different quality levels within a unified time axis and feature space. It adaptively reduces the weight of abnormal data channels, effectively suppressing interference from sensor noise, missing data, and local anomalies. This improves the accuracy and robustness of fault identification, fault severity assessment, and health index calculation, providing a more reliable data foundation for UAV health monitoring under complex operating conditions.

[0016] This invention constructs an improved PatchTST network, integrating a hierarchical health factor decoupling layer, an anomaly-sensitive dual-flow layer, and a cross-mission prototype memory enhancement layer within the same temporal coding framework. This enables hierarchical representation of control factors, environmental factors, and health factors; dual-flow sensitive modeling of normal trends and anomaly patterns; and the memorization and reuse of historical diagnostic knowledge from multiple flight missions. This not only improves the granularity of fault mode modeling across all flight phases and the ability to detect weak and hidden faults, but also allows diagnostic results to directly drive operational decisions such as return-to-base, operational condition limitations, payload adjustments, and maintenance recommendations. Furthermore, it continuously updates the cross-mission prototype memory and health records, overcoming the technical bottlenecks of insufficient utilization of diagnostic results and difficulty in supporting online intelligent operation and maintenance and closed-loop management in existing technologies. This provides effective support for the safe and reliable operation and intelligent maintenance of unmanned aerial vehicles (UAVs). Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows a method for intelligent fusion diagnosis and analysis of unmanned aerial vehicles based on multimodal fusion proposed in this invention. Figure 2 This is a structural block diagram of the multimodal fusion diagnostic pre-module of an intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion proposed in this invention; Figure 3This is a functional diagram of the improved PatchTST network, which is based on a multimodal fusion-based intelligent fusion diagnostic analysis method for unmanned aerial vehicles (UAVs) proposed in this invention. Detailed Implementation

[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0019] refer to Figure 1 , Figure 2 and Figure 3 A method for intelligent fusion diagnostic analysis of unmanned aerial vehicles based on multimodal fusion, comprising: Multi-source data is collected during the flight of the drone and preprocessed to obtain multi-source heterogeneous data; A multimodal fusion diagnostic pre-module is constructed. The module is divided into stages by aligning sample units, and the multi-source heterogeneous data is time-aligned and resampled. The quality fusion unit calculates the quality index and obtains the multi-channel time-series input tensor. The multi-channel temporal input tensor is divided into time segments along the time dimension, which are used as patch sequences. Linear mapping and positional encoding are then performed to obtain the initial patch feature sequence. An improved PatchTST network is constructed, and a hierarchical health factor decoupling structure is introduced to map the initial patch feature sequence to multi-factor features. It is decomposed into normal flow features and abnormal flow features through an anomaly-sensitive dual-stream structure. A cross-task prototype memory enhancement structure is used to set up a cross-task prototype memory library, and similarity retrieval and feature fusion are performed to generate fused diagnostic features. The fused diagnostic features are classified and regressed to obtain the fault type, fault severity, and health index of the UAV. The fault type, fault severity, and health index are compared with preset thresholds and operation and maintenance rules to generate diagnostic decision results; The cross-mission prototype memory is updated based on the diagnostic decision results, providing historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions.

[0020] In this embodiment, the multi-source data includes flight control data, power and energy data, structural and vibration data, environmental data, and load data.

[0021] In this embodiment, obtaining multi-source heterogeneous data includes: During the flight of the drone, multi-source data is collected in real time through airborne sensors and flight control platform to form a multi-source raw data stream; A timestamp based on a unified time base is appended to the multi-source raw data streams, and the data is written into the data buffer in the order of collection to form a multi-source data sequence with time stamps; The time-stamped multi-source data sequences are sequentially subjected to denoising filtering, outlier detection and removal, coordinate transformation and dimensional normalization to obtain multi-source heterogeneous data for multimodal fusion diagnostic analysis.

[0022] In this embodiment, obtaining the multi-channel temporal input tensor includes: The start and end times of flight are determined based on flight parameters such as altitude, speed, attitude, and thrust commands. The flight is then divided into takeoff, climb, cruise, maneuver, and landing phases according to preset rules, resulting in multiple time intervals for each flight phase. These preset rules are specifically as follows: The preset rules include the rules for determining the start and end of flight and the rules for dividing each flight phase. When the thrust command is continuously greater than 30% of the rated thrust, and the height difference between the UAV and the takeoff point is continuously greater than 5 meters, the vertical speed is continuously greater than 0.5 meters per second, and the duration is not less than 2 seconds, the current moment is determined as the start of flight. When the altitude drops to less than 2 meters from the takeoff point, and the horizontal speed is less than 1 meter per second, the absolute value of the vertical speed is less than 0.2 meters per second, and the duration is not less than 3 seconds, the current moment is determined as the end of flight. Between the start and end of the flight, the flight process is divided into different phases according to changes in altitude and attitude. The takeoff phase corresponds to an altitude increase from 0 meters to 20 meters, the climb phase corresponds to an altitude increase from 20 meters to 80 meters, when the altitude is stable within 5 meters above or below 80 meters, the absolute value of the vertical velocity is less than 0.3 meters per second, and the horizontal velocity is maintained between 8 meters per second and 15 meters per second, the current time interval is divided into the cruise phase, and when the altitude continuously decreases from around 80 meters to below 20 meters, it is divided into the landing phase. Within each phase, when the absolute value of the roll angle or pitch angle is greater than 15 degrees, or when the turning acceleration calculated based on the horizontal speed and turning radius is greater than one-third of the gravitational acceleration, the time interval that meets the condition is marked as the maneuver phase. For each flight phase time interval, time windows are slidably extracted from multi-source heterogeneous data according to the time window length and time step size to construct a time window sequence containing continuous time step data. Each time window is used as a time axis for a structured multimodal diagnostic sample, and the corresponding flight phase marker is recorded in the sample. The construction of the time window sequence containing continuous time step data is specifically as follows: After obtaining the time interval corresponding to the flight phase, the time interval is divided into discrete time steps arranged in chronological order according to a uniform sampling period. The time window length and time step size are set. Starting from the initial time step of the time interval, several consecutive time steps are extracted as the first time window. Then, the initial time step is shifted backward by the time step size, and new consecutive time steps are extracted as the next time window. This process is repeated until the end of the current time interval is reached, thus constructing a time window sequence composed of multiple time windows. Based on a unified time reference, a multimodal fusion diagnostic pre-module is constructed, consisting of an alignment sample unit and a quality fusion unit. The alignment sample unit performs interpolation and resampling processing on various types of sensor data in each structured multimodal diagnostic sample, so that various types of sensor data have the same time step interval and the same number of time steps on the time axis, forming a multi-channel time series data tensor on a unified time axis. The quality fusion unit calculates quality metrics for the time-series data tensor of each channel, including missing rate, outlier count, signal energy, and out-of-bounds frequency. Normalization is then applied to convert each quality metric into a modal quality score with values ​​ranging from zero to one, forming a corresponding quality score vector. Specifically, the calculation of the missing rate, outlier count, signal energy, and out-of-bounds frequency quality metrics involves: In the time series data of a channel, the number of missing data points in the current time window is counted. The ratio of the number of missing data points to the total number of sampling points in the channel within the time window is used as the missing rate of the current channel. All data points that do not meet the normal range conditions are detected based on the mean and standard deviation of historical time series data. The number of data points is used as the number of outliers in the current time window of the channel. In the time series data of a channel, the amplitude of each sampling point within the current time window is summed by squares, and then averaged according to the number of sampling points. The resulting average sum of squares is taken as the signal energy of the channel within the current time window. The number of times the statistical data curve enters the upper or lower limit range from the normal range within the current time window is taken as the channel's out-of-limit index. The quality fusion unit determines the channel weighting coefficients based on the modal quality scores of each channel in the quality score vector. A weighted summation operation is performed on the multi-channel time series data along the channel dimension to generate weighted fused time series features. These weighted fused time series features are then combined with the multi-channel time series data to form a multi-channel time series input tensor. Specifically, determining the channel weighting coefficients based on the modal quality scores of each channel in the quality score vector involves: The quality score vector within the current time window is statistically analyzed, and the modal quality scores of all channels are summed to obtain the total quality score. When the total quality score is greater than zero, the modal quality score of each channel is divided by the total quality score, and the result is used as the channel weighting coefficient. The sum of the weighting coefficients of all channels is made equal to one. When the total quality score is equal to zero, the weights are evenly distributed according to the number of channels, and the weighting coefficient of each channel is set to one and divided by the total number of channels. Even when all quality scores are zero, a set of channel weighting coefficients can still be obtained.

[0023] In this embodiment, obtaining the initial patch feature sequence includes: Based on the time segment length and time step, the multi-channel time series input tensor is truncated along the time dimension using a sliding window method to obtain a patch sequence composed of multiple time segments, each time segment containing multi-channel time series data of multiple consecutive time steps; The multi-channel time-series data of each time segment in the patch sequence is expanded into a one-dimensional vector according to a fixed order of time steps and channels. Each one-dimensional vector is linearly mapped to a preset feature dimension to generate a patch embedding vector sequence corresponding to each time segment. The preset feature dimension is set to 128. Assign a unique position index to each patch embedding vector in the patch embedding vector sequence, generate a corresponding position encoding vector based on the position index, and perform a vector addition operation between the position encoding vector and the patch embedding vector to obtain an initial patch feature sequence with temporal position information. Specifically, generating the corresponding position encoding vector based on the position index involves: Each possible time segment within the time window is sequentially numbered, starting from the beginning and increasing sequentially, serving as a position index. Based on the maximum number of time segments and the preset feature dimension, a position encoding matrix is ​​initialized. The number of rows in the position encoding matrix equals the maximum number of time segments, and the number of columns equals the preset feature dimension. Each row in the matrix corresponds to an encoding vector for a position index. During forward computation, for each patch embedding vector in the patch embedding vector sequence, the row in the position encoding matrix with the same number as the number in the sequence is searched, and the row vector is used as the current position encoding vector of the patch.

[0024] In this embodiment, the generation of fusion diagnostic features includes: An improved PatchTST network was constructed, which integrates a hierarchical health factor decoupling layer, an anomaly-sensitive dual-flow layer, and a cross-task prototype memory enhancement layer into the same temporal coding framework. The initial patch feature sequence is used as input and the fused diagnostic features are used as output. The hierarchical health factor decoupling layer inputs the initial patch feature sequence into the encoding unit. For each patch feature, it generates control factor subspace features, environmental factor subspace features, and health factor subspace features through three sets of linear mappings. The three sets of linear mappings are as follows: The first set of linear mappings uses preset control factor weight parameters and bias parameters to perform a matrix multiplication and addition operation on the input vector, and outputs a one-dimensional vector with a preset control factor dimension as the control factor subspace feature. The second set of linear mappings uses preset environmental factor weight parameters and bias parameters to perform a matrix multiplication and addition operation on the same input vector, and outputs a one-dimensional vector with a preset environmental factor dimension as the environmental factor subspace feature. The third set of linear mappings uses preset health factor weight parameters and bias parameters to perform a matrix multiplication and addition operation on the same input vector, and outputs a one-dimensional vector with a preset health factor dimension as the health factor subspace feature. The preset control factor weight parameters are set to a real matrix with a size equal to the preset feature dimension multiplied by the control factor dimension, the environmental factor weight parameters are set to a real matrix with a size equal to the preset feature dimension multiplied by the environmental factor dimension, and the health factor weight parameters are set to a real matrix with a size equal to the preset feature dimension multiplied by the health factor dimension. The initial values ​​of the three are generated according to a uniform random distribution in the interval from negative zero point 0.1 to positive zero point 0.1. The anomaly-sensitive dual-flow layer calculates the anomaly tendency score for each patch based on health factor branch features. According to the anomaly tendency score and a gating function, the patch features are decomposed into normal flow features and anomaly flow features. The normal flow features are input into the normal flow encoder branch, and the anomaly flow features are input into the anomaly flow encoder branch. Self-attention and feedforward operations are performed respectively to obtain the normal flow coding features and the anomaly flow coding features, where: The calculation of the abnormal tendency score for each patch is as follows: In the health factor branch features corresponding to the current set of time segments, the statistical mean and statistical standard deviation of the health factor branch features are calculated according to the feature dimensions to obtain the health feature benchmark of the current set of time segments. For each time segment, the difference between the health factor branch features and the statistical mean is calculated on each feature dimension. The square of the difference is summed on all feature dimensions and then divided by the square root of the number of feature dimensions to obtain a non-negative scalar reflecting the degree of deviation of the time segment from the normal health level. This scalar is then normalized to obtain the abnormal tendency score of the current time segment. The step of decomposing patch features into normal flow features and abnormal flow features based on abnormal tendency scores and gating functions is as follows: For each time segment, the abnormal tendency score is expanded along the feature dimension to form a gating coefficient vector with the same length as the original time segment features. The gating coefficient vector is multiplied element-wise with the time segment features along each feature dimension to obtain the abnormal flow features. At the same time, the abnormal tendency score is subtracted by one to obtain the normal flow gating coefficient. The normal flow gating coefficient is expanded along the feature dimension and multiplied element-wise with the time segment features to obtain the normal flow features. This completes the decomposition of normal flow features and abnormal flow features for each time segment. The anomaly-sensitive dual-stream layer inputs normal stream coding features and anomaly stream coding features into a cross-attention unit to perform cross-attention operations, and concatenates the normal stream coding features and anomaly stream coding features to generate anomaly-sensitive dual-stream fusion features. Specifically, the cross-attention operation involves: Using normal stream coding features as queries and abnormal stream coding features as comparison features, we first obtain representations for similarity calculation through linear transformation. For each time position of normal stream coding features, we calculate the Euclidean distance similarity between them and abnormal stream coding features at all time positions. We then perform exponential operations and normalization on the similarity to obtain a set of weight coefficients. We then use the weight coefficients to perform a weighted summation on the abnormal stream coding features to obtain normal stream update features that incorporate abnormal information. In the same way, we use abnormal stream coding features as queries and normal stream coding features as comparison features to calculate and obtain abnormal stream update features that incorporate normal information. The cross-task prototype memory enhancement layer constructs a cross-task prototype memory library based on health factor branch features and anomaly-sensitive dual-stream fusion features. Using the anomaly-sensitive dual-stream fusion features of the current sample as the query feature, it retrieves prototype features similar to the query feature from the cross-task prototype memory library. A weighted fusion operation is performed on the query feature and the retrieved prototype features, and the weighted fusion result is used as the fusion diagnostic feature. The construction of the cross-task prototype memory library specifically involves: Health factor branch features and anomaly-sensitive dual-stream fusion features corresponding to each time window are extracted from multiple historical flight missions. The features are grouped according to flight stage, fault type and health index interval. The features within each group are averaged one dimension at a time to obtain the prototype feature vector of the current group. The prototype feature vector and the corresponding flight stage, fault type and health index interval are written as a prototype record into the cross-mission prototype memory.

[0025] In this embodiment, obtaining the drone's fault type, fault severity, and health index includes: Fault type identification is performed on the fused diagnostic features. Multiple output channels are set, each corresponding to a fault type or normal state. Multi-class classification calculation is performed on the fused diagnostic features to obtain the probability value of each fault type. The fault type of each UAV is determined according to the category corresponding to the highest probability. The multi-class classification calculation of the fused diagnostic features specifically includes: Several fault categories and normal states are defined. The fused diagnostic features are weighted and summed with the corresponding classification weight parameters and a classification bias parameter to obtain a real score for each fault category and normal state. All scores are then exponentially calculated, and the exponential value of each score is divided by the sum of all score exponential values. The result is used as the probability value of the corresponding fault category or normal state. The severity of the fault is assessed by fusing diagnostic features and regression calculation is performed on the fusing diagnostic features to obtain a fault severity score ranging from zero to one. Zero is taken as the fault-free state and one is taken as the state of maximum fault severity. The fault severity score is taken as the fault severity of each UAV. Using fused diagnostic features and fault severity scores as input, a weighted calculation is performed according to a weighting rule to generate a health index ranging from zero to one. Zero is used as a complete failure state, and one is used as a complete health state. The health index corresponding to each UAV is output. The weighting rule is as follows: The first health score is calculated based on the mapping relationship of the fusion diagnostic features. A real weight is assigned to each feature component in the fusion diagnostic features. The weights are multiplied one by one by the feature components, the results are summed and a real bias is added. The real result is input into the Sigmoid function, which outputs the first health score with a value of zero to one. The fault severity score with a value of zero to one is read. The fault severity score is subtracted from one to obtain the second health score. The first health score and the second health score are weighted and calculated by multiplying the first health score by 0.6 and the second health score by 0.4 and adding them together to obtain the health index with a value range of zero to one.

[0026] In this embodiment, generating diagnostic decision results includes: The system reads the fault type, fault severity, and health index corresponding to each monitored object. It then retrieves the health index threshold, fault severity threshold, and operation and maintenance rule parameters corresponding to each monitored object from the storage medium. Finally, it sets the health index threshold and fault severity threshold for each monitored object. Specifically, the health index threshold, fault severity threshold, and operation and maintenance rule parameters are as follows: The health index warning threshold for each monitored object is set at 0.7, and the danger threshold is set at 0.4. A health index greater than or equal to 0.7 is considered normal, a health index less than 0.7 but greater than or equal to 0.4 is considered a warning, and a health index less than 0.4 is considered dangerous. The fault severity warning threshold is set at 0.3, and the danger threshold is set at 0.6. A fault severity score greater than or equal to 0.3 but less than 0.6 is considered a warning, and a score greater than or equal to 0.6 is considered dangerous. The operation and maintenance rule parameters are configured according to a combination of health index status and fault severity status. When any indicator is in a warning state, a maintenance suggestion or load adjustment command is triggered. When any indicator is in a dangerous state, a return command or emergency shutdown command is triggered. For each monitored object, the health index is compared with the corresponding health index threshold, and the fault severity is compared with the corresponding fault severity threshold. Based on the comparison results, the operating status of the monitored object is divided into normal state, warning state and dangerous state, and the diagnostic level of the monitored object is determined in combination with the fault type. Logical judgments are made on the overall operating status based on the diagnostic level and operation and maintenance rule parameters of each monitored object. When any monitored object is in a dangerous state, a return command or emergency shutdown command is generated. When there is a monitored object in a warning state but not in a dangerous state, a load adjustment command and maintenance suggestion are generated. The generated commands and maintenance suggestions are combined to form a diagnostic decision result.

[0027] In this embodiment, updating the cross-mission prototype memory based on diagnostic decision results to provide historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions includes: Obtain the fusion diagnostic features, fault type, fault severity, health index and diagnostic decision results corresponding to the current flight mission, associate the flight mission identifier, flight time, flight phase information with fault type, fault severity and health index and store them in the health record, and update the corresponding prototype features or add new prototype features in the cross-mission prototype memory using fusion diagnostic features, fault type, health index range and flight phase as conditions. In the cross-task prototype memory, when the number of prototype features exceeds the capacity limit, the prototype features are filtered according to the generation time order of the prototype features, and the prototype feature record at the bottom is deleted to keep the number of prototype features in the cross-task prototype memory within the capacity range.

[0028] Example 1: To verify the feasibility of this invention in practice, it was applied to a power company's UAV inspection project of transmission lines from April to June 2025. The inspection area included mixed terrain of mountains, hills, and suburbs, with typical flight altitudes ranging from 20 to 100 meters and single mission durations from 20 to 40 minutes. The UAV's onboard sensors included a flight control system, current, voltage, and speed sensors for four brushless motors, battery pack voltage, current, and temperature sensors, a three-axis accelerometer for the arms and fuselage, ambient temperature, humidity, wind speed, and direction sensors, and a dual-light camera (visible and infrared), forming a multi-source heterogeneous data channel. During the trial period, ten multi-rotor UAVs were deployed, performing over 500 inspection missions with a total flight time exceeding 300 hours. Various typical faults and abnormal operating conditions were observed on-site, including motor bearing wear, minor propeller cracks, battery capacity degradation, and strong crosswind interference. Traditional manual playback and threshold alarm methods suffered from low utilization of multi-source data, delayed fault detection, and high false alarm rates under complex conditions.

[0029] In the current scenario, the UAV's onboard computing unit and the ground station server jointly run the method of this invention to uniformly collect and preprocess multi-source data on flight control, power and energy, structure and vibration, environment, and load. The multi-modal fusion diagnostic pre-module completes flight phase division, time alignment, and resampling, encapsulating sensor data of different sampling rates and types into structured multi-modal diagnostic samples. Based on modal quality scores, channels with high noise and many missing data are automatically downweighted to generate multi-channel time-series inputs. Subsequently, the improved PatchTST network models each time window, distinguishes control, environmental, and health factors through a hierarchical health factor decoupling layer, and simultaneously represents long-term state changes and instantaneous anomalies through an anomaly-sensitive dual-flow layer. Combined with a cross-task prototype memory enhancement layer, it calls typical faults and health modes from historical flight missions to perform fusion diagnostic analysis on the current flight sample, outputting fault type, fault severity, and health index in real time. When the health index enters the warning or danger zone, it pushes diagnostic decision results such as return to base, load limitation, or maintenance suggestions to the ground operation and maintenance platform.

[0030] During a two-month trial run, the method of this invention achieved stable detection of multiple recorded events of early motor bearing failures, battery performance degradation, and propeller cracks in inspection scenarios. Compared with methods that use only a single vibration signal and a fixed threshold, the accuracy of fault identification was significantly improved, the number of false alarms was significantly reduced, and most structural faults were warned before the UAV showed obvious attitude abnormalities and a sharp increase in vibration. Under conditions of rapid wind speed changes and high temperatures, the method of this invention can still maintain a stable health index curve, and the alarm behavior is highly consistent with the actual maintenance conclusions. It solves the problem that traditional solutions are difficult to comprehensively utilize multi-source heterogeneous data throughout the entire flight phase and are difficult to support online intelligent diagnosis and operation and maintenance decisions, providing reliable technical support for the safety and operation and maintenance efficiency of UAV inspection of power transmission lines.

[0031] Table 1. Performance Comparison of Different UAV Fault Diagnosis Methods in Typical Scenarios

[0032] As shown in Table 1, the fault identification accuracy of the method of this invention reaches 96.8%, significantly higher than the 82.3% of the traditional threshold method and the 86.5%, 88.1%, and 89.7% of the single-modal RNN, single-modal CNN, and single-modal Transformer methods, respectively. It is also significantly improved compared to the 91.0% of the multimodal simple concatenation method. The false positive rate of the method of this invention is only 3.6%, while that of the traditional threshold method is 11.4% and that of the multimodal simple concatenation method is 7.1%. In terms of the false negative rate, the method of this invention is 4.1%, which is significantly lower than that of the traditional threshold method (18.9%) and the multimodal simple concatenation method (10.3%), directly demonstrating the comprehensive advantages of being more accurate, less wrong, and less missed.

[0033] In terms of diagnostic timeliness and adaptability to all operating conditions, the average diagnostic latency of the method of this invention is 4.2 seconds, shorter than the 9.5 seconds of the traditional threshold method and the 6.1-7.8 seconds range of single-modal and simple multimodal methods, significantly improving online response speed while ensuring high accuracy. Regarding full operating condition coverage, the method of this invention achieves 93.5%, while the traditional threshold method is only 65.0%, and the single-modal RNN, single-modal CNN, and single-modal Transformer are 71.4%, 74.2%, and 76.8% respectively, and the multimodal simple stitching method is 81.3%, indicating that the present invention can maintain stable diagnostic output in different phases of takeoff, climb, cruise, maneuvering, and landing.

[0034] In terms of early fault detection capability and system robustness, the method of this invention achieves an early detection rate of 81.4%, a significant improvement compared to the traditional threshold method's 48.6% and the multimodal simple splicing method's 66.7%. It is more sensitive to weak faults such as premature motor bearing damage, battery degradation, and structural microcracks. Regarding robustness to missing data, the method of this invention scores 9.2, higher than the traditional threshold method's 6.1 and the multimodal simple splicing method's 8.0. In terms of stable operating time, the method of this invention can run stably for 228 hours in continuous testing, while the traditional threshold method is 120 hours, the single-modal Transformer is 183 hours, and the multimodal simple splicing method is 196 hours. This demonstrates that in long-term, multi-tasking engineering operating environments, this invention has significant advantages in resisting data missing interference and overall stability.

[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multi-modal fusion-based unmanned aerial vehicle intelligent fusion diagnosis analysis method, characterized in that, include: Multi-source data is collected during the flight of the drone and preprocessed to obtain multi-source heterogeneous data; A multimodal fusion diagnostic pre-module is constructed. The module is divided into stages by aligning sample units, and the multi-source heterogeneous data is time-aligned and resampled. The quality fusion unit calculates the quality index and obtains the multi-channel time-series input tensor. The multi-channel temporal input tensor is divided into time segments along the time dimension, which are used as patch sequences. Linear mapping and positional encoding are then performed to obtain the initial patch feature sequence. An improved PatchTST network is constructed, and a hierarchical health factor decoupling structure is introduced to map the initial patch feature sequence to multi-factor features. It is decomposed into normal flow features and abnormal flow features through an anomaly-sensitive dual-stream structure. A cross-task prototype memory enhancement structure is used to set up a cross-task prototype memory library, and similarity retrieval and feature fusion are performed to generate fused diagnostic features. The fused diagnostic features are classified and regressed to obtain the fault type, fault severity, and health index of the UAV. The fault type, fault severity, and health index are compared with preset thresholds and operation and maintenance rules to generate diagnostic decision results; The cross-mission prototype memory is updated based on the diagnostic decision results, providing historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions. 2.The unmanned aerial vehicle intelligent fusion diagnosis analysis method based on multi-modal fusion of claim 1, characterized in that, The multi-source data includes flight control data, power and energy data, structural and vibration data, environmental data, and load data.

3. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The obtained multi-source heterogeneous data includes: During the flight of the drone, multi-source data is collected in real time through airborne sensors and flight control platform to form a multi-source raw data stream; A timestamp based on a unified time base is appended to the multi-source raw data streams, and the data is written into the data buffer in the order of collection to form a multi-source data sequence with time stamps; The time-stamped multi-source data sequences are sequentially subjected to denoising filtering, outlier detection and removal, coordinate transformation and dimensional normalization to obtain multi-source heterogeneous data for multimodal fusion diagnostic analysis.

4. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The process of obtaining the multi-channel temporal input tensor includes: The flight start and end times are determined based on flight parameters such as altitude, speed, attitude, and thrust commands. The flight phase is divided into takeoff, climb, cruise, maneuver, and landing phases according to preset rules, resulting in multiple flight phase time intervals. For each flight phase time interval, the data is slidably extracted on the multi-source heterogeneous data according to the time window length and time step, and a time window sequence containing continuous time step data is constructed. Each time window is used as a time axis of a structured multimodal diagnostic sample, and the corresponding flight phase marker is recorded in the sample. Based on a unified time reference, a multimodal fusion diagnostic pre-module is constructed, consisting of an alignment sample unit and a quality fusion unit. The alignment sample unit performs interpolation and resampling processing on various types of sensor data in each structured multimodal diagnostic sample, so that various types of sensor data have the same time step interval and the same number of time steps on the time axis, forming a multi-channel time series data tensor on a unified time axis. The quality fusion unit calculates the missing rate, outlier count, signal energy, and out-of-bounds count of the time-series data tensor for each channel. Normalization is then used to convert each quality indicator into a modal quality score with a value range between zero and one, thus forming the corresponding quality score vector. The quality fusion unit determines the channel weighting coefficients based on the modal quality scores of each channel in the quality scoring vector. Weighted summation is performed on the multi-channel time series data in the channel dimension to generate weighted fused time series features. The weighted fused time series features are then combined with the multi-channel time series data to form a multi-channel time series input tensor.

5. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The process of obtaining the initial patch feature sequence includes: Based on the time segment length and time step, the multi-channel time series input tensor is truncated along the time dimension using a sliding window method to obtain a patch sequence composed of multiple time segments, each time segment containing multi-channel time series data of multiple consecutive time steps; The multi-channel time-series data of each time segment in the patch sequence is unfolded into a one-dimensional vector according to the fixed order of time steps and channels. Each one-dimensional vector is linearly mapped to a preset feature dimension to generate a patch embedding vector sequence corresponding to each time segment. Assign a unique position index to each patch embedding vector in the patch embedding vector sequence, generate a corresponding position encoding vector based on the position index, and perform a vector addition operation between the position encoding vector and the patch embedding vector to obtain an initial patch feature sequence with time and position information.

6. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The generated fusion diagnostic features include: An improved PatchTST network was constructed, which integrates a hierarchical health factor decoupling layer, an anomaly-sensitive dual-flow layer, and a cross-task prototype memory enhancement layer into the same temporal coding framework. The initial patch feature sequence is used as input and the fused diagnostic features are used as output. The hierarchical health factor decoupling layer inputs the initial patch feature sequence into the encoding unit, and generates control factor subspace features, environmental factor subspace features and health factor subspace features for each patch feature through three sets of linear mappings. The abnormality-sensitive dual-flow layer calculates the abnormality tendency score of each patch based on the health factor branch features. According to the abnormality tendency score and the gating function, the patch features are decomposed into normal flow features and abnormal flow features. The normal flow features are input into the normal flow encoder branch, and the abnormal flow features are input into the abnormal flow encoder branch. Self-attention operation and feedforward operation are performed respectively to obtain normal flow coding features and abnormal flow coding features. The anomaly-sensitive dual-stream layer inputs normal stream coding features and anomaly stream coding features into the cross-attention unit to perform cross-attention operations, and concatenates the normal stream coding features and anomaly stream coding features to generate anomaly-sensitive dual-stream fusion features; The cross-task prototype memory enhancement layer constructs a cross-task prototype memory bank based on health factor branch features and abnormality-sensitive dual-stream fusion features. The abnormality-sensitive dual-stream fusion features of the current sample are used as query features. Prototype features similar to the query features are retrieved from the cross-task prototype memory bank. A weighted fusion operation is performed on the query features and the retrieved prototype features, and the weighted fusion result is used as the fusion diagnostic feature.

7. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The obtained fault type, fault severity, and health index of the drone include: Fault type identification is performed on the fused diagnostic features. Multiple output channels are set up, each corresponding to a fault type or normal state. Multi-class classification calculation is performed on the fused diagnostic features to obtain the probability value of each fault type. The fault type of each UAV is determined according to the category corresponding to the highest probability. The severity of the fault is assessed by fusing diagnostic features and regression calculation is performed on the fusing diagnostic features to obtain a fault severity score ranging from zero to one. Zero is taken as the fault-free state and one is taken as the state of maximum fault severity. The fault severity score is taken as the fault severity of each UAV. The system integrates diagnostic features and fault severity scores as input, performs weighted calculations according to weighting rules, and generates a health index with a value range of zero to one. Zero is used as a complete failure state, and one is used as a complete health state. The system outputs the health index corresponding to each UAV.

8. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The generation of diagnostic decision results includes: Read the fault type, fault severity and health index corresponding to each monitored object, retrieve the health index threshold, fault severity threshold and operation and maintenance rule parameters corresponding to each monitored object from the storage medium, and set the health index threshold and fault severity threshold for each monitored object. For each monitored object, the health index is compared with the corresponding health index threshold, and the fault severity is compared with the corresponding fault severity threshold. Based on the comparison results, the operating status of the monitored object is divided into normal state, warning state and dangerous state, and the diagnostic level of the monitored object is determined in combination with the fault type. Logical judgments are made on the overall operating status based on the diagnostic level and operation and maintenance rule parameters of each monitored object. When any monitored object is in a dangerous state, a return command or emergency shutdown command is generated. When there is a monitored object in a warning state but not in a dangerous state, a load adjustment command and maintenance suggestion are generated. The generated commands and maintenance suggestions are combined to form a diagnostic decision result.

9. The intelligent fusion diagnostic analysis method for unmanned aerial vehicles based on multimodal fusion according to claim 1, characterized in that, The process of updating the cross-mission prototype memory based on diagnostic decision results provides historical diagnostic information and prototype feature support for online diagnostics and operational decisions for flight missions, including: Obtain the fusion diagnostic features, fault type, fault severity, health index and diagnostic decision results corresponding to the current flight mission, associate the flight mission identifier, flight time, flight phase information with fault type, fault severity and health index and store them in the health record, and update the corresponding prototype features or add new prototype features in the cross-mission prototype memory using fusion diagnostic features, fault type, health index range and flight phase as conditions. In the cross-task prototype memory, when the number of prototype features exceeds the capacity limit, the prototype features are filtered according to the generation time order of the prototype features, and the prototype feature record at the bottom is deleted to keep the number of prototype features in the cross-task prototype memory within the capacity limit.