Unmanned aerial vehicle data anomaly detection method based on multi-modal time sequence modeling

Through multi-modal timing modeling, multi-stage preprocessing and feature extraction are performed on UAV data, combined with multi-stage information fusion and adaptive threshold detection, the problem of insufficient accuracy of UAV data abnormal detection is solved, and more efficient abnormal response capabilities are achieved.

CN120337084APending Publication Date: 2025-07-18SICHUAN UNIV
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510482005.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing drone data anomaly detection methods are insufficient in accuracy and generalization capabilities, and it is difficult to effectively respond to abnormal responses in complex or emergencies.

Method used

The multi-modal timing modeling method is used to perform multi-level pre-processing of drone data. The timing features are extracted through the gated attention Transformer module of continuous LSTM blocks and layer scaling, combined with multi-stage information fusion and mask screening, and fused spatiotemporal correlation features are generated, and anomaly detection is performed using attitude parameter prediction and adaptive thresholds.

Benefits of technology

It improves the accuracy and robustness of the abnormal detection of drone data, can better deal with abnormal responses in complex environments, and improves the safety and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337084A_ABST
    Figure CN120337084A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an unmanned aerial vehicle data anomaly detection method based on multi-modal time sequence modeling, and the method comprises the steps: carrying out the multi-stage data preprocessing of original data obtained in the working process of an unmanned aerial vehicle, and generating a standardized input data flow; extracting time sequence features of the standardized input data stream through continuous LSTM blocks in the multi-level feature extraction network, and generating a shallow time sequence feature set; performing depth time sequence modeling on the shallow time sequence feature set to generate a deep time sequence feature set; performing multi-stage information fusion and mask screening on the shallow time sequence feature set and the deep time sequence feature set to generate a fusion time-space correlation feature set; and generating an attitude parameter prediction result of the unmanned aerial vehicle according to the fusion time-space correlation feature set, and obtaining an anomaly detection decision result of the unmanned aerial vehicle by using a deviation comparison and multi-level anomaly judgment mechanism based on an adaptive threshold obtained based on the attitude parameter prediction result and prediction error statistical distribution. Therefore, the anomaly detection accuracy of the unmanned aerial vehicle data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unmanned aerial vehicles, and more specifically, to a method for detecting abnormal unmanned aerial vehicle data based on multimodal time series modeling. Background Technique

[0002] Unmanned Aerial Vehicles (UAVs) are widely used in medical transportation, precision agriculture, environmental monitoring, disaster response, search and rescue, etc. due to their advantages such as small size, light weight, low cost, strong mobility, flexible tasks, and low risk coefficient. With the continuous improvement of the task complexity and operating environment diversity of UAVs, higher requirements are put forward for the safety and operating reliability of UAVs, especially the abnormal response ability of UAVs in complex or emergency situations.

[0003] However, in the detection of abnormal UAV flight data, traditional methods have many problems due to the characteristics of UAV systems, and it is urgent to develop an efficient detection model. Currently, knowledge-based methods rely on experts to build a system to detect abnormalities. However, the knowledge system of UAVs is huge, and it takes a long time for experts to obtain and encode knowledge, and the application is limited. To make up for its deficiencies, model-based methods can combine domain knowledge and system mechanisms to build physical models. However, due to the complexity of UAV systems and the variability of operating states with flight environments and task types, it is almost impossible to establish an accurate global model covering all situations. Moreover, it is difficult to obtain abnormal samples, and the performance evaluation of the model has also become a problem. Although data-driven methods based on deep learning have strong generalization ability, their detection performance is greatly affected by prediction ability and error accumulation, and they focus on single parameters and are difficult to grasp overall abnormalities.

[0004] Therefore, how to improve the accuracy of abnormal detection of UAV data is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the purpose of this application is to provide a method for detecting abnormal UAV data based on multimodal time series modeling to improve the accuracy of abnormal detection of UAV data.

[0006] Combined with the first aspect of this application, a method for detecting abnormal UAV data based on multimodal time series modeling is provided. The method includes: Perform multi-level data preprocessing on the original data obtained during the operation of the UAV to generate a normalized input data stream. The multi-level data preprocessing includes noise elimination processing, dimensionality reduction processing, normalization processing, and time sliding window reconstruction. The normalized input data stream is a data stream containing fixed time steps; Extract the time series features of the normalized input data stream through continuous LSTM blocks in a multi-level feature extraction network to generate a shallow time series feature set; The gated attention Transformer module based on layer scaling performs deep temporal modeling on the shallow temporal feature set to generate a deep temporal feature set; The shallow temporal feature set and the deep temporal feature set are subjected to multi-stage information fusion and mask screening to generate a fused spatio-temporal correlation feature set; Based on the fused spatio-temporal correlation feature set, the attitude parameter prediction result of the UAV is generated, and based on the attitude parameter prediction result and the adaptive threshold obtained from the statistical distribution of the prediction error, the anomaly detection decision result of the UAV is obtained by using the deviation comparison and multi-level anomaly judgment mechanism.

[0007] Optionally, the multi-level data preprocessing of the original data obtained during the operation of the UAV to generate a normalized input data stream includes: Noise elimination is performed on the original data to generate preliminary smoothed time series data; According to the preliminary smoothed time series data, a multi-dimensional feature data set corresponding to the original data is obtained; A multi-dimensional feature dimensionality reduction operation is performed on the multi-dimensional feature data set to obtain a set of principal component feature vectors of the multi-dimensional feature data set; The eigenvalues of the principal component feature vectors in the set of principal component feature vectors are scaled to a preset interval through normalization processing to obtain a set of normalized principal component feature vectors; The set of normalized principal component feature vectors is reconstructed by a sliding window to form the normalized input data stream including the fixed time step.

[0008] Optionally, the temporal features of the normalized input data stream are extracted by consecutive LSTM blocks in the multi-level feature extraction network to generate a shallow temporal feature set, including: According to the sequence of LSTM blocks in the multi-level feature extraction network, the normalized input data stream is extracted layer by layer through the consecutive LSTM blocks. The LSTM block includes an LSTM cell and a dropout layer, and adjacent LSTM blocks are connected through a hidden state transfer mechanism. The LSTM cell uses a forget gate to control the retention ratio of historical information, an input gate to control the update ratio of new information, and an output gate to control the output ratio of the current state. The dropout layer is used to randomly mask the outputs of some neurons according to a preset probability; After the layer-by-layer extraction is completed according to the LSTM block sequence, a shallow temporal feature set including multi-level abstract features is obtained. The feature dimension of the shallow temporal feature set at each time step is related to the number of feature vectors in the set of normalized principal component feature vectors.

[0009] Optionally, the layer-scaling based gated attention Transformer module performs deep temporal modeling on the shallow temporal feature set to generate a deep temporal feature set, including: Embed sine-cosine positional encoding information in the shallow temporal feature set to generate a position-enhanced feature set, where the sine-cosine positional encoding information is related to the sliding window structure reconstructed by the sliding window; Input the position-enhanced feature set into a gated attention unit, dynamically adjust the interaction weight between the query vector and the key vector through learnable parameters, and generate an attention weight matrix; Perform weighted aggregation on the value vector based on the attention weight matrix to generate preliminary attention features; Perform layer-scaling processing on the preliminary attention features, adjust the feature distribution of the preliminary attention features through a trainable scaling factor, and generate scaled attention features; Merge the scaled attention features and the original input features through residual connection, and generate a deep temporal feature set based on activation function and normalization processing.

[0010] Optionally, the multi-stage information fusion and mask screening of the shallow temporal feature set and the deep temporal feature set to generate a fused spatio-temporal correlation feature set includes: Concatenate the shallow temporal feature set and the deep temporal feature set in the channel dimension to generate a joint feature tensor; Perform linear transformation on the joint feature tensor to compress the feature dimension and generate a gated weight vector; Use the gated weight vector to perform weighted fusion on the shallow features in the shallow temporal feature set and the deep features in the deep temporal feature set respectively to obtain preliminary fusion features; Apply a channel-level mask operation to the preliminary fusion features, activate the effective feature channels through a non-linear function, and generate a fused spatio-temporal correlation feature set. The channel-level mask operation adopts a dynamic threshold mechanism and automatically selects the activation threshold according to the importance score of the feature channels.

[0011] Optionally, generating the attitude parameter prediction result of the drone according to the fused spatio-temporal correlation feature set, and based on the attitude parameter prediction result and the adaptive threshold obtained from the statistical distribution of the prediction error, obtaining the abnormal detection decision result of the drone by using a deviation comparison and multi-level anomaly judgment mechanism, including: Input the fused spatio-temporal correlation feature set into the flatten layer of the multi-layer perceptron prediction head, and convert the fused spatio-temporal correlation feature set into a two-dimensional tensor feature set; Input the two-dimensional tensor feature set into the multi-layer fully connected network of the multi-layer perceptron prediction head for linear transformation processing of different dimensions to obtain a feature tensor to be mapped; Map the feature tensor to be mapped to the target dimensional space to obtain the pose parameter prediction result; Obtain a set of deviation values based on the pose parameter prediction result and reference data, where the reference data is obtained based on the real sensor data of the drone; Statistically calculate the deviation mean and deviation standard deviation of each pose parameter based on historical validation set data, and determine the initial threshold using a normal distribution; Dynamically adjust the weight coefficient of the initial threshold according to the current flight phase of the drone to generate an adaptive threshold set. Each pose parameter in the adaptive threshold set independently calculates an adjustment coefficient, and the adjustment coefficient is related to the flight speed, environmental interference intensity, and sensor confidence index; Obtain the anomaly detection decision result based on the set of deviation values and the comparison result of the adaptive threshold. The anomaly detection decision result includes multi-level anomaly alarms, and the multi-level anomaly alarms include single-point instantaneous alarms, continuous anomaly alarms, and trend deviation alarms. Different-level alarms correspond to different flight control system response strategies.

[0012] Optionally, it further includes: Inject at least one of the following anomaly patterns into the normalized input data stream: insertion anomaly pattern, offset anomaly pattern, drift anomaly pattern, static anomaly pattern. The insertion anomaly pattern is injected by randomly selecting a time point in the normalized input data stream to add a first data point, and the first data point is a data point whose value exceeds the preset data value range. The offset anomaly pattern is injected by adding a fixed offset to the data points within the first target time period in the normalized input data stream. The drift anomaly pattern is injected by adding an offset that increases or decreases with time to the data points within the second target time period in the normalized input data stream. The static anomaly pattern is injected by setting the values of the data points within the third target time period in the normalized input data stream to the same value.

[0013] Optionally, it further includes: Establish an anomaly pattern feature library, where the anomaly patterns in the anomaly pattern feature library include anomaly patterns of sensor hardware failures, signal interference, and environmental impacts; Generate a set of fault causes by performing similarity matching between the abnormal data in the anomaly detection decision result and each anomaly pattern in the anomaly pattern feature library. The similarity matching uses the dynamic time warping algorithm to align time series of different lengths and calculates the cosine similarity score; Analyze the abnormal propagation paths of each of the fault causes in the set of fault causes based on the attention weight matrix output by the layer-scaling based gated attention Transformer module, determine the root sensor nodes, and the analysis of the abnormal propagation paths is realized through visualizing the attention heat map. The high-weight connection paths in the abnormal propagation paths point to potential fault sources. Combine the flight state parameters of the drone and the historical maintenance records, and calculate the probability distribution of each of the fault causes in the set of fault causes. Generate hierarchical warning information and maintenance suggestions according to the probability distribution.

[0014] Optionally, it further includes: Based on the original data corresponding to the multi-scene flight states, obtain the normalized input data stream corresponding to the multi-scene flight states, where the multi-scene flight states include at least two of stable flight, high-speed maneuvering, and strong wind interference. Configure the adjustment weights of the normalized input data streams corresponding to each of the scene flight states according to the recognition difficulty of each scene flight state in the multi-scene flight states. Adjust the normalized input data stream input into the multi-level feature extraction network according to the adjustment weights.

[0015] Optionally, it further includes: Dynamically adjust the adaptive threshold according to the influence of the flight environment change, current flight state, and historical flight data of the drone on the predicted result of the attitude parameters, and the prediction accuracy of the predicted result of the attitude parameters. And / or According to the historical prediction results of the normalized input data stream input into the multi-level feature extraction network after adjustment, obtain the typical scene features corresponding to the historical prediction results during the prediction process. Determine the typical scene features corresponding to the current scene flight state according to the current scene flight state of the drone. Pre-input the typical scene features for prediction when predicting the predicted result of the attitude parameters of the drone.

[0016] Combined with the second aspect of the present application, there is provided a drone data anomaly detection system based on multi-modal time series modeling. The drone data anomaly detection system based on multi-modal time series modeling includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the drone data anomaly detection system based on multi-modal time series modeling implements the aforementioned drone data anomaly detection method based on multi-modal time series modeling.

[0017] In combination with the third aspect of the present application, there is provided a computer-readable storage medium storing computer-executable instructions, which when executed, implement the aforementioned method for detecting abnormal drone data based on multimodal time series modeling.

[0018] In combination with the fourth aspect of the present application, there is provided a computer program product, which when executed by a processor, implements the aforementioned method for detecting abnormal drone data based on multimodal time series modeling.

[0019] In combination with any of the above aspects, by performing multi-level data preprocessing on the original data obtained during the operation of the drone, a normalized input data stream is generated; the time series features of the normalized input data stream are extracted by consecutive LSTM blocks in a multi-level feature extraction network to generate a shallow time series feature set; deep time series modeling is performed on the shallow time series feature set to generate a deep time series feature set; the shallow time series feature set and the deep time series feature set are subjected to multi-stage information fusion and mask screening to generate a fused spatio-temporal correlation feature set; based on the fused spatio-temporal correlation feature set, a prediction result of the attitude parameters of the drone is generated, and based on the prediction result of the attitude parameters and the adaptive threshold obtained from the statistical distribution of the prediction error, an abnormal detection decision result of the drone is obtained by using a deviation comparison and multi-level abnormal judgment mechanism, thereby improving the accuracy of abnormal detection of drone data. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained in combination with these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of a method for detecting abnormal drone data based on multimodal time series modeling provided by an embodiment of the present application. Detailed Embodiments

[0022] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0023] The terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal comprising a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or terminals.

[0024] Reference to "embodiment" herein means that a particular feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0025] Figure 1 The figure is a schematic flow chart of a method for detecting abnormal drone data based on multimodal time series modeling provided for an embodiment of the present application. It should be understood that in other embodiments, the order of some steps of the method for detecting abnormal drone data based on multimodal time series modeling in this embodiment can be shared according to actual needs, or some of the steps can also be omitted or maintained. The details of the method for detecting abnormal drone data based on multimodal time series modeling include: S101. Perform multi-level data preprocessing on the raw data obtained during the operation of the drone to generate a normalized input data stream.

[0026] Among them, the multi-level data preprocessing includes noise elimination processing, dimensionality reduction processing, normalization processing, and time sliding window reconstruction, and the normalized input data stream is a data stream containing fixed time steps.

[0027] Raw data refers to the data collected by various sensors (such as GPS, inertial measurement unit IMU, barometer, etc.) during the flight of the drone. This data usually has the characteristics of high noise, high dimension, and complexity. Exemplarily, it can include the attitude data of the drone (for example, the rotation state of the drone is characterized by quaternions q[0], q[1], q[2], q[3]), speed data (for example, it can include the magnitude of the speed, speed components, etc.), acceleration data (for example, it includes the acceleration measured by the accelerometer, acceleration components, etc.), angular velocity data, position data (for example, longitude, latitude, altitude, position deviation, etc.), navigation and positioning data (for example, it includes positioning type, number of satellites, dilution of precision, position accuracy, etc.), sensor data (for example, accelerometer data, speedometer data, barometric altitude, satellite signal quality, etc.), time data (for example, timestamp), and other data (for example, reference point data, rate of change of altitude, etc.).

[0028] Noise cancellation processing is used to remove the noise in the raw data, improve the signal-to-noise ratio of the data, and provide a high-quality data basis for subsequent data processing and analysis. For example, the sensor signals can be smoothed by a Kalman filter. In addition, due to the high dimension of the raw data, direct processing will increase the computational complexity. Therefore, dimensionality reduction processing (such as dimensionality reduction by principal component analysis) can be used to retain the main feature information in the data while reducing the data dimension and improving the processing efficiency. Normalization processing refers to scaling the data to a specific interval (such as [-1,1]), eliminating the dimensional differences between different features, and improving the convergence speed and performance of the model. Time sliding window reconstruction refers to dividing the continuous data stream into data segments with a fixed time step to form the basic data units for model input, facilitating time series analysis and prediction.

[0029] In this step, a Kalman filter can be used to smooth the sensor signals in the raw data. The Kalman filter can calculate the optimal estimate of the current state based on the current measurement value and the estimated value at the previous moment through the state transition matrix and the observation matrix, thereby removing the noise.

[0030] Methods such as principal component analysis or factor analysis are used to reduce the dimension of the multi-dimensional features included in the raw data. Taking principal component analysis as an example, the covariance matrix between the features in the raw data can be calculated to find several principal component directions with the largest variances, and the projections in these directions are used as the new feature vectors, reducing the data dimension while retaining the main information.

[0031] By using the sliding window technique, the original data in the form of time series data can be segmented into data segments with a fixed time step. For example, if the window size is set to 10, each window contains 10 data samples at each time step. The sliding window technique can cover the entire time series and provide a standardized input data stream for subsequent model training.

[0032] By performing data preprocessing, a standardized input data stream can be obtained, which can improve the prediction efficiency and accuracy of subsequent models.

[0033] Exemplarily, suppose there is a drone used for agricultural mapping flying over farmland. During the flight, position information can be obtained through GPS, attitude and acceleration information can be obtained through an inertial measurement unit (IMU), and altitude information can be obtained through a barometer, etc. Due to factors such as environmental interference and sensor accuracy, there is a lot of noise in this data, and the data dimension is relatively high. For example, the GPS signal may be blocked by clouds, buildings, etc., resulting in jumps in the position data; the IMU data may generate noise due to vibrations.

[0034] For example, at a certain moment, the position recorded by the drone's GPS is the longitude and latitude (116.4074, 39.9042), and the altitude is 100 meters. However, due to signal interference, the position recorded at the next moment becomes (116.4076, 39.9044), and the altitude becomes 102 meters, with relatively large fluctuations. Through the Kalman filter, based on the current measurement value and the estimated value at the previous moment, the optimal estimated value of the current state is calculated to obtain smoother position data, such as the longitude and latitude (116.4075, 39.9043), and the altitude is 100.5 meters, effectively removing the noise.

[0035] Dimensionality reduction processing: Since the original data contains features of multiple dimensions, such as attitude quaternion (4 dimensions), acceleration (3 dimensions), angular velocity (3 dimensions), etc. Suppose after PCA analysis, it is found that the first 3 principal component directions can retain most of the data information. Then, the projections in these directions are used as new feature vectors, reducing the data dimension from 10 dimensions to 3 dimensions and reducing the computational complexity.

[0036] Normalization processing: Scale the data after dimensionality reduction to the interval [-1, 1]. For example, if a feature value obtained after dimensionality reduction is 50, after normalization processing, it is scaled to the interval [-1, 1]. Suppose the result after normalization is 0.3.

[0037] Time-sliding window reconstruction: Set the window size to 10, and divide the continuous data stream into data segments with a fixed time step. For example, if the drone collects data once per second, then each window contains 10 seconds of data samples. The data from the 1st second to the 10th second is taken as one window, and the data from the 2nd second to the 11th second is taken as the next window, and so on, covering the entire time series to form a normalized input data stream.

[0038] S102. Extract the temporal features of the normalized input data stream through consecutive LSTM blocks in the multi-level feature extraction network to generate a set of shallow temporal features.

[0039] Long Short-Term Memory (LSTM) blocks are used to capture long-term dependencies in time series data. In this application, a sequence structure can be formed by multiple LSTM blocks to extract the temporal features of the normalized input data stream layer by layer, thereby increasing the depth of the temporal features extracted from the normalized input data stream and obtaining a set of shallow temporal features obtained by deeply mining the normalized input data stream. In this application, each LSTM block contains an LSTM unit with a fixed hidden layer dimension and a dropout layer. The LSTM unit is used to capture temporal features, while the dropout layer is used to prevent overfitting.

[0040] In this step, the normalized input data stream can be input into the first LSTM block to extract the corresponding set of temporal features, and then the set of temporal features output by the first LSTM block is input into the next LSTM block for further extraction according to the LSTM sequence, and so on, until the output result of the last LSTM block, which is the set of shallow temporal features.

[0041] Exemplarily, assume that the normalized input data stream contains 10 time steps, and each time step has 3 feature dimensions. Taking the 1st time step as an example, the 3 feature values are 0.22, 0.33, and 0.11 respectively; the feature values of the 2nd time step are 0.25, 0.36, and 0.13; the feature values of the 3rd time step are 0.28, 0.39, and 0.15, and so on.

[0042] Input this data segment into the first LSTM block, which contains an LSTM cell with a fixed hidden layer dimension of 5 and a dropout layer (dropout rate of 0.2). At the first time step, the LSTM cell receives feature values 0.22, 0.33, 0.11, combines with the hidden state of the previous time step (initially a zero vector), and updates the current hidden state. Assume the updated hidden state is 0.12, 0.23, 0.34, 0.45, 0.56. After passing through the dropout layer, some neurons are randomly set to zero, and assume the resulting output is 0.12, 0, 0.34, 0.45, 0.

[0043] The set of temporal features output by the first LSTM block is input into the second LSTM block (also containing an LSTM cell with a hidden layer dimension of 5 and a dropout layer with a dropout rate of 0.2) for further extraction. After being processed by multiple LSTM blocks, a set of shallow temporal features is finally output. The set of shallow temporal features contains 10 time steps, and each time step has 5 feature dimensions. For example, the 5 feature values at the first time step are 0.21, 0.32, 0.13, 0.44, 0.25; the 5 feature values at the second time step are 0.24, 0.35, 0.16, 0.47, 0.28, and so on.

[0044] S103. Perform deep temporal modeling on the set of shallow temporal features using a layer-scaled gated attention Transformer module to generate a set of deep temporal features.

[0045] The Transformer module refers to a deep learning model based on the self-attention mechanism, which can capture global dependencies in data through the self-attention mechanism. The layer-scaled gated attention Transformer (LSGA-T) module refers to an improved Transformer module that realizes deeper feature modeling and more efficient computation by introducing a layer-scaling mechanism and a gated attention unit.

[0046] In this application, the LSGA-T module contains multiple Blocks, and each Block consists of components such as a gated attention unit (GAU), a layer-scaling mechanism, a residual connection, a GELU activation function, and LayerNorm. The GAU realizes adaptive feature interaction by introducing learnable scaling factors and bias parameters. Based on the self-attention mechanism, the GAU can more effectively capture global dependencies in data. The layer-scaling mechanism finely adjusts features through a parameterized layer-scaling factor to enhance the stability of model training.

[0047] In this step, the shallow temporal feature set can be input into the LSGA-T module. After being processed by multiple Blocks, a deep temporal feature set is output. Among them, in each Block, first, self-attention calculation can be performed on the input features through GAU to capture global dependencies. Then, the layer scaling mechanism is used to finely adjust the features to enhance the training stability of the model. Next, components such as residual connections, GELU activation functions, and LayerNorm are applied to optimize the gradient propagation path and training stability. Finally, the outputs of multiple Blocks are used as the deep temporal feature set for subsequent steps. Optionally, multiple LSGA-T modules can also be set in this step. These multiple LSGA-T modules form an LSGA-T sequence. By inputting the shallow temporal feature set into the first LSGA-T module for temporal modeling, the corresponding output result is obtained, and this output result is used as the input for the next LSGA-T module for further temporal modeling, and so on, until the output result of the last LSGA-T module, which is the deep temporal feature set.

[0048] Exemplarily, assume that the above-mentioned shallow temporal feature set is input into the LSGA-T module, and the LSGA-T module contains 2 Blocks.

[0049] In the first Block: GAU performs self-attention calculation on the input shallow temporal features. For example, for the features 0.21, 0.32, 0.13, 0.44, 0.25 at the first time step, it will consider the relationships with the features at other time steps. Assume that after being calculated by GAU, the features at the first time step are updated to 0.31, 0.42, 0.23, 0.54, 0.35.

[0050] The layer scaling mechanism finely adjusts the features through a parameterized layer scaling factor. Assume that the layer scaling factor is 0.8, then the features at the first time step become 0.31 * 0.8 = 0.248, 0.42 * 0.8 = 0.336, 0.23 * 0.8 = 0.184, 0.54 * 0.8 = 0.432, 0.35 * 0.8 = 0.28.

[0051] Residual connection, GELU activation function, and LayerNorm: Add the original first-time-step features 0.21, 0.32, 0.13, 0.44, 0.25 to the layer-scaled features for residual connection, obtaining 0.21 + 0.248 = 0.458, 0.32 + 0.336 = 0.656, 0.13 + 0.184 = 0.314, 0.44 + 0.432 = 0.872, 0.25 + 0.28 = 0.53. After being processed by the GELU activation function, assume we get 0.4, 0.6, 0.2, 0.8, 0.4. Then, through LayerNorm for normalization, the features at the first time step finally become 0.32, 0.51, 0.12, 0.72, 0.33.

[0052] After being processed by the first Block, a new feature set is obtained. Input this new feature set into the second Block for the same processing. Finally, a deep temporal feature set is output. For example, the 5 feature values at the first time step are 0.32, 0.51, 0.12, 0.72, 0.33; the 5 feature values at the second time step are 0.43, 0.62, 0.23, 0.83, 0.44, and so on.

[0053] S104. Perform multi-stage information fusion and mask screening on the shallow temporal feature set and the deep temporal feature set to generate a fused spatio-temporal correlation feature set.

[0054] In this step, according to the correlation of the features in the shallow temporal feature set and the deep temporal feature set, the features in the shallow temporal feature set and the deep temporal feature set can be subjected to multi-stage information fusion, and then the fused features are screened through a mask mechanism, retaining the valid features and removing the invalid or redundant features to improve the efficiency and accuracy of feature representation.

[0055] Specifically, the shallow temporal feature set and the deep temporal feature set can be concatenated in the channel dimension to form a fused feature representation. Use a linear layer to effectively compress and transform the information of the fused feature representation to obtain gating weights. Assign the gating weights to the shallow information and the deep information respectively to achieve weighted fusion. During the weighted fusion process, the weight allocation can be dynamically adjusted according to the importance and correlation of the features. Then, apply a channel-level feature mask to the fused features to achieve effective screening of information. During the mask screening process, the mask value can be dynamically adjusted according to the importance and correlation of the features, retaining the valid information and removing the invalid or redundant information to generate a fused spatio-temporal correlation feature set.

[0056] For example, assume that the shallow time series feature set has 5 feature dimensions per time step, and the deep time series feature set also has 5 feature dimensions per time step, and after connection, a fused feature representation containing 10 feature dimensions is formed. Taking the first time step as an example, the fused features are 0.21, 0.32, 0.13, 0.44, 0.25, 0.32, 0.51, 0.12, 0.72, and 0.33. Assume that the linear layer compresses the features of 10 channels to 3 channels and obtains the gating weights. For the fused features of the first time step of 0.21, 0.32, 0.13, 0.44, 0.25, 0.32, 0.51, 0.12, 0.72, and 0.33, after being processed by the linear layer, the gating weights are 0.22, 0.33, and 0.45. Multiply the shallow information of the first time step 0.21, 0.32, 0.13, 0.44, 0.25 by the gate weight 0.22, and the deep information 0.32, 0.51, 0.12, 0.72, 0.33 by the gate weights 0.33 and 0.45, and add them together. Assume that the fused features are 0.26, 0.37, and 0.17. Assuming the mask values are 1, 0, and 1, the first and third features in the fused features of the first time step are retained, and the second feature is removed. Finally, the features of the first time step become 0.26 and 0.17. After processing all time steps, a fused spatiotemporal correlation feature set is generated. For example, the feature values of the first time step are 0.26 and 0.17; the feature values of the second time step are 0.37 and 0.28, and so on.

[0057] S105. Generate the attitude parameter prediction result of the UAV according to the fused spatiotemporal correlation feature set, and obtain the abnormal detection decision result of the UAV by using deviation comparison and multi-level abnormal judgment mechanism based on the attitude parameter prediction result and the adaptive threshold obtained based on the statistical distribution of the prediction error.

[0058] Among them, the attitude parameter prediction result refers to the result obtained by predicting the attitude parameters of the drone through the model. The adaptive threshold is a threshold that is dynamically adjusted based on the statistical distribution of the prediction error, and is used to determine whether the prediction result is abnormal. Deviation comparison refers to comparing the deviation between the current sensor data and the predicted value with the adaptive threshold. The multi-level anomaly judgment mechanism refers to a mechanism that combines multiple anomaly types for judgment. The anomaly detection decision result refers to the judgment result on whether the drone has an anomaly and the type and severity of the anomaly.

[0059] In this step, the predicted head of the pre-constructed multi-layer perceptron can be used to process the fused spatio-temporal correlation feature set to obtain the predicted result of the attitude parameters of the UAV. Then, the predicted result of the attitude parameters can be compared with the data obtained by the sensors of the UAV to obtain the deviation between the two. Then, the adaptive threshold obtained based on the statistical distribution of the prediction error is compared with this deviation, and the multi-level anomaly judgment mechanism is used to determine whether there is an anomaly in the flight data of the UAV. If there is an anomaly, the anomaly detection decision result corresponding to this anomaly is obtained to support the decision-making of the flight control system and ensure the working safety of the UAV.

[0060] Exemplarily, it is assumed that the pitch angle of the UAV at a certain future moment is predicted to be 7°, the roll angle is 5°, and the yaw angle is 13°. It is assumed that the pitch angle actually measured by the sensor is 9°, the roll angle is 7°, and the yaw angle is 16°. Then the pitch angle deviation is 2°, the roll angle deviation is 2°, and the yaw angle deviation is 3°. It is assumed that the adaptive thresholds for the pitch angle, roll angle, and yaw angle are 3°, 3°, and 4° respectively. Comparing the deviation with the adaptive threshold, since the pitch angle deviation of 2° is less than the adaptive threshold of 3°, the roll angle deviation of 2° is less than the adaptive threshold of 3°, and the yaw angle deviation of 3° is less than the adaptive threshold of 4°, at this moment, according to the multi-level anomaly judgment mechanism, it is judged that there is no anomaly in the flight data of the UAV. If a certain deviation is greater than the corresponding adaptive threshold, it is judged that there is an anomaly, and the anomaly detection decision result corresponding to this anomaly is obtained, such as the anomaly type is attitude deviation anomaly and the severity is mild anomaly, to support the decision-making of the flight control system and ensure the working safety of the UAV.

[0061] The method provided by the embodiment of the present application generates a standardized input data stream by performing multi-level data preprocessing on the original data obtained during the operation of the UAV; extracts the temporal features of the standardized input data stream through consecutive LSTM blocks in the multi-level feature extraction network to generate a shallow temporal feature set; performs deep temporal modeling on the shallow temporal feature set to generate a deep temporal feature set; performs multi-stage information fusion and mask screening on the shallow temporal feature set and the deep temporal feature set to generate a fused spatio-temporal correlation feature set; generates the predicted result of the attitude parameters of the UAV according to the fused spatio-temporal correlation feature set, and based on the predicted result of the attitude parameters and the adaptive threshold obtained based on the statistical distribution of the prediction error, uses deviation comparison and multi-level anomaly judgment mechanism to obtain the anomaly detection decision result of the UAV, thereby improving the accuracy of anomaly detection of UAV data.

[0062] Next, for Figure 1 the specific content of each step in the embodiment is introduced in detail: S1011. Eliminate the noise of the original data to generate preliminary smoothed time series data.

[0063] In this step, the Kalman filter can be used to eliminate noise from the original data and generate preliminary smoothed time series data. This process can include the following formula: Among them, the meanings of the parameters in the formula are introduced as follows. : State vector at time t (true value of UAV attitude parameters), : State transition matrix, describing the dynamic characteristics of the system, : Control input matrix, : Control input (such as UAV control commands), : Process noise, assumed to be Gaussian white noise, : Observation value at time t (sensor readings), : Observation matrix, : Observation noise, assumed to be Gaussian white noise, : State estimate at time t, : Kalman gain, : Prior estimate error covariance, : Observation noise covariance.

[0064] S1012. Obtain the multi-dimensional feature dataset corresponding to the original data according to the preliminary smoothed time series data.

[0065] In this step, various data features of the data can be extracted and integrated from the preliminary smoothed time series data. These features can include, for example, but are not limited to, position features, speed features, change trend features of various data changes (such as change trend features of position changes, acceleration features of speed changes), etc. Then, the above-mentioned extracted and integrated features are arranged in chronological order to obtain the multi-dimensional feature dataset corresponding to the original data.

[0066] S1013. Perform multi-dimensional feature dimensionality reduction operation on the multi-dimensional feature dataset to obtain the set of principal component feature vectors of the multi-dimensional feature dataset.

[0067] In this step, first, according to the multi-dimensional feature dataset, calculate the covariance matrix between the features in the multi-dimensional feature dataset. The covariance matrix reflects the correlation between the features. Perform eigenvalue decomposition on the covariance matrix to obtain eigenvalues and corresponding eigenvectors. The eigenvalues represent the variance magnitudes of each principal component, and the eigenvectors represent the directions of the principal components. Then, sort the eigenvectors in descending order of eigenvalues, and select the top k eigenvectors, where k is the number of principal components that are desired to be retained. The selection of k can be determined according to the cumulative variance contribution rate. Usually, select the top k principal components whose cumulative variance contribution rate reaches a certain threshold (such as 90%). Combine the selected k eigenvectors together to form a set of principal component feature vectors. This set contains the main information in the multi-dimensional feature dataset while reducing the dimensionality.

[0068] S1014. Scale the eigenvalues of the principal component feature vectors in the set of principal component feature vectors to a preset interval through normalization processing to obtain a set of normalized principal component feature vectors.

[0069] Scale the eigenvalues in the set of principal component feature vectors to a preset interval (such as [-1, 1] or [0, 1]) to obtain a set of normalized principal component feature vectors. For example, the eigenvalues in the set of principal component feature vectors can be scaled to the preset interval through linear normalization (such as Min-Max normalization) or Z-score normalization, etc. Specifically, how to perform the normalization processing can refer to the prior art and will not be elaborated here.

[0070] S1015. Perform sliding window reconstruction on the set of normalized principal component feature vectors to form a standardized input data stream containing fixed time steps.

[0071] In this step, first determine the size of the sliding window according to actual needs, that is, the length of the fixed time step. For example, the window size is n, indicating that each window contains data of n time steps.

[0072] Starting from the starting position of the set of normalized principal component feature vectors, successively select consecutive data segments of length n as a window. After each selection, slide the window backward by one time step and continue to select the next window until the entire set of normalized principal component feature vectors is traversed.

[0073] Arrange all the selected windows in chronological order to form a standardized input data stream. Each window in this data stream contains data of fixed time steps, which is suitable for subsequent time series analysis and model training.

[0074] The method of this embodiment, through data preprocessing of the original data of the unmanned aerial vehicle, obtains a standardized input data stream that can improve the prediction efficiency and accuracy of subsequent models, thereby improving the efficiency and accuracy of subsequent predictions.

[0075] Optionally, at least one of the following abnormal patterns may also be injected into the normalized input data stream: insertion abnormal pattern, offset abnormal pattern, drift abnormal pattern, static abnormal pattern. By injecting these abnormal patterns, the model's recognition ability for various types of abnormalities is enhanced.

[0076] Among them, the insertion abnormal pattern is injected by randomly selecting a time point in the normalized input data stream to add a first data point, and the first data point is a data point whose value exceeds the preset data value range. For example, when a drone is flying normally, the value range of the altitude sensor is usually between 0 and 1000 meters (assumed). If an insertion abnormal pattern is selected to be injected at a certain time point, the value of the altitude sensor at that time point may be replaced with 1500 meters (a value exceeding the normal range), so as to simulate the situation where an abnormal value suddenly appears in the sensor. By randomly inserting values outside the normal range in this way, data containing such abnormal situations is provided for model training, thereby enhancing the model's recognition ability for this type of single-point sudden abnormality.

[0077] The offset abnormal pattern is injected by adding a fixed offset to the data points within the first target time period in the normalized input data stream. Among them, the first target time period can be selected according to actual needs, and this application does not limit it. For the sensor values involved in each data point of the first target time period, a fixed offset is added. This fixed offset is a preset constant. For example, for the speed sensor of a drone, within the first target time period, a fixed offset of 5 (assumed unit: m / s) is added to the speed value at each time point, so that the original normal speed value is offset, simulating the situation where the sensor is affected by a certain continuous fixed interference and the data as a whole deviates from the normal range. By adding a fixed offset within a specific time period, data samples containing such abnormalities are generated to help the model learn and recognize this offset abnormal pattern.

[0078] The drift abnormal pattern is injected by adding an offset that increases or decreases with time to the data points within the second target time period in the normalized input data stream. Among them, the second target time period can be selected according to actual needs, and this application does not limit it. For the sensor values in each data point within the second target time period, an offset that increases or decreases with time is added. The specific offset calculation can refer to the following logic: as time progresses from the abnormal start time to the end time, the offset gradually changes according to a certain ratio.

[0079] For example, assume the abnormal start time is , the end time is , the current time is , and a coefficient related to the sensor is preset in advance, then the change of the offset is related to is associated with, and as changes, the offset changes accordingly. Taking the acceleration sensor of a drone as an example, within the second target time period, as time goes by, an offset is gradually added to the acceleration value to simulate the situation where sensor data drifts slowly. The abnormal data generated in this way enables the model to learn the characteristics during data drift and improves the recognition ability for drift abnormal patterns.

[0080] The static abnormal pattern is injected by setting the values of data points within the third target time period in the normalized input data stream to the same value. Among them, the third target time period can be selected according to actual needs, and this application does not limit it. Within this third target time period, the values of all data points are set to the same value. This same value can be a preset constant. For example, for the attitude angle sensor of a drone, within the third target time period, the attitude angle values at all time points are set to a fixed value. For example, assuming the pitch angle is set to 30 degrees, it simulates the situation where the sensor is in a faulty state and the data does not change for a period of time. By generating data containing static anomalies in this way, the model can identify the abnormal pattern where the sensor data does not change for a long time, thereby enhancing the comprehensive recognition ability for various anomalies.

[0081] Comprehensively applying the above abnormal patterns enhances the model's recognition ability for various anomalies and improves the recognition accuracy and sensitivity of the system to different types of anomalies.

[0082] S1021. According to the LSTM block sequence in the multi-level feature extraction network, the normalized input data stream is extracted layer by layer through consecutive LSTM blocks.

[0083] Among them, the LSTM block includes an LSTM cell and a dropout layer, and adjacent LSTM blocks are connected through a hidden state transfer mechanism. The LSTM cell uses a forget gate to control the retention ratio of historical information, an input gate to control the update ratio of new information, and an output gate to control the output ratio of the current state. The dropout layer is used to randomly mask the outputs of some neurons according to a preset probability.

[0084] Each LSTM block is implemented based on the content included in the following formula: Among them, : time, : Forget gate, which controls the proportion of information to be discarded, : Input gate, which controls the proportion of information to be updated, : Candidate memory cell, : Current memory cell, : Output gate, which controls the proportion of information to be output, : Hidden state, which serves as the output at the current time step, 、 、 、 : Weight matrix, 、 、 、 : Bias vector, : Sigmoid activation function, : Hadamard product (element-wise multiplication).

[0085] In this step, the normalized input data stream is input into the first LSTM block of the multi-level feature extraction network. Inside this LSTM block, the LSTM cells process the input data at each time step. Inside the LSTM cell, based on the current input and the previous hidden state, the output of the forget gate is calculated. The role of the forget gate is to determine which information in the previous cell state needs to be forgotten. The output value is mapped to the interval [0,1] through a Sigmoid function. The closer the value is to 0, the higher the degree of forgetting; the closer it is to 1, the higher the degree of retention. Based on the current input and the previous hidden state, the output of the input gate is calculated. The input gate controls the update proportion of new information. First, the Sigmoid function is used to determine which values need to be updated, and then the Tanh function is used to create new candidate values. Combining the outputs of the forget gate and the input gate, the cell state is updated. That is, first multiply the previous cell state by the output of the forget gate, and then add the product of the output of the input gate and the candidate value. Based on the current input, the previous hidden state, and the updated cell state, the output of the output gate is calculated. The output gate controls the output proportion of the current state. The Sigmoid function is used to determine which parts of the cell state will be output to the hidden state. Finally, multiply the output of the output gate by the cell state processed by the Tanh function to obtain the hidden state at the current moment.

[0086] Then, the hidden state output by the LSTM cell is input into the dropout layer. The dropout layer will randomly mask the outputs of some neurons according to a preset probability (such as 0.2). This can prevent the model from overfitting and improve the generalization ability of the model.

[0087] After that, the result output by the first LSTM block (the hidden state after passing through the dropout layer) is used as the input of the next LSTM block. Adjacent LSTM blocks are connected through a hidden state transfer mechanism, that is, the hidden state output by the previous LSTM block at each time step is used as the input of the corresponding time step of the next LSTM block. Repeat the above process of LSTM units and dropout layers until the hierarchical extraction of all LSTM blocks is completed.

[0088] S1022. After completing the hierarchical extraction according to the LSTM block sequence, obtain a shallow temporal feature set containing multi-level abstract features.

[0089] Among them, the feature dimension of the shallow temporal feature set at each time step is related to the number of feature vectors in the candidate input data stream. Specifically, during the process of being processed by a series of LSTM blocks, the feature dimension will change according to the hidden layer dimension of the LSTM unit. Generally, the final feature dimension is equal to the hidden layer dimension of the LSTM unit in the last LSTM block. And the hidden layer dimension of the LSTM unit can be reasonably set according to the complexity of the candidate input data stream and the number of feature vectors when designing the model to ensure that the feature information in the data can be fully extracted. Combine the hidden states at each time step to form a complete shallow temporal feature set. This set can be represented as a two-dimensional matrix, where the number of rows is equal to the number of time steps, and the number of columns is equal to the feature dimension of each time step.

[0090] When the normalized input data stream undergoes hierarchical extraction through all LSTM blocks, the hidden states output by the last LSTM block at each time step constitute a feature sequence. This feature sequence contains multi-level abstract features extracted from the input data.

[0091] The method of the embodiment of the present application, according to the LSTM block sequence in the multi-level feature extraction network, performs hierarchical extraction on the normalized input data stream through consecutive LSTM blocks. After completing the hierarchical extraction according to the LSTM block sequence, obtain a shallow temporal feature set containing multi-level abstract features, thereby effectively capturing the long-term dependencies in the normalized input data stream, extracting more in-depth and abstract feature information, providing a rich and high-quality data basis for subsequent deep temporal modeling, and improving the model's understanding and processing ability of temporal data.

[0092] S1031. Embed sine-cosine position encoding information in the shallow temporal feature set to generate a position-enhanced feature set.

[0093] Among them, the sine-cosine position encoding information is related to the sliding window structure reconstructed by the sliding window.

[0094] In the foregoing embodiments, the data has been reconstructed through the sliding window technique to form a sliding window with a specific structure. Based on this sliding window structure, sine-cosine position encoding information is generated. The purpose of position encoding is to introduce the position information of the elements in the sequence into the model because the order of the elements is crucial when processing time series data.

[0095] The principle of sine-cosine position encoding is to utilize the periodicity of trigonometric functions to generate a unique encoding vector for each position. For a sequence of length L, the j-th dimension of the encoding vector at position i can be calculated by the following formula: For even dimensions 2j: For odd dimensions 2j + 1: Where, is the dimension of the feature.

[0096] After generating the sine-cosine position encoding information, it is fused with the shallow time series feature set by addition. That is, for each feature vector in the shallow time series feature set, and the corresponding position encoding vector , a new position-enhanced feature vector is obtained. In this way, a position-enhanced feature set is generated.

[0097] In this step, by introducing position encoding, the model can better capture the temporal dependencies in the time series and improve the sensitivity to the time order.

[0098] S1032: Input the position-enhanced feature set into the gated attention unit, and dynamically adjust the interaction weight between the query vector and the key vector through learnable parameters to generate an attention weight matrix.

[0099] First, calculate the intermediate feature representation: Use the position-enhanced feature set as the input , and perform a linear transformation through the formula to obtain the intermediate feature representation . Where, is a learnable weight matrix, and its dimension is determined according to the input feature dimension and the expected intermediate feature dimension; is a bias vector, and its dimension is the same as that of . This step maps the input features to a new feature space.

[0100] Then generate the query vector and the key vector : Further process the intermediate feature representation through the formula to obtain an intermediate result, where and are learnable scaling and offset parameters. Then, through the split operation, is split into a query vector and a key vector according to certain rules. These learnable parameters enable the model to adaptively adjust and in the generation manner.

[0101] Next, calculate the attention weight matrix: Using the query vector and the key vector , through the formula calculate the attention weight matrix . Among them, is the dimension of the key vector , represents and the product of the transposes, is the rectified linear unit function, used to introduce non-linearity. The generated attention weight matrix reflects the attention distribution among positions. The function is used to ensure that all attention weights are normalized to a probability distribution and sum to 1. Specifically, for the input vector ,

[0102] This step dynamically adjusts the interaction weights through learnable parameters, enabling the model to more flexibly focus on important parts of the input features, enhancing the model's expressive ability and the ability to capture key information.

[0103] S1033. Weighted aggregation of the value vector based on the attention weight matrix to generate preliminary attention features.

[0104] Perform a linear transformation on the input set of position-enhanced features X again, through the formula to obtain the value vector and the control vector of the gating mechanism. Among them, is the learnable weight matrix, is the bias vector, and the split operation splits the transformed result into and according to preset rules. The value vector contains the information of the input features, and the gating vector is used to control the flow of information.

[0105] Using the attention weight matrix obtained from the previous calculation , perform weighted aggregation on the value vector , and combine it with the gating vector . Through the formula , calculate the preliminary attention feature. Among them, represents matrix multiplication, represents the Hadamard product (element-wise multiplication). This process weights the value vector according to the attention weight matrix and adjusts the result using the gating mechanism to generate the preliminary attention feature.

[0106] S1034. Perform layer scaling on the preliminary attention feature, and adjust the feature distribution of the preliminary attention feature through a trainable scaling factor to generate the scaled attention feature.

[0107] First, perform a normalization operation on the preliminary attention feature using the formula where, is the preliminary attention feature, is the mean of is the variance of is a very small constant used to ensure that the denominator is not zero; and are learnable scaling and offset parameters. The purpose of normalization is to make the feature distribution more stable, which is helpful for the training of the model.

[0108] Perform layer scaling through the formula . Among them, is a trainable scaling factor, and its initial value is set to a small constant; is the output of the gated attention unit (i.e., the preliminary attention feature). Layer scaling can adjust the feature distribution of the preliminary attention feature and enhance the stability of model training. After this step of processing, the scaled attention feature is generated.

[0109] The layer scaling mechanism can stabilize the model training process, prevent the training instability problem caused by the feature distribution being too large or too small, and at the same time improve the feature expression ability, which is helpful for enhancing the generalization ability of the model.

[0110] S1035. Combine the scaled attention feature and the original input feature through a residual connection, and generate a deep temporal feature set based on the activation function and normalization processing.

[0111] The scaled attention features are residually connected to the original input features, that is, the two are added together. The role of residual connection is to help the model better learn the differences and connections between features without adding too much computational complexity. At the same time, it also helps to alleviate the vanishing gradient problem, enabling the model to be trained deeper.

[0112] The merged features are processed using the Gaussian Error Linear Unit (GELU) activation function. The GELU activation function can introduce non-linearity according to the probability distribution of the input. Compared with traditional activation functions such as ReLU, it can better capture complex patterns in the data, making the model have stronger expressive power.

[0113] Finally, the features processed by the activation function are normalized again using LayerNorm to further stabilize the feature distribution and make the training of the model more stable. Through the above series of operations, a deep temporal feature set is generated.

[0114] Among them, the residual connection can maintain information flow and prevent the vanishing gradient problem in deep neural networks; the activation function and normalization processing can enhance the non-linear expressive power of features, improving the generalization ability and stability of the model.

[0115] The method provided by the embodiments of this application generates a position-enhanced feature set by embedding sine-cosine position encoding information in the shallow temporal feature set. The position-enhanced feature set is input into the gated attention unit, and the interaction weight between the query vector and the key vector is dynamically adjusted through learnable parameters to generate an attention weight matrix. Based on the attention weight matrix, the value vectors are weighted and aggregated to generate preliminary attention features. The preliminary attention features are subjected to layer scaling processing, and the feature distribution of the preliminary attention features is adjusted through a trainable scaling factor to generate scaled attention features. The scaled attention features are merged with the original input features through residual connection, and based on the activation function and normalization processing, a deep temporal feature set is generated, thereby capturing more complex temporal dependencies and potential abnormal patterns, and providing a highly abstract temporal representation for prediction in subsequent steps.

[0116] Optionally, before inputting the normalized input data stream into the multi-level feature extraction network, the normalized input data stream corresponding to the multi-scenario flight states can also be obtained based on the original data corresponding to the multi-scenario flight states, where the multi-scenario flight states include at least two of stable flight, high-speed maneuvering, and strong wind interference. For example, in the stable flight state, the original data is collected when the flight attitude of the UAV is stable and the speed change is small; in the high-speed maneuvering state, the original data is collected when the UAV quickly changes its flight speed and direction; in the strong wind interference state, the original data is collected when the UAV is affected by strong wind. Then, the original data of these different scenarios is processed according to the multi-level data preprocessing method in the aforementioned step S101, so as to obtain the normalized input data stream corresponding to each multi-scenario flight state.

[0117] Configure the adjustment weight of the normalized input data stream corresponding to each scenario flight state according to the recognition difficulty of each scenario flight state in the multi-scenario flight states.

[0118] The multi-scenario flight state weight adjustment algorithm can be expressed as: Wherein, It can be calculated in the following way: represents the historical recognition error rate of this scenario, represents the complexity of the data of this scenario, which can be calculated by the entropy value or variance of the data, 、 are adjustable parameters used to balance the influence of the two factors Specifically, analyze the characteristics of the data under different multi-scenario flight states and the difficulty in the model training and recognition process. For example, in the strong wind interference scenario, due to the uncertainty and complexity of the wind, the attitude and sensor data of the UAV may fluctuate greatly, increasing the difficulty for the model to recognize abnormal and normal states; while in the stable flight state, the data is relatively stable and the recognition difficulty may be lower. According to these recognition difficulties, configure the adjustment weight for the normalized input data stream corresponding to each scenario flight state. The higher the recognition difficulty of the scenario, the higher the allocated adjustment weight; the lower the recognition difficulty of the scenario, the lower the allocated adjustment weight. For example, if the recognition difficulty of the strong wind interference scenario is high, its adjustment weight can be set to 0.8; if the recognition difficulty of the stable flight scenario is low, the adjustment weight can be set to 0.2.

[0119] Adjust the normalized input data stream input into the multi-level feature extraction network according to the adjustment weight.

[0120] S1041. Concatenate the shallow temporal feature set and the deep temporal feature set in the channel dimension to generate a joint feature tensor.

[0121] The shallow temporal feature set contains shallow features extracted from the original data, retaining more local temporal detail information; the deep temporal feature set contains complex spatio-temporal correlation features. These two feature sets are concatenated in the channel dimension, that is, their channels are connected. For example, if the shape of the shallow temporal feature set is (batch_size, sequence_length, num_channels1) and the shape of the deep temporal feature set is (batch_size, sequence_length, num_channels2), the shape of the resulting combined feature tensor after concatenation is (batch_size, sequence_length, num_channels1 + num_channels2).

[0122] In this step, different feature information from the shallow and deep layers is integrated through concatenation in the channel dimension, providing a rich feature basis for subsequent fusion operations.

[0123] S1042. Perform a linear transformation on the combined feature tensor to compress the feature dimension and generate a gating weight vector.

[0124] In S1041, the combined feature tensor has been obtained, and its shape is assumed to be (batch_size, sequence_length, num_channels), where batch_size represents the number of samples in one input data, sequence_length represents the length of the temporal sequence in each sample, and num_channels represents the number of channels of the features.

[0125] To generate the gating weight vector, a linear layer (i.e., a fully connected layer) needs to be applied to this combined feature tensor. The core components of the linear layer are the learnable weight matrix W and the bias vector b.

[0126] The dimension of the weight matrix W is (num_channels, output_dim), where num_channels is the channel dimension of the combined feature tensor and output_dim is a pre-set output dimension, which determines the final dimension size of the gating weight vector. The dimension of the bias vector b is a one-dimensional vector with a length of output_dim.

[0127] The calculation process of the linear transformation is as follows: The joint feature tensor is unfolded along the channel dimension so that it can be regarded as a two-dimensional matrix (if batch_size and sequence_length are regarded as the overall sample number dimension). For each "sample" in the joint feature tensor (i.e., a sequence of feature vectors of length sequence_length and number of channels num_channels), the following calculations are performed: Let this "sample" be X (with shape (sequence_length, num_channels)), and the output after linear transformation be Y (with shape (sequence_length, output_dim)). Then the calculation formula is . In actual calculation, for each element in the channel dimension of the joint feature tensor, it is multiplied and summed with the corresponding column elements of the weight matrix , and then added with the corresponding element of the bias vector to obtain the transformed result.

[0128] Through the above linear transformation, the channel dimension of the joint feature tensor is compressed from num_channels to output_dim, generating a gating weight vector with shape (batch_size, sequence_length, output_dim). This gating weight vector will be used for subsequent weighted fusion operations on shallow features and deep features to determine the weight distribution of different features in the fusion process.

[0129] This step compresses the feature dimension through linear transformation, reduces the number of parameters, reduces the computational overhead, and at the same time generates a gating weight vector for subsequent weighted fusion.

[0130] S1043. Use the gating weight vector to perform weighted fusion on the shallow features in the shallow temporal feature set and the deep features in the deep temporal feature set respectively to obtain preliminary fusion features.

[0131] For each shallow feature in the shallow temporal feature set and each deep feature in the deep temporal feature set, perform weighted operations with the gating weight vector respectively.

[0132] Specifically, the gated weight vector is multiplied with the shallow features and the deep features in a certain way (for example, element-wise multiplication in the channel dimension, etc.), and then the weighted shallow features and deep features are added together to obtain the preliminary fusion features. Suppose the shape of the shallow features is (batch_size, sequence_length, num_channels1), the shape of the deep features is (batch_size, sequence_length, num_channels2), and the shape of the gated weight vector is (batch_size, sequence_length, output_dim). After the weighted fusion operation, the shape of the preliminary fusion features is the same as that of the shallow features and the deep features (for example, (batch_size, sequence_length, num_channels1) or (batch_size, sequence_length, num_channels2), which specifically depends on the implementation of the weighted fusion).

[0133] In this step, through weighted fusion, the contribution degrees of the shallow features and the deep features are automatically adjusted according to the gated weight vector, effectively combining the local temporal details of the shallow layer and the complex spatio-temporal correlations of the deep layer.

[0134] S1044. Apply a channel-level masking operation to the preliminary fusion features, activate the effective feature channels through a non-linear function, and generate a set of fused spatio-temporal correlation features.

[0135] Among them, the channel-level masking operation adopts a dynamic threshold mechanism to automatically select the activation threshold according to the importance scores of the feature channels.

[0136] First, calculate the importance scores of each channel in the preliminary fusion features. The specific calculation method can be based on statistics such as the variance and mean of the features, or calculated through an additional network module. Then, according to the calculated importance scores, adopt a dynamic threshold mechanism to automatically select the activation threshold. For the feature values of each channel, input them into a non-linear function (such as the Tanh function) for activation operation. If the activated feature value is greater than the automatically selected activation threshold, then the channel is considered an effective channel and its feature value is retained; otherwise, the feature value of the channel is processed accordingly (such as set to 0). Through such a channel-level masking operation, a set of fused spatio-temporal correlation features is finally generated.

[0137] In this step, the channel-level masking operation combines the dynamic threshold mechanism and the non-linear activation function, which can effectively filter out the effective temporal information, remove the unimportant channel features, further optimize the feature representation, and improve the model's ability to process spatio-temporal correlation information.

[0138] S1051. Input the fused spatio-temporal correlation feature set into the flatten layer of the multi-layer perceptron prediction head to convert the fused spatio-temporal correlation feature set into a two-dimensional tensor feature set.

[0139] The data structure of the fused spatio-temporal correlation feature set is usually a three-dimensional tensor with a shape of (batch_size, sequence_length, num_channels). Here, batch_size represents the number of samples input in one training or inference process, which can be understood as multiple groups of UAV flight data instances processed simultaneously; sequence_length represents the length of the time series, that is, the number of time steps for the UAV to continuously collect data within a period of time, reflecting the time dimension information of the data; num_channels is the number of feature channels, which contains various types of feature information such as the position, speed, and attitude angle of the UAV.

[0140] The role of the flatten layer is to perform a structural transformation on the three-dimensional tensor. While keeping the batch_size dimension unchanged, the flatten layer combines the sequence_length and num_channels dimensions. Specifically, it flattens the structure where there are num_channels features at each of the sequence_length time steps into a one-dimensional vector, and then forms a two-dimensional tensor with batch_size such one-dimensional vectors. After being processed by the flatten layer, the shape of the output two-dimensional tensor feature set becomes (batch_size, sequence_length * num_channels). This transformation is to make the data format meet the input requirements of the subsequent multi-layer fully connected network, because the fully connected network usually processes data in the form of two-dimensional matrices.

[0141] S1052. Input the two-dimensional tensor feature set into the multi-layer fully connected network of the multi-layer perceptron prediction head for linear transformation processing of different dimensions to obtain the feature tensor to be mapped.

[0142] The multi-layer fully connected network is composed of multiple fully connected layers connected in sequence. When the two-dimensional tensor feature set is input into the first fully connected layer, this layer contains a set of learnable weight matrices W and bias vectors b. The dimension of the weight matrix W is (input_dim, output_dim), where input_dim is the dimension of the input feature vector, that is, the second dimension of the two-dimensional tensor after flattening in the previous step (sequence_length * num_channels); output_dim is the dimension of the feature vector that this layer hopes to output, which is preset according to the model design.

[0143] During the calculation process of the fully connected layer, for each input feature vector (with dimension input_dim), the weighted sum is calculated through matrix multiplication , and then the bias vector is added to obtain the output of this layer , with dimension output_dim. In this way, a linear transformation is completed, mapping the input feature vector to a new dimensional space.

[0144] To introduce non-linearity and enable the model to learn more complex functional relationships, after the linear transformation of each fully connected layer, an activation function can be connected. Common activation functions such as ReLU have the following calculation formula , that is, when the input is greater than 0, the output is ; when is less than or equal to 0, the output is 0. The activation function can perform non-linear conversion on the result of the linear transformation, enhancing the expression ability of the model.

[0145] In addition, to prevent the model from overfitting during training, a dropout layer is also connected after the activation function. The dropout layer will randomly set some elements in the input feature vector to 0 with a certain probability (e.g., 0.2), which can reduce the complex co-adaptation relationship between neurons and improve the generalization ability of the model.

[0146] After being processed by multiple such fully connected layers, activation functions, and dropout layers in sequence, the two-dimensional tensor feature set undergoes multiple linear transformations, non-linear conversions, and overfitting suppression, and finally the feature tensor to be mapped is obtained.

[0147] S1053: Map the feature tensor to be mapped to the target dimensional space to obtain the pose parameter prediction result.

[0148] After being processed by the multi-layer fully connected network, the feature tensor to be mapped enters the output layer of the multi-layer perceptron prediction head. The output layer is also a fully connected layer, whose role is to map the feature tensor to be mapped from the current feature dimensional space to the target dimensional space, and this target dimensional space corresponds to the dimension of the key pose parameters of the UAV.

[0149] Assume that the key attitude parameters of the drone include pitch angle, roll angle, yaw angle, etc. Then the dimension of the target dimensional space corresponds to the number of these attitude parameters. The output layer contains a weight matrix \(W_{out}\) of a specific dimension and a bias vector \(b_{out}\). The dimension of the weight matrix \(W_{out}\) is \((current\_dim, num\_targets)\), where \(current\_dim\) is the dimension of the feature tensor to be mapped, and \(num\_targets\) is the dimension of the target dimensional space, that is, the number of the drone's key attitude parameters.

[0150] Through matrix multiplication and addition operations, the feature tensor to be mapped is multiplied by the weight matrix \(W_{out}\) and added with the bias vector \(b_{out}\) to obtain the final output result, that is, the attitude parameter prediction result. This prediction result contains the predicted values of the model for each key attitude parameter of the drone, providing a basis for subsequent analysis and decision-making.

[0151] S1054. Obtain a set of deviation values based on the attitude parameter prediction result and reference data, where the reference data is obtained based on the real sensor data of the drone.

[0152] The attitude parameter prediction result is the estimated value of the drone's attitude parameters by the model, while the reference data is the actual attitude parameter values collected by the real sensors installed on the drone. Since the data collected by the sensors reflects the real state of the drone during actual flight, it is used as a reference standard.

[0153] For each attitude parameter prediction value in the attitude parameter prediction result, it is compared with the corresponding actual attitude parameter value in the reference data. The specific comparison method is to do a difference operation, that is, deviation value = attitude parameter prediction value - actual attitude parameter value. By performing such difference calculations on all attitude parameters, a set of deviation values is obtained. Summing up these deviation values forms a set of deviation values. This set of deviation values reflects the difference between the model prediction result and the actual situation, and subsequent anomaly detection will be based on these deviation values for judgment.

[0154] S1055. Statistically calculate the deviation mean and deviation standard deviation of each attitude parameter based on the historical validation set data, and determine the initial threshold using the normal distribution.

[0155] The historical validation set data is accumulated during the model training and validation process, which contains the deviation data between the prediction results and the true values of the model on the validation set in previous times. For each attitude parameter of the drone (such as pitch angle, roll angle, yaw angle, etc.), statistical analysis is performed on all the deviation values of it on the historical validation set.

[0156] The method for calculating the mean deviation is to sum up all the deviation values of the attitude parameters in the historical validation set and then divide by the number of deviation values, i.e., mean deviation = sum of all deviation values / number of deviation values. The calculation of the standard deviation of the deviation is to first calculate the square of the difference between each deviation value and the mean deviation, sum up these squared values and divide by the number of deviation values, and then take the square root of the result.

[0157] According to the 3σ statistical principle of the normal distribution, in a normal distribution, approximately 99.7% of the data will fall within the range of the mean plus or minus 3 times the standard deviation. Therefore, the initial threshold is set as the calculated mean deviation plus or minus 3 times the standard deviation of the deviation. This initial threshold can theoretically cover the distribution of most normal data and provide a preliminary judgment criterion for subsequent anomaly detection.

[0158] S1056. Dynamically adjust the weight coefficient of the initial threshold according to the current flight phase of the unmanned aerial vehicle (UAV) to generate an adaptive threshold set.

[0159] Among them, the calculation of the adaptive threshold can use the following formula: Among them, is the initial threshold calculated based on the historical validation set data, that is is the adjustment coefficient calculated according to the current flight phase. For example, it can be determined through multiple adjustable weight parameters according to the importance of the influence of various factors on the threshold (such as indicators such as flight speed, environmental interference intensity, and sensor confidence). These parameters can determine the optimal values through historical data analysis. The adjustment coefficients for each attitude parameter in the adaptive threshold set are calculated independently and are related to the indicators of flight speed, environmental interference intensity, and sensor confidence.

[0160] Exemplarily, the flight phases of the UAV can be divided into takeoff phase, cruise phase, hover phase, maneuver phase, and landing phase. The flight phase identification is based on the following rules: 1. Takeoff phase: The vertical speed is greater than the threshold T_v_up, and the altitude continues to increase.

[0161] 2. Cruise phase: The horizontal speed is greater than the threshold T_v_cruise, and the altitude change is less than the threshold T_h_stable.

[0162] 3. Hover phase: Both the horizontal speed and the vertical speed are less than the threshold T_v_hover.

[0163] 4. Maneuver phase: The angular velocity is greater than the threshold T_ω, or the acceleration is greater than the threshold T_a.

[0164] 5. Landing phase: The vertical speed is less than the threshold T_v_down (negative value), and the altitude continues to decrease.

[0165] Different adaptive threshold adjustment strategies are adopted for different flight phases. For example, in the maneuvering phase, due to the drastic attitude changes, a larger threshold adjustment coefficient is used; while in the hovering phase, the attitude is relatively stable, and a smaller threshold adjustment coefficient is used. Specifically, the flight speed usually changes significantly during the takeoff and landing phases, while it is relatively stable during the cruise phase; the intensity of environmental interference also varies at different flight altitudes and regions. For example, there may be interference from buildings when flying at low altitudes, and relatively less interference during high-altitude cruising; the sensor confidence index is also affected by various factors. For example, the accuracy of the sensor may decrease in complex environments.

[0166] For each attitude parameter in the adaptive threshold set, the adjustment coefficient is calculated independently. The calculation of the adjustment coefficient comprehensively considers factors such as the above-mentioned flight speed, environmental interference intensity, and sensor confidence index. The relationship between the adjustment coefficient and these factors can be determined by establishing a corresponding mathematical model or using machine learning methods. For example, a function can be designed that takes the flight speed, environmental interference intensity, and sensor confidence index as inputs and calculates an adjustment coefficient through the function.

[0167] According to the calculated adjustment coefficient, the weight of the initial threshold is dynamically adjusted. For example, if the calculated adjustment coefficient is large in a certain flight phase, it means that a large adjustment of the initial threshold is required to adapt to the characteristics of this phase; conversely, if the adjustment coefficient is small, the adjustment amplitude of the initial threshold is also small. In this way, a threshold adapted to its current flight phase is generated for each attitude parameter, and finally an adaptive threshold set is formed.

[0168] S1057. Obtain the abnormal detection decision result according to the deviation value set and the comparison result of the adaptive threshold.

[0169] Among them, the abnormal detection decision result includes multi-level abnormal alarms. The multi-level abnormal alarms include single-point instantaneous alarms, continuous abnormal alarms, and trend deviation alarms. Different-level alarms correspond to different flight control system response strategies.

[0170] Compare each deviation value in the deviation value set with the corresponding attitude parameter threshold in the adaptive threshold set. In order to comprehensively detect different types of abnormal situations, a multi-level abnormal judgment mechanism is designed.

[0171] Exemplarily, the specific algorithm of the multi-level abnormal judgment mechanism is as follows: 1. Single-point instantaneous alarm judgment: If and and , a single - point instantaneous alarm is triggered. Among them, is the deviation value at the

[0172] 2. Continuous abnormal alarm judgment: If the deviation values at consecutive time points all exceed the adaptive threshold, that is: are all greater than , a continuous abnormal alarm is triggered. Among them, is the preset window size for continuous abnormal judgment, which can be set to 3 - 5, for example.

[0173] 3. Trend deviation alarm judgment: The slope of the deviation values within the window can be calculated: If , a trend deviation alarm is triggered. Among them, is the window size for trend analysis, which can be set to 8 - 10, for example.

[0174] When a certain deviation value exceeds the corresponding adaptive threshold and this exceeding occurs instantaneously at a single point, that is, the deviation value exceeds the threshold only at a certain time point, a single - point instantaneous alarm is triggered at this time. This kind of alarm usually indicates that there may be a short - term and accidental abnormal situation, such as momentary noise interference of the sensor.

[0175] When the deviation values at consecutive multiple time points all exceed the adaptive threshold, a continuous abnormal alarm is triggered. The continuous abnormal alarm indicates that the attitude parameters of the UAV continuously deviate from the normal range within a period of time, and there may be relatively serious problems, such as system failures or continuous external interference, etc.

[0176] When the deviation values show a certain trend of deviation, such as gradually increasing or decreasing, and exceed a certain range, a trend deviation alarm is triggered. The trend deviation alarm can detect potential problems that may occur in the attitude parameters of the UAV in advance, which helps to take measures in time for adjustment and prevention.

[0177] Output different abnormal warning levels and corresponding recommended operations according to different alarm types and severities. For example, for single-point instantaneous alarms, a relatively low warning level may be output, and it is recommended to continue observing; for continuous abnormal alarms and trend deviation alarms, a relatively high warning level may be output, and operations such as adjusting the flight attitude, reducing the flight speed, or performing a system check may be recommended. These results constitute the abnormal detection decision results, providing key judgment basis for the flight control system, and can help the flight control system make decisions in a timely manner to ensure the navigation safety of the UAV.

[0178] In a possible implementation manner, the method may further include: Establish an abnormal mode feature library, where the abnormal modes in the abnormal mode feature library include abnormal modes of sensor hardware failures, signal interference, and environmental impacts. Specifically, a large amount of data related to UAV flight can be collected and sorted out, including but not limited to the output data of sensors when hardware failures occur, data under signal interference, and data under different environmental impacts. Analyze and process these data, and extract the characteristic patterns that can characterize sensor hardware failures, signal interference, and environmental impacts. For example, for sensor hardware failures, specific numerical abnormal change patterns, data fluctuation rules, etc. may be extracted; for signal interference, characteristic changes in aspects such as frequency and intensity may be analyzed; for environmental impacts, the correlation characteristics between environmental factors such as temperature, air pressure, and wind speed and sensor data may be considered. Organize and store these extracted characteristic patterns to establish an abnormal mode feature library.

[0179] By matching the abnormal data in the abnormal detection decision results with each abnormal pattern in the abnormal pattern feature library, a set of fault causes is generated. The similarity matching is obtained by using the dynamic time warping algorithm to align time series of different lengths and calculating the cosine similarity score. Specifically, when matching the abnormal data in the abnormal detection decision results with the abnormal patterns in the abnormal pattern feature library, since different time series may have different lengths, the dynamic time warping algorithm is used to align these time series. The dynamic time warping algorithm finds the optimal matching path between two time series, enabling them to be aligned as much as possible on the time axis. For example, a distance matrix can be constructed, where each element in the matrix represents the distance between corresponding points of two time series. Then, through the method of dynamic programming, a path with the minimum cumulative distance is found in this matrix to achieve the alignment of the time series. After using the dynamic time warping algorithm to align the time series, for the aligned abnormal data sequence and the abnormal pattern sequence in the feature library, calculate the cosine similarity score between them. The cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. The closer the value is to 1, the more similar the two sequences are. By setting a similarity threshold, the fault causes corresponding to the abnormal patterns with similarity scores exceeding the threshold are included in the set of fault causes.

[0180] Based on the attention weight matrix output by the gated attention Transformer module based on layer scaling, analyze the abnormal propagation paths of each fault cause in the set of fault causes to determine the root sensor node. The analysis of the abnormal propagation path is achieved through visualizing the attention heatmap. The specific steps are as follows: 1. Extract the attention weight matrix from the LSGA-T module; 2. Normalize the weight matrix so that the value range is within [0, 1]; 3. Use a heatmap color mapping function (such as "jet") to map the normalized weight values to colors; 4. Generate a heatmap, where the color depth represents the correlation strength between different sensor nodes; 5. Identify the high-weight connection paths in the heatmap (such as dark regions); 6. Analyze the starting point and propagation direction of the high-weight connections to determine the root sensor node. Among them, the high-weight connection paths in the abnormal propagation path point to potential fault sources.

[0181] Specifically, the attention weight matrix reflects the correlation degree between features at different positions. For each fault cause in the set of fault causes, this attention weight matrix can be used to analyze the abnormal propagation path. By visualizing the attention weight matrix, an attention heatmap is generated. In the heatmap, the color depth represents the weight size. The high-weight connection paths are the darker connection parts, and the sensor nodes pointed to by these paths are considered potential fault sources. Through the analysis of the heatmap, gradually trace the abnormal propagation path to determine the root sensor node.

[0182] Combine the flight status parameters of the UAV with the historical maintenance records to calculate the probability distribution of each fault cause in the fault cause set. Specifically, the flight status parameters of the UAV can be collected, such as flight altitude, speed, attitude angle, flight phase, etc., as well as the historical maintenance records, including information such as the types of faults that have occurred in the past, the time of fault occurrence, and maintenance measures. For each fault cause in the fault cause set, analyze the association between these flight status parameters and historical maintenance records and this fault cause. For example, certain faults may be more likely to occur at specific flight altitudes or flight phases, or certain faults have a higher frequency of occurrence in the past maintenance records. Through statistical analysis and machine learning algorithms (such as Bayesian classifiers, etc.), calculate the probability of each fault cause based on this association information, so as to obtain the probability distribution of each fault cause in the fault cause set.

[0183] Generate hierarchical warning information and maintenance suggestions according to the probability distribution. Specifically, different probability thresholds can be set to divide the warning levels according to the calculated probability distribution of the fault causes. For example, when the probability of a certain fault cause exceeds 0.8, a first-level warning is issued; when it is between 0.5 and 0.8, a second-level warning is issued, etc. For different warning levels, combine the characteristics of the fault causes and historical maintenance experience to generate corresponding maintenance suggestions. A first-level warning may recommend immediately stopping the flight and conducting a comprehensive fault investigation and repair; a second-level warning may recommend conducting a detailed inspection after the flight ends, etc. Provide this hierarchical warning information and maintenance suggestions to relevant personnel so as to take timely measures to ensure the safe operation of the UAV.

[0184] In a possible implementation manner, this method may further include: A trend analysis algorithm can be developed based on the normalized input data stream. By learning and pattern recognition of the historical normalized input data stream, capture the change trend of the data in the normalized input data stream over time. For example, observe indicators such as the change slope and fluctuation range of the UAV altitude sensor data over a period of time, judge whether there is an abnormal change trend, and then predict possible abnormal situations in the future. For the short-term data change characteristics, focus on analyzing the data with a short time step within the sliding window to capture information such as the rapid change of the attitude angle when the UAV is instantaneously disturbed by air flow. Capturing the medium-term data change characteristics focuses on the data trend over a relatively long period of time, such as the stable change of the flight speed of the UAV during a cruise phase. The analysis of the long-term data change characteristics focuses on a more macroscopic time range.

[0185] Specifically, different structures and parameter settings can be used to adapt to feature extraction at different time scales. For short-term feature extraction, a smaller sliding window and a simpler network layer structure may be used to quickly process and capture local detail information; for medium-term and long-term feature extraction, a larger sliding window, a deeper network layer or a special time series processing module may be used to mine more complex trends and patterns.

[0186] Then, the severity of the anomaly is assessed based on its potential impact on the safety and performance of the drone. For example, if an abnormality is predicted in the drone’s attitude angle sensor, it may cause loss of flight control, and its severity is high; while abnormalities in some secondary sensors may have less impact on flight, and their severity is correspondingly low. The probability is based on the output of the trend analysis algorithm and the multi-time scale feature extraction network, combined with historical data and statistical analysis to estimate the probability of anomaly occurrence.

[0187] Afterwards, these factors are combined through the fusion rules determined according to actual needs to calculate the comprehensive risk score. For example, different abnormality types, severity and likelihood can be given corresponding weights, and then the comprehensive risk score can be obtained by weighted summation.

[0188] Next, set different risk score thresholds to divide the warning level, such as low risk (risk score 0-0.3), medium risk (risk score 0.3-0.6), and high risk (risk score greater than 0.6). When the risk score is in the low risk range, the first-level warning is triggered, and only continuous monitoring and data recording are carried out; if it is in the medium risk range, the second-level warning is triggered, prompting the operator to pay attention and prepare to take preventive measures; if it reaches the high risk range, the third-level warning is triggered, and an alarm is immediately issued and emergency response measures are recommended. According to the abnormal root cause determined in step 12 (such as sensor hardware failure, signal interference, environmental impact, etc.), combined with the warning level, detailed abnormal information and potential impact descriptions are provided to the operator.

[0189] Finally, according to different warning levels and root cause analysis results, formulate corresponding preventive operation suggestions. If the warning is caused by the risk of sensor hardware failure, it is recommended to self-check, calibrate or replace the relevant sensors; if it is a signal interference risk, it is recommended to adjust the flight route to avoid the interference source, or switch to the backup signal channel; if the risk is caused by environmental influences (such as bad weather), it is recommended to change the flight altitude, speed, or postpone the flight mission. These suggestions are communicated to the flight control system in a timely manner, so that it can automatically or assist the operator to take measures in advance to reduce the possibility of abnormalities or mitigate the impact caused by abnormalities, and realize the transformation from passive response to abnormalities to active prevention of abnormalities.

[0190] In a possible implementation, the method may further include: The adaptive threshold is dynamically adjusted according to the changes in the UAV's flight environment, the current flight status, the impact of historical flight data on the attitude parameter prediction results, and the prediction accuracy of the attitude parameter prediction results.

[0191] Among them, the flight environment of drones is complex and diverse, including different weather conditions (such as sunny days, rainy days, windy days, etc.) and different geographical environments (such as cities, mountainous areas, plains, etc.). For example, in windy weather, drones are subject to greater airflow interference, and the fluctuation of attitude parameters will be more obvious; when flying in mountainous areas, the terrain may affect the GPS signal and cause instability. These environmental factors will affect the actual flight state of the drone, and thus affect the accuracy and reliability of the attitude parameter prediction results. Therefore, it is necessary to monitor the changes in the flight environment in real time, such as obtaining information such as wind speed, wind direction, precipitation, etc. through meteorological sensors, and understanding the terrain characteristics of the flight area through satellite positioning and terrain data.

[0192] The current flight status covers a variety of dynamic information of the drone, such as flight speed, flight altitude, attitude angle (pitch angle, roll angle, yaw angle) and flight mode (steady flight, high-speed maneuver, sharp turn, hovering, etc.). In different flight states, the attitude parameters of the drone show different characteristics. For example, in high-speed maneuvering flight, the acceleration and attitude of the drone change more dramatically; while in the hovering state, the attitude parameters are relatively stable. Real-time acquisition of the current flight status data of the drone helps to accurately determine the normal range and change trend of the attitude parameters.

[0193] Historical flight data records the flight information of the drone in various situations in the past and the corresponding attitude parameter prediction results. By analyzing historical data, some rules and patterns can be found. For example, in certain specific flight environments and flight state combinations, attitude parameter predictions may often have large deviations; or flight data in certain time periods may show specific trends, which can provide references for current attitude parameter predictions. Statistical analysis methods can be used, such as calculating statistics such as the mean and standard deviation under different conditions, to quantify the impact of historical data on prediction results.

[0194] Prediction accuracy can be measured by a variety of indicators, such as mean square error, mean absolute error, etc. By calculating these indicators, we can understand the accuracy of the current model's prediction of posture parameters. If the prediction accuracy is low, it means that the model's prediction results deviate greatly from the actual situation, and it may be necessary to adjust the adaptive threshold to more accurately judge abnormal situations.

[0195] Taking into account the above factors, a dynamic adjustment model or algorithm can be established. This model can adjust the adaptive threshold using machine learning or rule-based methods based on real-time information such as changes in the flight environment, the current flight state, historical flight data, and the prediction accuracy index of the attitude parameter prediction results. For example, a regression model can be used to fit the relationship between these factors and the adaptive threshold, and the appropriate threshold adjustment amount can be predicted based on the input real-time data; or some rules can be set, and when certain factors meet specific conditions, the threshold can be adjusted according to the predetermined rules. Through this dynamic adjustment mechanism, the adaptive threshold can better adapt to different flight conditions, improving the accuracy and reliability of anomaly detection.

[0196] In a possible implementation manner, this method may further include: Based on the historical prediction results of the normalized input data stream of the adjusted input multi-level feature extraction network, obtain the typical scenario features corresponding to the historical prediction results during the prediction process. For example, the historical prediction results of the normalized input data stream of the adjusted input multi-level feature extraction network, as well as various relevant data during the prediction process, including but not limited to flight environment data, flight state data, sensor data, etc., can be collected. The collected data is deeply analyzed to extract typical features that can represent different scenarios. These features can be numerical, such as specific ranges of flight speed and altitude; or categorical, such as the type of flight environment (sunny, rainy, etc.), flight mode (steady flight, high-speed maneuvering, etc.). Through methods such as clustering analysis and principal component analysis, find the representative feature combinations in different scenarios, and these feature combinations are the typical scenario features. For example, in the scenario of high-speed maneuvering and strong wind interference, the typical scenario features may include the flight speed being greater than a certain threshold, the wind speed exceeding a certain value, and a relatively large change rate of the attitude angle, etc.

[0197] Based on the current scenario flight state of the unmanned aerial vehicle, determine the typical scenario features corresponding to the current scenario flight state. For example, the current scenario flight state information of the unmanned aerial vehicle can be obtained in real time, including the current flight environment (such as the current weather condition, geographical area where it is located, etc.), flight mode (steady flight, high-speed maneuvering, etc.), and other relevant state parameters (such as speed, altitude, attitude angle, etc.). Match and analyze these real-time state information with the previously extracted typical scenario features to determine the typical scenario features corresponding to the current scenario flight state. For example, if the current unmanned aerial vehicle is in an environment of high-speed flight and strong wind, by comparing with the historical typical scenario features, it can be determined that the corresponding typical scenario features are the feature combination in the scenario of high-speed maneuvering and strong wind interference.

[0198] When predicting the prediction results of the attitude parameters of an unmanned aerial vehicle (UAV), typical scenario features are pre-input for prediction. For example, when predicting the attitude parameters of a UAV, the typical scenario features corresponding to the determined current scenario flight state are used as additional input information and input into a prediction model (such as the multi-layer perceptron prediction head mentioned in the foregoing embodiments) together with other original input data. In this way, when the model makes a prediction, it can make full use of the information contained in these typical scenario features, better capture the variation rules of attitude parameters in different scenarios, and thus improve the accuracy and reliability of the attitude parameter prediction results. For example, after the model learns the relationship between the typical features and attitude parameters in scenarios of high-speed maneuvering and strong wind interference, when encountering a similar scenario again, it can more accurately predict the attitude parameters based on the pre-input typical scenario features.

[0199] Optionally, when using the method described above for prediction, the prediction efficiency can be further improved by one or more of the following methods.

[0200] An incremental inference method can be designed based on the sliding window structure described above, and only the data newly entering the window is processed each time. Among them, the traditional inference method may process the data in the entire window each time, but the incremental inference method focuses on the data newly entering the window. When new data arrives, only the inference calculation is performed on the data newly entering the sliding window, without reprocessing the old data that has been processed in the window. For example, assume that the sliding window size is 10 time steps. When the data at the 11th time step enters the window, the incremental inference method only analyzes and infers the data at the 11th time step, and uses the relevant information (such as feature representations) that has been calculated from the data in the previous window to quickly obtain the prediction results related to the new data. This can reduce the amount of calculation, especially in scenarios where data flows continuously, avoiding ineffective calculations on a large amount of duplicate data, and thus improving the speed and efficiency of inference.

[0201] The repeated calculation of processed data features can be avoided through a feature caching mechanism. After the feature extraction is performed for the first time, the obtained feature results are stored in the cache. If the same or partially the same data is encountered again later (in practical applications, this situation is relatively common due to the continuity and repeatability of data), the features calculated previously are directly read from the cache instead of performing complex feature extraction calculations again. For example, during the flight of a UAV, the flight state is relatively stable in some time periods, and the corresponding input data features are also relatively similar. When this similar data enters the system again, the existing feature representations can be quickly obtained through the feature caching mechanism, reducing repeated calculations, saving computing resources and time, and thus improving the overall prediction efficiency.

[0202] Model quantization techniques can be applied to convert model parameters from high-precision floating-point numbers to low-precision integers, thereby reducing computational and storage requirements.

[0203] Inference pipeline parallelization can be achieved to perform data preprocessing, feature extraction, and anomaly detection simultaneously, maximizing the utilization of hardware resources and enabling the entire system to run in real time on drones, providing immediate anomaly detection and early warning.

[0204] Optionally, a model evaluation method and a performance metric analysis process can also be designed to provide guidance for improving the prediction model. For example, the mean squared error and mean absolute error can be used to measure the error between the predicted value and the true value of the multi-layer perceptron prediction head, determining the ability to explain data variation. A smaller error and a coefficient of determination close to 1 indicate better performance. The confusion matrix is used to present the distribution of anomaly detection decision results, from which the true positive rate, false positive rate, true negative rate, and false negative rate are calculated to evaluate the accuracy of the anomaly detection decision results. The detection latency of different types of anomalies can also be analyzed to evaluate the real-time performance of the prediction.

[0205] In addition, during the training process of the aforementioned prediction model, a composite loss function can be designed to improve the training effect of the prediction model, thereby enhancing the accuracy of the prediction model for anomaly detection in drones. For example, for a regression task, the mean squared error loss can be used as the main loss function to accurately model the error between the predicted value and the true value. Then, the mean absolute error loss is introduced as an auxiliary loss to enhance the robustness of the model to outliers. Next, a contrastive loss is designed to measure the relationship between features through cosine similarity, enhancing the clustering effect of similar samples. Multiple individually acting loss functions can be designed separately, or multiple loss functions can be weighted and fused to obtain a fused loss function to optimize the training of each step of the model (such as the LSTM and LSGA-T processes).

[0206] Moreover, an adaptive optimizer and a cosine annealing learning rate strategy can be adopted to dynamically adjust the learning rate during the training process, ensuring that the model converges to the optimal solution and improving the reliability of adaptive threshold generation. For example, an adaptive optimizer (such as the Adam optimizer) and a cosine annealing learning rate strategy can be used during the model training process. The Adam optimizer can dynamically adjust the learning rate for different parameters based on the first-moment estimate (mean) and second-moment estimate (uncentered variance) of the gradients of each parameter. In the initial stage of training, it can quickly move towards the optimal solution; when approaching the optimal solution, it can automatically reduce the learning rate to avoid overshooting the optimal solution. For example, when the model updates certain parameters with a large amplitude, Adam will automatically reduce the learning rate of these parameters to make the updates more stable. As the number of training epochs increases, the learning rate decays in the form of a cosine function. In the beginning stage of training, the learning rate is relatively high, and the model can quickly learn the main features in the data; as training progresses, the learning rate gradually decreases, and the model can learn the detailed features more precisely.

[0207] In the above embodiments, the multi-modal time series modeling-based UAV data anomaly detection system for executing the above method embodiments includes at least one processor, a control module (chipset) coupled to at least one of the (at least one) processors, a memory coupled to the control module, a non-volatile memory (NVM) / storage device coupled to the control module, at least one input / output device coupled to the control module, and a network interface coupled to the control module.

[0208] The processor may include at least one single-core or multi-core processor, and the processor may include any combination of general-purpose processors or dedicated processors (such as graphics processors, application processors, baseband processors, etc.). For some alternative embodiments, the multi-modal time series modeling-based UAV data anomaly detection system can be used as a UAV data anomaly detection system device such as a gateway in the embodiments of the present application.

[0209] For some alternative embodiments, the multi-modal time series modeling-based UAV data anomaly detection system may include at least one computer-readable medium having instructions (such as a memory or an NVM / storage device) and at least one processor integrated with the at least one computer-readable medium and configured to execute the instructions to implement modules to perform the actions in the present disclosure.

[0210] For one embodiment, the control module may include any suitable interface controller to provide any suitable interface to at least one of the (at least one) processors and / or any suitable device or component communicating with the control module.

[0211] The control module may include a memory controller module to provide an interface to the memory. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0212] The memory may be used to load and store data and / or instructions for, e.g., a multi-modal time series modeling-based UAV data anomaly detection system. For one embodiment, the memory may include any suitable volatile memory, e.g., a suitable DRAM.

[0213] For one embodiment, the control module may include at least one input / output controller to provide an interface to the NVM / storage device and the (at least one) input / output device.

[0214] For example, the NVM / storage device may be used to store data and / or instructions. The NVM / storage device may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (at least one) non-volatile storage device (e.g., at least one hard disk drive (HDD), at least one compact disc (CD) drive, and / or at least one digital versatile disc (DVD) drive).

[0215] The NVM / storage device may include storage resources that are physically part of the device on which the multi-modal time series modeling-based UAV data anomaly detection system is installed, or it may be accessible by the device without being part of the device. For example, the NVM / storage device may be accessed via the (at least one) input / output device according to a network.

[0216] The (at least one) input / output device may provide an interface for the multi-modal time series modeling-based UAV data anomaly detection system to communicate with any other suitable device. The input / output device may include a communication component, a phonetic component, a sensor component, etc. The network interface may provide an interface for the multi-modal time series modeling-based UAV data anomaly detection system to communicate according to at least one network. The multi-modal time series modeling-based UAV data anomaly detection system may wirelessly communicate with at least one component of a wireless network according to any standard and / or protocol in at least one wireless network standard and / or protocol, e.g., access a wireless network according to a communication standard.

[0217] For one embodiment, at least one of the (at least one) processors may be loaded with the logic of at least one controller of a control module (e.g., a memory controller module). For one embodiment, at least one of the (at least one) processors may be loaded with the logic of at least one controller of a control module to form a system-level load. For one embodiment, at least one of the (at least one) processors may be integrated with the logic of at least one controller of a control module on the same die. For one embodiment, at least one of the (at least one) processors may be integrated with the logic of at least one controller of a control module on the same die to form a system-on-chip (SoC).

[0218] The embodiments of the present application have been introduced in detail above. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

[0219] An embodiment of the present invention discloses a computer-readable storage medium that stores a computer program for electronic data exchange, wherein the computer program causes a computer to execute the steps in the method for detecting abnormal drone data based on multimodal time series modeling in the foregoing embodiments.

[0220] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the method for detecting abnormal drone data based on multimodal time series modeling in the foregoing embodiments.

[0221] The device embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0222] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used for a computer to have or store data.

[0223] Finally, it should be noted that: the above-disclosed are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for detecting abnormal drone data based on multimodal time series modeling, characterized in that Including: Performing multi-level data preprocessing on the raw data obtained during the operation of the drone to generate a normalized input data stream. The multi-level data preprocessing includes noise elimination processing, dimensionality reduction processing, normalization processing, and time-sliding window reconstruction. The normalized input data stream is a data stream containing fixed time steps; Extracting the temporal features of the normalized input data stream through consecutive LSTM blocks in a multi-level feature extraction network to generate a shallow temporal feature set; Performing deep temporal modeling on the shallow temporal feature set based on a layer-scaled gated attention Transformer module to generate a deep temporal feature set; Performing multi-stage information fusion and mask screening on the shallow temporal feature set and the deep temporal feature set to generate a fused spatio-temporal correlation feature set; Generating a prediction result of the attitude parameters of the drone according to the fused spatio-temporal correlation feature set, and based on the prediction result of the attitude parameters and an adaptive threshold obtained based on the statistical distribution of the prediction error, obtaining an anomaly detection decision result of the drone by using a deviation comparison and multi-level anomaly judgment mechanism.

2. The method according to claim 1, wherein The performing multi-level data preprocessing on the raw data obtained during the operation of the drone to generate a normalized input data stream includes: Performing noise elimination on the raw data to generate preliminary smoothed time series data; According to the preliminary smoothed time series data, obtaining a multi-dimensional feature data set corresponding to the raw data; Performing a multi-dimensional feature dimensionality reduction operation on the multi-dimensional feature data set to obtain a set of principal component feature vectors of the multi-dimensional feature data set; Scaling the eigenvalues of the principal component feature vectors in the set of principal component feature vectors to a preset interval through normalization processing to obtain a set of normalized principal component feature vectors; Performing sliding window reconstruction on the set of normalized principal component feature vectors to form the normalized input data stream containing the fixed time steps.

3. The method according to claim 2, wherein The extracting the temporal features of the normalized input data stream through consecutive LSTM blocks in a multi-level feature extraction network to generate a shallow temporal feature set includes: According to the sequence of LSTM blocks in the multi-level feature extraction network, performing layer-by-layer extraction on the normalized input data stream through the consecutive LSTM blocks. The LSTM block includes an LSTM unit and a dropout layer, and adjacent LSTM blocks are connected through a hidden state transfer mechanism. The LSTM unit uses a forget gate to control the retention ratio of historical information, an input gate to control the update ratio of new information, and an output gate to control the output ratio of the current state. The dropout layer is used to randomly mask the outputs of some neurons according to a preset probability; After completing the layer-by-layer extraction according to the sequence of LSTM blocks, obtaining a shallow temporal feature set containing multi-level abstract features. The feature dimension of the shallow temporal feature set at each time step is related to the number of feature vectors in the set of normalized principal component feature vectors.

4. The method according to claim 1, wherein The performing deep temporal modeling on the shallow temporal feature set based on a layer-scaled gated attention Transformer module to generate a deep temporal feature set includes: Embed sine-cosine position encoding information in the shallow temporal feature set to generate a position-enhanced feature set, where the sine-cosine position encoding information is related to the sliding window structure reconstructed by the sliding window; Input the position-enhanced feature set into the gated attention unit, and dynamically adjust the interaction weight between the query vector and the key vector through learnable parameters to generate an attention weight matrix; Based on the attention weight matrix, perform weighted aggregation on the value vector to generate preliminary attention features; Perform layer scaling on the preliminary attention features, and adjust the feature distribution of the preliminary attention features through a trainable scaling factor to generate scaled attention features; Merge the scaled attention features and the original input features through residual connection, and based on the activation function and normalization processing, generate a deep temporal feature set.

5. The method according to claim 1, wherein The multi-stage information fusion and mask screening of the shallow temporal feature set and the deep temporal feature set to generate a fused spatio-temporal correlation feature set includes: Concatenate the shallow temporal feature set and the deep temporal feature set in the channel dimension to generate a joint feature tensor; Perform a linear transformation on the joint feature tensor to compress the feature dimension and generate a gated weight vector; Use the gated weight vector to perform weighted fusion on the shallow features in the shallow temporal feature set and the deep features in the deep temporal feature set respectively to obtain preliminary fusion features; Apply a channel-level mask operation to the preliminary fusion features, activate the effective feature channels through a non-linear function, and generate a fused spatio-temporal correlation feature set. The channel-level mask operation adopts a dynamic threshold mechanism to automatically select the activation threshold according to the importance score of the feature channels.

6. The method according to claim 1, characterized in that, The method for generating the attitude parameter prediction result of the UAV according to the fused spatio-temporal correlation feature set, and based on the attitude parameter prediction result and the adaptive threshold obtained from the statistical distribution of the prediction error, using the deviation comparison and multi-level anomaly judgment mechanism to obtain the anomaly detection decision result of the UAV includes: Input the fused spatio-temporal correlation feature set into the flattening layer of the multi-layer perceptron prediction head, and convert the fused spatio-temporal correlation feature set into a two-dimensional tensor feature set; Input the two-dimensional tensor feature set into the multi-layer fully connected network of the multi-layer perceptron prediction head for linear transformation processing of different dimensions to obtain a to-be-mapped feature tensor; Map the to-be-mapped feature tensor to the target dimension space to obtain the attitude parameter prediction result; According to the attitude parameter prediction result and the reference data, obtain a set of deviation values, where the reference data is obtained based on the real sensor data of the UAV; Based on the historical validation set data, statistically calculate the deviation mean and deviation standard deviation of each attitude parameter, and determine the initial threshold using the normal distribution; Dynamically adjust the weight coefficient of the initial threshold according to the current flight stage of the UAV to generate an adaptive threshold set, where each attitude parameter in the adaptive threshold set independently calculates the adjustment coefficient, and the adjustment coefficient is related to the flight speed, environmental interference intensity and sensor confidence index; Based on the comparison results of the set of deviation values and the adaptive threshold, obtain the abnormal detection decision result, where the abnormal detection decision result includes multi-level abnormal alarms, and the multi-level abnormal alarms include single-point instantaneous alarms, continuous abnormal alarms, and trend deviation alarms. Different levels of alarms correspond to different flight control system response strategies.

7. The method according to claim 1, wherein It further includes: Inject at least one of the following abnormal patterns into the normalized input data stream: insertion abnormal pattern, offset abnormal pattern, drift abnormal pattern, static abnormal pattern. The insertion abnormal pattern is injected by randomly selecting a time point in the normalized input data stream to add a first data point, where the first data point is a data point whose value exceeds the preset data value range. The offset abnormal pattern is injected by adding a fixed offset to the data points within the first target time period in the normalized input data stream. The drift abnormal pattern is injected by adding an offset that increases or decreases with time to the data points within the second target time period in the normalized input data stream. The static abnormal pattern is injected by setting the values of the data points within the third target time period in the normalized input data stream to the same value.

8. The method according to claim 1, characterized in that It further includes: Establish an abnormal pattern feature library, where the abnormal patterns in the abnormal pattern feature library include abnormal patterns of sensor hardware failures, signal interference, and environmental impacts; By performing similarity matching between the abnormal data in the abnormal detection decision result and each of the abnormal patterns in the abnormal pattern feature library, generate a set of failure causes. The similarity matching is obtained by using the dynamic time warping algorithm to align time series of different lengths and calculating the cosine similarity score; Based on the attention weight matrix output by the gated attention Transformer module based on layer scaling, analyze the abnormal propagation paths of each of the failure causes in the set of failure causes, and determine the root sensor nodes. The analysis of the abnormal propagation paths is realized through a visual attention heat map, and the high-weight connection paths in the abnormal propagation paths point to potential failure sources; Combined with the flight state parameters and historical maintenance records of the UAV, calculate the probability distribution of each of the failure causes in the set of failure causes; Generate hierarchical warning information and maintenance suggestions according to the probability distribution.

9. The method according to claim 1, wherein It further includes: Based on the original data corresponding to the multi-scenario flight states, obtain the normalized input data stream corresponding to the multi-scenario flight states, where the multi-scenario flight states include at least two of steady flight, high-speed maneuvering, and strong wind interference; According to the recognition difficulty of each scenario flight state in the multi-scenario flight states, configure the adjustment weights of the normalized input data stream corresponding to each scenario flight state; Adjust the normalized input data stream input into the multi-level feature extraction network according to the adjustment weights.

10. The method according to claim 9, wherein It further includes: According to the historical prediction results of the normalized input data stream input into the multi-level feature extraction network after adjustment, obtain the typical scenario features corresponding to the historical prediction results during the prediction process; According to the current scenario flight state of the UAV, determine the typical scenario features corresponding to the current scenario flight state; When predicting the prediction result of the attitude parameters of the drone, the typical scenario features are pre-input for prediction.

Citation Information

Cited By

  • Unmanned aerial vehicle hoisting load state prediction method and system based on multi-modal fusion

    CN120745437A

  • Transmission optimization method and system of unmanned aerial vehicle-mounted high and low orbit Ku frequency band satellite terminal

    CN120750403A

  • Tower crane tower body damage abnormity detection method

    CN120805081A

  • A tower crane tower damage anomaly detection method

    CN120805081B

  • Coal spontaneous combustion early warning method based on deep learning and physical constraint

    CN120823700A