Method, storage medium and application for identifying risks in positioning and tracking of unmanned aerial vehicles

By adopting enhanced Bi-LSTM network and multi-head attention mechanism in the drone system, the state data within the sliding time window is processed, and the robustness and adaptability of drones' attitude recognition in complex environments is solved, achieving higher risk identification accuracy and system security.

CN120011898BActive Publication Date: 2025-07-01QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510486699.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-01
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing drone attitude recognition technology is poor in complex environments, especially in high dynamic flight or large environmental interference, which is prone to errors and affects flight safety.

Method used

The drone positioning tracking risk identification method based on enhanced Bi-LSTM is adopted to obtain drone status data in real time, and use sliding time window technology, multi-dimensional convolutional neural network, Bi-LSTM network and multi-head attention mechanism to extract state timing characteristics and perform risk identification and prediction.

Benefits of technology

Effectively extracting the drone's status timing characteristics improves the accuracy and robustness of attitude risk identification, reduces the risk of drone autonomous driving systems in complex environments, and improves the safety of the drone control system and the real-time accuracy of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011898B_ABST
    Figure CN120011898B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of autonomous driving technology, and specifically relates to a method for identifying risks in UAV positioning and tracking, a storage medium, and an application. State data during the flight of the UAV is received in real time; after data preprocessing and standardization, a sliding time window is used to reorganize the state data to construct the state tensor format of the UAV at the next moment; a one-dimensional convolutional neural network combined with a bidirectional long short-term memory network is used for feature extraction; feature integration and final state prediction are performed. This method constructs a prediction relationship by sampling and sliding with a sliding time window for the multi-dimensional motion state information of the UAV, and combines an efficient sequence space and time two-dimensional feature extraction mechanism to improve the accuracy and stability of prediction, and improve the safety of the UAV control system and the real-time accuracy of decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to a method for identifying risks in UAV positioning and tracking, a storage medium, and an application. Background Art

[0002] Existing UAV attitude recognition technologies mostly rely on inertial measurement units (IMUs) or vision sensors for real-time attitude estimation. However, these traditional methods are relatively less robust and adaptable in complex environments. Especially in the case of high-dynamic flight or large environmental interference, errors are likely to occur, thus affecting flight safety. Therefore, how to accurately identify risks based on real-time data of the UAV's own attitude and make timely responses is the core technical issue for improving the flight safety of UAVs.

[0003] Currently, the research on UAV attitude recognition mainly focuses on two aspects: sensor data processing and algorithm optimization. Traditional attitude recognition methods are based on inertial measurement unit data, and estimate the attitude of the UAV through sensors such as accelerometers and gyroscopes. However, these methods are easily affected by factors such as sensor noise, drift error, and magnetic field interference during flight. Especially in the case of fast flight and dynamically changing environments, it is difficult to guarantee the accuracy. And the vision-based attitude estimation method, although it can provide relatively accurate spatial positioning information, its performance significantly degrades in low light, occlusion, or complex backgrounds. Therefore, traditional methods often cannot fully meet the high-precision requirements for attitude recognition in complex environments.

[0004] With the rapid development of deep learning technology, the pose recognition algorithm based on artificial intelligence has gradually become an effective solution. Especially the long short-term memory network (LSTM), due to its advantages in processing time series data, can overcome the deficiencies of traditional methods in dealing with long-term dependent data in dynamic environments. The LSTM network can analyze the time patterns in the historical flight data of drones and identify potential risks in the flight postures. For example, by analyzing the pose data of a drone at a certain moment and combining its historical flight states, LSTM can predict possible pose anomalies in the next few seconds and thus make a reaction in advance. In addition, combined with the multi-sensor data fusion technology, LSTM can effectively fuse the data from different sensors such as IMU, visual sensors, GPS, etc., further improving the accuracy and robustness of pose risk recognition. Although these methods have improved the accuracy of pose recognition to a certain extent, there are still some problems. First, deep learning models usually require a large amount of labeled data for training, and the collection and labeling of these data are both cumbersome and costly. Especially in special flight scenarios, the lack of labeled data may affect the training effect of the model. Second, the computational complexity of the LSTM model is relatively high. For a drone system with high real-time requirements, it may lead to processing delays and affect the timeliness of flight decisions. In addition, although the multi-sensor fusion technology can reduce the errors of a single sensor, the different error characteristics and data format differences of different sensors may still cause inconsistencies in the data fusion process, thus affecting the accuracy and robustness of the recognition results. In short, the existing methods still face challenges such as large data requirements, poor real-time performance, and high difficulty in sensor fusion, and need to be further optimized and innovated to meet the requirements of complex flight environments. Summary of the Invention

[0005] To solve the problems existing in the prior art, the present invention provides a method, a storage medium, and an application for risk recognition of drone positioning and tracking.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: The method for risk recognition of drone positioning and tracking includes the following steps:

[0007] S1. Receiving data, receiving the state data during the flight of the drone in real time;

[0008] S2. Preprocessing the data, after preprocessing and normalizing the data, reorganize the state data using a sliding time window to construct a state tensor format representing the state of the drone at the next moment with a fixed-length window, as the input data;

[0009] S3. Feature extraction: First, use the convolutional kernels from multiple one-dimensional convolutional neural networks to extract features of different spatial dimensions from the input data respectively. Then, use the Bi-LSTM network to further extract time-series features from the convolved data to capture the long-term dependence information in the state sequence, and combine the multi-head attention mechanism to splice the outputs to obtain an enhanced feature representation;

[0010] S4. Feature integration and final state prediction: Perform the final integration of features through the long short-term memory network LSTM, and output the state prediction value or risk classification result of the drone at the next moment.

[0011] Preferably, in step S1, the state data includes but is not limited to the heading angle, roll angle, yaw angular velocity, total velocity, throttle opening, and rudder angle.

[0012] Preferably, in step S2, first represent the state quantity of the drone at the next moment through a non-linear discrete mapping based on the multi-dimensional state data, and then use the z-score normalization method to normalize the state data; then process it in the way of sliding a time window from beginning to end to form new samples in turn.

[0013] Preferably, in step S2, according to the heading angle , yaw angular velocity , roll angle , total velocity , throttle opening and rudder angle , represent the state quantity of the drone at the next moment as:

[0014] ;

[0015] wherein, represents the state quantity at moment; represents the input quantity at moment, is its non-linear discrete mapping;

[0016] Considering that the motion state of the drone has a periodic law, sampling is performed using a sliding time window, and the prediction relationship formula for the drone state information is expressed as:

[0017] ;

[0018] That is, the input for drone model identification and prediction: ;

[0019] In the formula: , is the state quantity of the drone at moment after being processed by the sliding time window; It is a UAV prediction model that can predict the state at the next moment based on the current state and control input; is the state quantity after sliding time window processing; is the input quantity after sliding time window processing.

[0020] Preferably, in step S3, multiple groups of one-dimensional convolutional neural networks are used to perform multi-dimensional feature extraction according to different dimensions of the state data.

[0021] Preferably, in step S3, the Bi-LSTM is composed of multiple LSTM units, and each unit includes a forget gate, an input gate, and an output gate to perform information screening and update; the combination layer superimposes the output states calculated by the front and rear LSTM units in vectors.

[0022] Preferably, in the multi-head attention mechanism, scaled dot-product attention is used. After using the Softmax layer to normalize the weight values obtained by the dot-product operation of the query matrix and the key matrix, weighted summation is performed with the value matrix to obtain the attention result.

[0023] Preferably, the baseline method is used to set the safety risk threshold, and the risk envelope distribution of each state variable is calculated and determined for risk classification.

[0024] A storage medium stores a program capable of executing the above-mentioned UAV positioning and tracking risk identification method.

[0025] The application of the UAV positioning and tracking risk identification method deploys the above-mentioned storage medium in a modular manner in an embedded target platform or deploys it to an external computing module.

[0026] Compared with the prior art, the beneficial effects of the present invention are:

[0027] 1. The present invention clearly proposes a UAV positioning and tracking risk identification method based on enhanced Bi-LSTM, which can effectively extract state time series features, has strong prediction accuracy and robustness. This method determines the safety range of the UAV motion state by real-time acquiring the state data of the UAV. For the multi-dimensional motion state information of the UAV, the baseline method is used to set the safety risk threshold, clarify the boundary of the risk area, and calculate and determine the risk envelope distribution of each state variable to identify potential risks;

[0028] 2. The prediction relationship constructed by sampling and sliding with a sliding time window can effectively capture the time series dynamic features of the UAV state data and improve the accuracy and stability of prediction;

[0029] 3. A method of jointly using a one-dimensional convolutional neural network and an enhanced Bi-LSTM network is used for feature extraction. The one-dimensional convolutional network is used to effectively extract the spatial features between different time series, and the enhanced Bi-LSTM captures the long-term time-dependent features of the sequence. The two work together to form an enhanced bidirectional long short-term memory network model IBLSTM with an efficient sequence space and time two-dimensional feature extraction mechanism, effectively reducing the risk of the UAV autonomous driving system in complex environments and improving the safety of the UAV control system and the real-time accuracy of decision-making.

[0030] In summary, the risk identification method in this application can effectively capture the temporal dynamic features of UAV state data, improve the accuracy and stability of prediction, effectively reduce the risk of the UAV autonomous driving system in complex environments, and improve the safety of the UAV control system and the real-time accuracy of decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic diagram of the application of the sliding time window method in the present invention;

[0032] Figure 2 is a schematic diagram of the one-dimensional convolutional neural network in the present invention;

[0033] Figure 3 is a schematic diagram of the LSTM unit structure in the present invention;

[0034] Figure 4 is a schematic diagram of the Bi-LSTM structure in the present invention;

[0035] Figure 5 is a schematic diagram of the multi-head self-attention mechanism in the present invention;

[0036] Figure 6 is a structural diagram of the IBLSTM model in the present invention;

[0037] Figure 7 is a bar chart of the prediction correct rate of the heading angle of three models;

[0038] Figure 8 is a bar chart of the prediction correct rate of the yaw angular velocity of three models;

[0039] Figure 9 is a bar chart of the prediction correct rate of the roll angle of three models. DETAILED DESCRIPTION OF THE INVENTION

[0040] To facilitate the understanding of the present invention, the present invention will be described in more detail below with reference to the accompanying drawings and specific embodiments. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present invention more thorough and comprehensive.

[0041] The method for identifying the risks of UAV positioning and tracking based on enhanced Bi-LSTM is as follows:

[0042] Step 1: UAV model identification and data preprocessing.

[0043] Collect the state data of the UAV through real flight experiments. After data preprocessing and standardization, use the sliding time window technology to reorganize the state data, so as to construct the data samples for the training of the risk identification model.

[0044] Data collection and standardization:

[0045] Use real UAVs or computer simulations to conduct experiments under different environments and different influencing factors, obtain the real state information, position information and positioning and tracking error data of the UAV from the state data, and design different risk areas through different tracking point positions. The state information of the UAV motion model includes: heading angle , yaw angular velocity , roll angle , total speed , input quantity: throttle opening , rudder angle of the multi-input and multi-output high-dimensional system. The relationship formula for UAV model identification and prediction is:

[0046] ;

[0047] Let represent the state quantity at ; represent the input quantity at , is its non-linear discrete mapping.

[0048] First, perform data preprocessing on the collected state information and position information data of the UAV in a large number of real environments, and use the z-score standardization method to normalize the state information: ; where is the original data set, is the mean of the original data set, is the standard deviation of the original data set, is the value after standardization.

[0049] Reorganize data with a sliding time window:

[0050] Considering that the UAV motion state has a periodic law, that is, has obvious periodicity. Considering the time correlation of the state data, use the method of the sliding time window shown in Figure 1 to process the UAV motion state time series obtained through experiments.

[0051] The sliding time window technique is used to process the UAV state sequence data, that is, on the original data, a window with a fixed length (such as length ) slides forward from the beginning to the end at a certain step length, and a new data sample is formed each time it slides. The data within each window is used to construct input features, and the next state at the end of the window is used as the prediction target of the model. In this way, the temporal dynamic characteristics of the UAV state data can be effectively captured, and the accuracy and stability of the model prediction can be improved.

[0052] Model identification and prediction:

[0053] After being processed by the sliding time window, the state sequence of the UAV will be presented in a more temporally characteristic form. The UAV state information model after being processed by the sliding time window algorithm can be expressed as:

[0054] ;

[0055] That is, the relationship formula for UAV model identification and prediction problems is rewritten as:

[0056] ;

[0057] In the formula: , is the state quantity of the UAV at time after being processed by the sliding time window; is the UAV prediction model, which can predict the state at the next moment based on the current state and control input; is the state quantity after the sliding time window processing; is the input quantity after the sliding time window processing.

[0058] Through this method, the model can not only capture the time dependence of the UAV state data, but also effectively perform continuous single-step predictions, gradually approaching the true motion state of the UAV. In this way, accurate identification of the UAV motion model can be achieved, providing a basis for subsequent risk identification and control decision-making.

[0059] Step 2: Feature extraction and optimization

[0060] Use a one-dimensional convolutional neural network (1D Convolutional Neural Network, 1D-CNN) to extract spatial dimension features from the state sequence data obtained in Step 1, effectively reducing the data dimension and optimizing the feature representation, and improving the subsequent model calculation efficiency.

[0061] Considering that the spatial attributes of the state information volume of the UAV motion model have multi-dimensional characteristics, in order to improve the calculation speed, 1D-CNN is used for optimization to reorganize features among the state information at different times. The 1D-CNN one-dimensional convolutional neural network is asFigure 2 As shown, the 1D-CNN performs a convolution operation on the input data by sliding the convolutional kernel along the time axis, extracting local features between different time steps.

[0062] The convolutional kernel (filter) slides along the time axis of the input data, calculates the dot product between each local region (time window) and the convolutional kernel, and obtains the convolutional feature map. Each convolutional kernel corresponds to extracting a specific feature, such as the change trend or local fluctuation of the UAV state at a certain moment. By using multiple convolutional kernels, feature information of different granularities can be extracted, and this feature information helps to describe the motion characteristics of the UAV at different time periods.

[0063] The feature information processed by the 1D-CNN will provide more compact and highly distinguishable input data for the subsequent model, helping to improve the training efficiency and prediction accuracy of the subsequent model. Through this step, the features of the original UAV state data are effectively extracted and optimized, providing clearer input data for capturing time-dependent features in the subsequent steps, thereby enhancing the performance of the model.

[0064] Step 3: Time series feature extraction and application of Long Short-Term Memory (LSTM) network

[0065] An enhanced Bidirectional Long Short-Term Memory (Bi-LSTM) network is used to further extract time series features from the convolved data, capturing long-term dependency information in the state sequence. Compared with traditional Recursive Neural Network (RNN), it can better capture the time correlation in the sequence, improving the accuracy of the model's prediction of the state change trend.

[0066] LSTM (Long Short-Term Memory network) is a special type of RNN. The structure of the LSTM unit is as Figure 3 shown, and it is commonly used to process and predict time series data. The core advantage of LSTM is its ability to effectively capture long-term dependencies and avoid the common problem of gradient disappearance in traditional RNNs. Each LSTM unit consists of the following key components:

[0067] Forget Gate: The forget gate determines which information needs to be discarded from the cell state. As shown in the following formula represents the sigmoid activation function, is the forget gate weight matrix, is the forget gate bias matrix. It receives the hidden state from the previous moment and the current input , and generates an output with a value between 0 and 1 through the sigmoid activation function. This output represents the proportion of information retained, where 0 means complete forgetting and 1 means complete retention:

[0068] ;

[0069] Input Gate: The input gate determines how the information at the current time should be added to the cell state. It first generates an update signal through the sigmoid activation function to determine which values will be updated; at the same time, the tanh activation function generates a candidate value representing the impact of the current input on the cell state:

[0070] ;

[0071] where tanh represents the hyperbolic tangent activation function, is the input gate weight matrix, is the input gate bias matrix.

[0072] Update Cell State: Update the cell state through the combination (multiplication unit) of the forget gate and the input gate. The cell state is the core of the LSTM unit, which stores important information at the current time and past times and is passed in the sequence:

[0073] ;

[0074] ;

[0075] where, is the cell state at the current time, is the cell state at the previous time, is the input gate candidate value, is the LSTM multiplication unit weight matrix, is the multiplication unit bias matrix.

[0076] Output Gate: The output gate determines the hidden state (i.e., output) at the next time. It generates the hidden state based on the current cell state and is ultimately used for the calculation at the next time:

[0077] ;

[0078] ;

[0079] where, is the hidden state at the current time, which serves as the input at the next time and the final model output. is the weight matrix of the output gate of the LSTM cell, is the bias matrix of the output gate.

[0080] As Figure 4 shown, Bi-LSTM is an improved LSTM network. By adding a backward propagation layer and a combination layer on the basis of the traditional LSTM, it can capture the bidirectional dependencies of time series more comprehensively. Specifically, Bi-LSTM extracts information from two directions, past (forward propagation) and future (backward propagation), simultaneously during the processing of time series. The input of each LSTM cell includes the state information at the current moment as well as the state information from the past and the future, enabling it to capture the temporal characteristics in the data more precisely. Bi-LSTM consists of multiple LSTM cells, and each cell contains structures such as forget gates, input gates, and output gates, which are responsible for determining which information needs to be retained, which information needs to be forgotten, and how to update the memory cell.

[0081] The added combination layer performs a vector superposition operation on the output states calculated by the forward and backward LSTM cells as follows:

[0082] ;

[0083] where, represents the output of the Bi-LSTM layer at time t; are the output states calculated by the forward and backward LSTM cells respectively.

[0084] Step 4: Multi-Head Self-Attention (MHSA) enhances the feature extraction ability

[0085] Introduce the features output by the Bi-LSTM network into the multi-head self-attention mechanism (MHSA) to further enhance the model's ability to understand long time series information, reduce the loss of long-term sequence information, and improve the feature extraction accuracy.

[0086] Overview of the multi-head self-attention mechanism:

[0087] As a variant of the attention mechanism, the multi-head attention mechanism (MHSA) uses multiple independent attention weights to obtain the correlation of the sequence. By capturing the dependencies between different ranges through each independent head, finally, the calculation results of these multiple attention heads are concatenated or weighted and merged to form the final output. This mechanism enables the model to understand the structure of sequence data from multiple perspectives, improving the flexibility and robustness of information processing.

[0088] In step 3, Bi-LSTM has successfully captured the bidirectional temporal dependencies in time series data. However, for modeling long-term dependency information, there may still be some information loss or insufficient processing. Therefore, in this step, the computational results of Bi-LSTM are input into MHSA to reduce the loss of sequence relevance during long-term transmission and improve the ability to extract features in the time dimension.

[0089] On the output of the sequence features processed by Bi-LSTM, MHSA is applied to calculate the attention scores for each time step and weight the input features according to these scores. Specifically, for each time step, MHSA calculates the similarity between this time step and other time steps and dynamically adjusts the weights of the features based on the calculation results.

[0090] Combination of MHSA and Bi-LSTM:

[0091] As Figure 5 shown, the adopted ScaledDot-ProductAttention uses the Softmax layer to normalize the weight values obtained by taking the dot product operation of the query matrix ( ) and the key matrix ( ), and then performs a weighted sum with the value matrix ( ) to obtain the attention result. The calculation process of MHSA generally includes the following steps:

[0092] Calculation of query, key, and value: First, the features output by Bi-LSTM are transformed into query (Query), key (Key), and value (Value) matrices through linear transformation (weight matrix).

[0093] Among them,

[0094] ;

[0095] ;

[0096] ;

[0097] .

[0098] Calculation of attention scores: Then, calculate the dot product similarity between the query and the key, and normalize it through the softmax function to obtain the attention weights:

[0099] ;

[0100] Among them, is the dimension of the key vector, used to scale the dot product result to avoid overly large numerical values.

[0101] Calculation and combination of multiple heads: The attention results of each head are calculated in parallel through multiple independent attention heads. Finally, the outputs of all heads are concatenated to obtain an enhanced feature representation: Among them, is the calculation result of Bi-LSTM: 、 、 、 are respectively 、 、 weight matrices and dimensions. MHSA is calculated based on a single independent attention head, and the calculation results of n attention heads are concatenated to obtain the total attention result , and the calculation is as follows:

[0102] ;

[0103] ;

[0104] where 、 、 respectively represent the th attention head's 、 、 weight matrices, and is the MHSA weight matrix.

[0105] Through the processing of the MHSA mechanism, the model can more accurately focus on the key time points in the sequence and capture time-dependent information in different ranges from different perspectives (multiple attention heads).

[0106] The feature representation after MHSA processing has stronger representation ability, can effectively enhance the subsequent model's ability to model long-term dependence information, and provide more detailed and efficient feature input for UAV state prediction. The introduction of the multi-head self-attention mechanism enables the model to better handle long-term dependence relationships, avoid information loss, and improve the ability to focus on information in different time periods when processing complex time series data. By combining MHSA with Bi-LSTM, the accuracy and robustness of UAV state prediction can be significantly improved.

[0107] Step 5: Feature integration and final state prediction.

[0108] The final integration of features is performed through the LSTM network, and the state prediction value of the UAV at the next moment is output. This step realizes the real-time prediction of the UAV's future state.

[0109] The features are passed using an LSTM with 128 hidden units, and the calculation result of the last LSTM unit is output as 4D data as the risk prediction result. The structure of the enhanced Bi-LSTM model (IBLSTM model) proposed by the invention is as Figure 6 shown, which consists of an input layer, a CNN layer, a Bi-LSTM layer, an MHSA layer and a decoding layer, and the structure is as Figure 6 shown.

[0110] Motion model recognition and prediction can be regarded as a sequence regression problem, and the state at the next moment is predicted according to the current state at each sampling point. To evaluate the prediction accuracy, the mean square error (MSE) is selected as the reference performance evaluation index, and its calculation is as follows:

[0111] ;

[0112] The root mean square error (RMSE), which is more sensitive to data outliers, is selected as the core performance evaluation index, and its calculation is as follows:

[0113] ;

[0114] where and are the predicted value and the actual value of the th sample in the dataset respectively, and is the number of samples.

[0115] Step 6: Model training and cross-validation

[0116] The EuRoC MAV Dataset is processed, and experts annotate different data segments to ensure that the dataset is suitable for the risk classification task. Through this annotation, we can provide samples with clear labels for training the network so that the model can learn the UAV state features under different risk levels. The annotated dataset is used to train the IBLSTM model, and its generalization ability and prediction accuracy under different flight conditions are evaluated through cross-validation.

[0117] After completing the model structure design, to train the attitude risk recognition neural network model based on the enhanced Bi-LSTM structure, EuRoC MAV Dataset is selected as the training data source. EuRoC MAV Dataset is a UAV micro flight dataset provided by ETH Zurich, which contains high-frequency IMU data (acceleration, angular velocity), attitude ground truth, position data and image data collected under multiple typical indoor flight scenarios, and has the characteristics of multi-modal, multi-scene and high time accuracy, and is suitable for tasks such as UAV attitude estimation, state prediction and risk recognition.

[0118] To adapt to the risk classification task, the training data selects the IMU data and attitude state variables in the EuRoC dataset that contain different flight trajectories and action types as feature inputs, including information such as heading angle, yaw angular velocity, roll angle, three-axis acceleration, three-axis angular velocity, and timestamp. First, the original data is cleaned and synchronized, and the sliding time window algorithm is used to restructure the time series data to form input-output paired samples. Subsequently, the data is normalized using the z-score normalization method to improve the stability and generalization ability of model training.

[0119] Next, expert data annotation is carried out. According to the state changes in different time periods during the flight, the expert annotates the corresponding risk level for each data segment. The risk levels are divided into six levels, increasing from level one (lowest risk) to level six (highest risk). This annotation process is carried out manually to ensure the accuracy and reliability of the data labels. A total of 400 groups of data are annotated, and each group of data contains the state characteristics of the UAV in a specific time period and the corresponding risk level label.

[0120] After the dataset annotation is completed, each data segment not only contains the state information of the UAV but also corresponds to a clear risk level. This enables us to classify based on different risk levels when training the model, so that the network can learn the characteristics under different risk levels and ultimately improve the accuracy of risk identification.

[0121] The training of the neural network model is carried out in the Matlab Deep Learning Toolbox environment. The network structure includes an input layer, a one-dimensional convolutional layer (used to extract state space features), a bidirectional LSTM layer (used to capture the time-dependent characteristics in state evolution), a multi-head self-attention mechanism layer (used to enhance the long-term dependence modeling ability), and an output decoding layer. The model output is the risk level of the current UAV, and the training objective is to minimize the error between the predicted value and the true state.

[0122] During the training process, indicators such as MSE (Mean Squared Error) and RMSE (Root Mean Squared Error) are used to evaluate the model performance. The model parameters are updated through the Adam optimizer, and the training process dynamic monitoring graph is used to evaluate the model convergence. The hyperparameters are set with an initial learning rate of 0.001, a mini-batch size of 128, a maximum number of learning batches of 200, a learning rate decay of 0.15, and a decay period of 20.

[0123] By using multiple subsequences in the labeled EuRoC MAV Dataset for cross-validation, the generalization ability of the proposed IBLSTM model under different flight conditions is verified. Two metrics, MSE and RMSE, are used to conduct a comparative analysis of three prediction and classification models: the IBLSTM, the common LSTM model for prediction and classification tasks, and the Support Vector Machine (SVM) regression model. Through the heading angle of the unmanned aerial vehicle , yaw angular velocity , and roll angle of these three attitude data for risk classification, the results are shown in Table 1. The prediction accuracy of the three models under different types of data is as shown in Figures 7 - 9 . The results show that the IBLSTM model can effectively extract state time-series features compared with other networks, and has strong prediction and classification accuracy and robustness, providing an algorithm basis for subsequent system deployment and risk determination.

[0124] The results of cross-validation further prove that the IBLSTM model has good generalization ability under different flight conditions and can effectively process time-series data in various flight environments.

[0125] Table 1 Comparison of prediction performance of the EuRoC MAV Dataset for different models

[0126]

[0127] Step 7: Model deployment and application

[0128] Introduce the actual deployment method of the IBLSTM model, including simulation verification in the Simulink environment and the actual application plan on the external computing module, providing real-time risk identification ability for the unmanned aerial vehicle autonomous driving system.

[0129] Model deployment and usage plan. After completing the training of the IBLSTM model composed of a one-dimensional convolutional neural network, a bidirectional long short-term memory network (Bi-LSTM), and a multi-head self-attention mechanism (MHSA), the model can be applied to the system integration of attitude risk identification during the autonomous flight of the unmanned aerial vehicle in the following two ways:

[0130] One way is to import the trained deep neural network model into the Simulink environment. The trained network structure is introduced into the Simulink model in a supported format (such as DAGNetwork or dlnetwork) through the Matlab Deep Learning Toolbox, and the model is encapsulated and simulated by combining the sliding time window construction module, data preprocessing module, and prediction output module in Simulink. Further, C / C++ code is generated using Simulink Coder, enabling the deployment of the model in an embedded target platform or software-in-the-loop (SiL) testing, and performing joint simulation verification or actual operation in cooperation with the PX4 flight control system.

[0131] Another way is to deploy the IBLSTM model to an external computing module. The external computing module is an independent embedded intelligent processing device with an operating system and deep learning inference capabilities, such as Jetson Nano, Raspberry Pi, or industrial-grade edge computing terminals. The model deployment process includes the following steps:

[0132] Use the exportONNXNetwork function in Matlab or a Python-side conversion tool to export the trained IBLSTM model to ONNX or other common formats; deploy an adapted inference framework (such as PyTorch) in the external processing device and load the model file; build a data acquisition interface, and the external module establishes a connection with the PX4 flight control system through communication protocols such as serial port, UDP, or MAVROS to receive real-time state information during the flight of the drone, including heading angle, roll angle, yaw angular velocity, total velocity, throttle opening, rudder angle, etc.; the received state information enters the sliding time window cache module and is constructed into an input tensor format consistent with the model training stage; the preprocessed state data is input into the IBLSTM model for real-time inference, and the predicted value of the next moment state or the risk classification result is output; the recognition result is fed back to the PX4 flight control system through a feedback link (serial port, UDP packet, or ROS topic) for policy correction of the mission controller or attitude controller; according to actual application requirements, a risk level threshold or an anomaly detection trigger mechanism can be set. When the output result exceeds the set risk limit, the PX4 can be linked to execute control logics such as attitude stabilization and augmentation, mission interruption, and return to base.

[0133] Among them, a large amount of state information and position information of the drone in the real environment are collected through experiments using the baseline method. The normal fluctuation range of each state variable is determined by statistical analysis, and a risk envelope is formed based on this as the baseline to clarify the boundary of the risk area and identify potential risks. The risk envelope is specifically manifested as setting upper and lower threshold values for each motion state variable. When the real-time detected state data exceeds this threshold range, it is determined that there is an abnormal risk, thereby realizing the accurate extraction and clustering of risk information of the drone state data.

[0134] As an auxiliary risk identification unit of the main flight control system, the external module does not change the original internal controller structure of PX4, but only provides additional risk perception and information feedback functions. It has the characteristics of structural decoupling, flexible deployment, and strong processing ability, and is suitable for the actual scenario where complex neural network models cannot be run on resource-constrained platforms.

[0135] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

Claims

1. A method for identifying risks of drone positioning and tracking, characterized in that: Here are the steps: S1. Data reception: real-time reception of the status data of the UAV during flight, including but not limited to heading angle, roll angle, yaw angular velocity, total speed, throttle opening and rudder angle; S2. Data preprocessing: First, the state quantity of the drone at the next moment is represented by nonlinear discrete mapping based on the multidimensional state data, and then the state data is normalized by using the z-score standardization method. The state data is reorganized by sliding the sliding time window from the beginning to the end to form new samples in sequence, and the state tensor format of the drone at the next moment is constructed by demarcating the state data with a fixed-length window as the input data; According to the heading angle , yaw angular velocity , Roll Angle , total speed , throttle opening and rudder angle , the state of the drone at the next moment It is expressed as: ; in, express The state quantity at the moment; express The input amount at time, is its nonlinear discrete mapping; Considering that the UAV motion state has a periodic law, a sliding time window sampling is adopted, and the UAV state information prediction relation is expressed as: ; That is, drone model identification and prediction input: ; Where: , which is processed by sliding time window The state of the drone at the moment; It is a prediction model for drones, which can predict the state at the next moment based on the current state and control input; It is the state quantity after sliding time window processing; is the input quantity after sliding time window processing; S3, feature extraction, first use the convolution kernels from multiple one-dimensional convolutional neural networks to extract features of different spatial dimensions of the input data, then use the Bi-LSTM network to further extract time series features from the convolved data to capture the long-term dependency information in the state sequence, and combine the multi-head attention mechanism to splice the output to obtain enhanced feature representation; S4, feature integration and final state prediction, the final integration of features is performed through the long short-term memory network LSTM, and the state prediction value or risk classification result of the drone at the next moment is output.

2. The method for identifying risk of UAV positioning and tracking according to claim 1, characterized in that: In step S3, multiple groups of one-dimensional convolutional neural networks are used to perform multi-dimensional feature extraction according to different dimensions of the state data.

3. The method for identifying risk of UAV positioning and tracking according to claim 2, characterized in that: In step S3, the Bi-LSTM is composed of multiple LSTM units, each of which contains a forget gate, an input gate, and an output gate to perform information screening and updating; the combination layer vectorizes the output states calculated by the previous and next LSTM units.

4. The method for identifying risk of UAV positioning and tracking according to claim 3, characterized in that: In the multi-head attention mechanism, scaled dot product attention is adopted. The Softmax layer is used to perform dot product operation on the query matrix and the key matrix to obtain the weight value, which is then normalized and weightedly summed with the value matrix to obtain the attention result.

5. The method for identifying risk of UAV positioning and tracking according to claim 4, characterized in that: The baseline method is used to set the safety risk threshold, and the risk envelope distribution of each state variable is calculated and determined for risk classification.

6. A storage medium, characterized in that A program capable of executing the drone positioning and tracking risk identification method as described in any one of claims 1 to 5 is stored.

7. Application of the UAV positioning and tracking risk identification method, characterized in that: The storage medium as described in claim 6 is modularly deployed in an embedded target platform or in an external computing module.

Citation Information

Patent Citations

  • Intelligent memory leak predicting and tracking method and system for micro-service architecture

    CN119690725A

  • A security risk analysis system and method based on multimodal data processing

    CN119784145A