Unmanned aerial vehicle dynamic path control method based on multi-source sensing fusion

By setting the inertial measurement unit on the drone as the time reference, interpolating and mapping the visual and lidar perception units, and combining deep neural networks for feature fusion and delay compensation, the problem of path estimation deviation during high-dynamic maneuvers of the drone is solved, achieving more accurate and stable path control.

CN120722928AActive Publication Date: 2025-09-30FUJIAN POLYTECHNIC OF INFORMATION TECH

Patent Information

Application Number
CN202511134074.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-30
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

When a drone performs high-dynamic maneuvers such as high-speed dives and sharp turns, the synchronization error and fusion lag of the multimodal perception system cause the path estimation to deviate seriously from the actual trajectory, causing control oscillation or instability of the flight control system.

Method used

By setting the inertial measurement unit time series as the main time axis, interpolation resampling and nonlinear time mapping are performed on the visual and lidar perception units, combined with deep neural networks for feature fusion, and estimating system delays, delay compensation control signals are generated to drive the drone's path control.

Benefits of technology

The fusion error is significantly reduced, the spatiotemporal accuracy and robustness of path estimation are enhanced, the control lag under high dynamic conditions is reduced, and the stable flight control performance of the UAV in complex flight missions is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120722928A_ABST
    Figure CN120722928A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle dynamic path control method based on multi-source sensing fusion, and relates to the technical field of unmanned aerial vehicle path control, and the method comprises the following steps: collecting original sensing data; performing interpolation resampling on the data acquired by the visual perception unit to realize time alignment; inputting the multi-modal sensing data after time alignment into a time interpolation sensing fusion network model; estimating system delay from sensing data acquisition to control instruction output, constructing a feedforward prediction model based on a current state, predicting a future state, and further generating a delay compensation control signal; the delay compensation control signal is input into a flight controller, attitude, speed and acceleration instructions are calculated and output, the unmanned aerial vehicle is driven to execute a path control task, and parameter adaptive adjustment is performed based on flight control feedback information; according to the method, the space-time accuracy of path estimation under the high-dynamic maneuvering condition is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field related to unmanned aerial vehicle (UAV) path control, and specifically to a UAV dynamic path control method based on multi-source perception fusion. Background Art

[0002] When a drone performs highly dynamic maneuvers such as high-speed dives and evasive turns, its multimodal perception systems (such as vision, inertial measurement units, and lidar) will expose a dual failure phenomenon in the time domain: On the one hand, due to the timestamp drift of the sensor itself, there is synchronization error in the multi-source data before fusion; on the other hand, there is an overall system-level lag in the fusion algorithm when responding to these high-dynamic actions (such as visual frame rate delay and inertial measurement unit data hysteresis), which causes additional response delay between perception and action.

[0003] The combined effect of synchronization drift and fusion lag will cause the path estimation to deviate seriously from the true trajectory in time and space, causing control oscillation or instability of the flight control system, especially in flight missions that require a fast and real-time perception-decision-making-control closed loop. Summary of the Invention

[0004] In order to solve the defects of the existing technology, the present invention provides a UAV dynamic path control method based on multi-source perception fusion.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a method for controlling a dynamic path of an unmanned aerial vehicle (UAV) based on multi-source perception fusion, comprising the following steps: The multimodal perception data synchronous collection step, during the flight of the UAV, the inertial measurement unit, visual perception unit and lidar perception unit synchronously collect raw perception data; The perception data time alignment step uses the time series of the inertial measurement unit as the unified main time axis, and interpolates and resamples the data collected by the visual perception unit to achieve time alignment; In the perception feature fusion modeling step, the time-aligned multimodal perception data is input into the time interpolation perception fusion network model, feature representations of the inertial measurement unit, visual perception unit, and lidar perception unit are extracted, and feature fusion is performed on a unified time axis. A sensing delay compensation control signal generation step estimates the system delay from sensing data acquisition to control command output, builds a feedforward prediction model based on the current state, predicts the future state, and then generates a delay compensation control signal; The dynamic path control execution step inputs the delay compensation control signal into the flight controller, calculates and outputs attitude, speed and acceleration instructions, drives the UAV to perform path control tasks, and performs parameter adaptive adjustment based on the flight control feedback information.

[0006] As a preferred technical solution of the present invention, in the perception data time alignment step, the interpolation resampling of the data collected by the visual perception unit includes: Set the inertial measurement unit time series as the main time axis as the alignment reference for the visual image frames; According to the original timestamp of the visual image frame and the time point of the adjacent inertial measurement unit data, a time interpolation function is constructed to resample the visual image data; Choose linear interpolation, Lagrange interpolation or spline interpolation strategy according to sampling density; Generate a visual perception unit feature sequence corresponding to an inertial measurement unit time node; Outputs a visual perception unit feature data sequence that is consistent with the inertial measurement unit timestamp.

[0007] As a preferred technical solution of the present invention, in the perception data time alignment step, the nonlinear time mapping of the data collected by the lidar perception unit includes: Extract the original timestamp sequence of the inertial measurement unit and lidar perception unit; Construct a cost matrix to measure local time differences; Use dynamic time warping algorithm to search for the minimum cost path in the cost matrix; Output the nonlinear time mapping index between the inertial measurement unit and the lidar perception unit; The time index of the lidar perception unit is adjusted according to the mapping index to achieve alignment with the main time axis of the inertial measurement unit.

[0008] As a preferred technical solution of the present invention, the perceptual feature fusion modeling step includes: Input the time-aligned data of the inertial measurement unit, visual perception unit, and lidar perception unit into the time interpolation perception fusion network model; Build deep neural network structures for processing time series data, including multilayer perceptrons or Transformer networks; A temporal continuity constraint mechanism is introduced to complement the missing features in low-frequency sampling modes; Jointly generate multimodal fusion feature output results on a unified timeline.

[0009] As a preferred technical solution of the present invention, the perceptual feature fusion modeling step further includes a main modality identification process, which includes the following steps: Send the fused features to the modal analysis module; Statistically analyze the residual value, information entropy and related distribution characteristics of each perception modality; Construct a modal confidence evaluation function based on residual value, information entropy and related distribution characteristics; Dynamically select the perception modality with the highest current confidence as the main modality; When the confidence of the main modality is lower than the threshold, it switches to the suboptimal modality and updates the fusion network structure.

[0010] As a preferred technical solution of the present invention, the step of generating the perceptual delay compensation control signal includes: Calculate the immediate control signal based on the current fused perception features; Input the current state sequence into the feedforward prediction model to predict the state values ​​at multiple time points in the future; Estimate the total system latency from when sensory data is collected until control commands take effect; Calculate the delay compensation control signal by combining the current control quantity, the predicted state and the estimated delay value; The control signal is output to the flight control system to achieve path tracking control.

[0011] As a preferred technical solution of the present invention, the method for estimating system delay includes the following steps: Calculate the error between the desired flight control output and the feedback signal; Extract the current moment fusion perception state vector; Using extended Kalman filter or neural network-based function structure as delay estimator; The control error and the current state are input into the delay estimator to model the delay value; The system delay estimation result is updated in each control cycle and fed back to the perception delay compensation control signal generation step.

[0012] As a preferred technical solution of the present invention, the construction of the feedforward prediction model includes: Collect the fusion perception state data of the current moment and several historical moments to form a time series; Build a recurrent neural network, gated recurrent unit, or Transformer prediction structure based on the sequence; Input the time series into the prediction network to establish the state transition relationship; Output the state estimation results within the specified time window in the future; The predicted state is used as a feedforward control variable to participate in the perception delay compensation control signal generation step.

[0013] As a preferred technical solution of the present invention, the method further includes a model self-correction step, which includes: Compare the flight control system output with the desired control quantity of the fusion perception system to obtain the control deviation; Taking the control deviation as input, a loss function is constructed to evaluate the error; Calculate the gradient of the loss function with respect to the parameters of the time interpolation perception fusion network model and the delay modeling module; Update model parameters according to the set learning rate; When the error exceeds the set threshold, the model is updated and a closed-loop optimization mechanism is formed.

[0014] The beneficial effects of the present invention are: 1. In this invention, by setting the inertial measurement unit time series as the main time axis, interpolating and resampling the visual perception units separately, and using a dynamic time warping algorithm to perform nonlinear mapping on the lidar perception unit timestamps, the three perception modalities are unified in the time domain. This perception data time alignment step effectively eliminates the synchronization error caused by sensor clock drift between perception sources and provides a unified time reference for the subsequent perception feature fusion modeling step. This design significantly reduces the accumulation of fusion errors and enhances the spatiotemporal accuracy of path estimation under highly dynamic maneuvering conditions.

[0015] 2. The present invention uses a modal analysis module to statistically analyze the residual values, information entropy, and distribution characteristics of each perception mode, establish a modal confidence assessment function, and dynamically select the primary mode. When the confidence level of the primary mode drops below a threshold, the system automatically switches to the suboptimal mode and updates the fusion network structure. This mechanism ensures that even if a perception mode experiences signal interruption, reduced accuracy, or other failures, the fusion output remains usable, effectively enhancing the system's perceptual robustness in complex flight missions.

[0016] 3. In the present invention, through the combined design of the perception delay compensation control signal generation step and the control compensation execution module, the system can estimate the total system delay from perception data acquisition to control signal execution in real time, and predict the future state based on the feedforward prediction model, thereby generating a delay-compensated control signal; the perception delay compensation control signal generation step also combines the estimation of system delay with the construction of the feedforward prediction model, and adaptively updates the delay within each control cycle. This link-level response optimization significantly reduces the control lag under high-dynamic actions, allowing the UAV to maintain stable flight control performance in tasks such as high-speed dives and sharp turns to avoid obstacles. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 Schematic diagram of the workflow of the UAV dynamic path control method of the present invention.

[0018] Figure 2 Schematic diagram comparing the actual trajectory of the aircraft of the present invention and the trajectory without delay compensation.

[0019] Figure 3 Schematic diagram comparing the trajectory after compensation according to the present invention and the trajectory after no delay compensation. DETAILED DESCRIPTION

[0020] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0021] Example 1:

[0022] like Figure 1 As shown, a method for dynamic path control of a UAV based on multi-source perception fusion includes the following steps: In the step of synchronously collecting multimodal perception data, during the flight of the drone, the inertial measurement unit (IMU), visual perception unit and lidar perception unit synchronously collect raw perception data. The inertial measurement unit collects acceleration, angular velocity and attitude angle data, with a frequency of usually 200~1000Hz and high time accuracy. The visual perception unit collects continuous image frames, mainly used to obtain forward or downward image frames, with a frequency of generally 30~60Hz, for environmental understanding and attitude estimation. The lidar perception unit obtains environmental three-dimensional point cloud data and adds timestamp information to each set of data, with a frequency of usually 10~20Hz.

[0023] In the perception data time alignment step, the time series of the inertial measurement unit is used as the unified main time axis. Based on its high sampling rate and low jitter characteristics, it is used as an alignment reference, and the data collected by the visual perception unit is interpolated and resampled to achieve time alignment. Linear interpolation, Lagrange interpolation or spline interpolation methods are used to match the image frame data with the inertial measurement unit time node; then the dynamic time warping algorithm is used to map the timestamp of the lidar perception unit to the inertial measurement unit time axis to achieve the unification of the three perception modalities in the time domain. The algorithm searches for the optimal path by constructing a cost matrix, thereby mapping the original timestamp of the lidar perception unit to the inertial measurement unit time axis to achieve alignment under non-uniform sampling.

[0024] In the perception feature fusion modeling step, the time-aligned multimodal perception data is input into the time interpolation perception fusion network model, the feature representations of the inertial measurement unit, visual perception unit and lidar perception unit are extracted, and feature fusion is performed on a unified time axis. A deep neural network structure (such as MLP or Transformer) can be used to process inputs of different modalities, and end-to-end feature fusion is achieved through alignment of the time dimension. Shared encoders and modality-specific sub-channels can be designed in the network structure to take into account modal differences and joint expression capabilities. During the fusion process, the current main information source is dynamically identified based on the confidence mechanism, and a confidence mechanism is introduced. By analyzing indicators such as residuals and information entropy of different modalities in the current scenario, its reliability is dynamically evaluated, and the current "main modality" is selected to dominate the fusion weight.

[0025] The step of generating a delay-compensated control signal for perception estimates the system delay from the acquisition of perception data to the output of control commands. A feedforward prediction model is constructed based on the current state to predict the future state and generate a delay-compensated control signal. In this step, the system-level delay in the perception control link is considered, and a feedforward prediction and lag compensation mechanism is designed to ensure the timeliness of the control signal.

[0026] The dynamic path control execution step inputs the delay compensation control signal into the flight controller, calculates and outputs attitude, speed and acceleration instructions, drives the UAV to perform path control tasks, and performs parameter adaptive adjustment based on the flight control feedback information.

[0027] The present invention can be widely used in the following typical scenarios: urban low-altitude logistics flight, high-speed drone competitive flight, etc.

[0028] Furthermore, in the perception data time alignment step, the interpolation resampling of the data collected by the visual perception unit includes: Set the inertial measurement unit time series as the main time axis as the alignment reference for the visual image frames; Considering that inertial measurement units usually have a high frequency (generally above 100 Hz) and a stable sampling period, as well as low transmission delay and system time drift, the data timestamp collected by the inertial measurement unit is selected as the main data time axis of the entire system. This main time axis serves as the benchmark reference for the alignment of all subsequent perception modal data.

[0029] According to the original timestamp of the visual image frame and the time points of the adjacent inertial measurement unit data, a time interpolation function is constructed to resample the visual image data, and linear interpolation, Lagrange interpolation or spline interpolation strategy is selected according to the sampling density; The visual perception unit usually has a low sampling frequency. The timestamp corresponding to each frame of the image may not match the main time axis of the inertial measurement unit. As a result, the time interval between adjacent frames is much larger than the sampling period of the inertial measurement unit, making direct correspondence impossible. Therefore, a time interpolation function needs to be constructed to remap the visual image data on the inertial measurement unit time axis. The construction process of the interpolation function includes: Original image frame timestamp extraction: Extract the timestamp sequence of consecutive image frames, denoted as {t (i)}; Inertial measurement unit time point selection: at every two consecutive visual frame time points t (i) With t (i+1) Select all the inertial measurement unit time points {t (j) IMU}; Interpolation strategy selection: Select the interpolation method based on the time point spacing and system real-time requirements: If the data changes linearly, linear interpolation can be used; If the image frame has nonlinear displacement characteristics, Lagrange interpolation can be used; If you want to improve the fitting accuracy and smooth the curve, you can use cubic spline interpolation.

[0030] Generate a visual perception unit feature sequence corresponding to an inertial measurement unit time node; After completing the interpolation function construction, the visual image frame data is resampled using the selected interpolation strategy according to the time nodes of the main time axis of the inertial measurement unit to obtain a "pseudo-synchronous" visual frame feature value sequence corresponding to each inertial measurement unit timestamp. This feature value sequence may include the pixel intensity distribution, edge extraction features, key point matching features, etc. of the original image frame, which is used as input in the subsequent feature fusion modeling step; In particular, to improve the robustness of the interpolation results, a smoothing factor or weighting term can be introduced during the interpolation process to suppress interpolation anomalies caused by image quality jitter or frame loss. At the same time, a mapping table between the original timestamp and the interpolation result needs to be established during the resampling process for traceability and system debugging.

[0031] Output the visual perception unit feature data sequence consistent with the inertial measurement unit timestamp. Finally, output the visual perception feature sequence that is completely consistent with the inertial measurement unit time axis. This sequence is synchronized with the original data such as acceleration and angular velocity of the inertial measurement unit to ensure that in the subsequent time interpolation perception fusion network, the features of each modality are under a consistent time reference, thereby improving the fusion accuracy and timing continuity.

[0032] Furthermore, in the perception data time alignment step, the nonlinear time mapping of the data collected by the lidar perception unit includes: Extract the original timestamp sequence of the inertial measurement unit and lidar perception unit; The inertial measurement unit timestamp sequence is as follows: ; The lidar timestamp sequence is as follows: ; Among them, N>M, and the sampling frequency of the inertial measurement unit is much higher than that of the lidar.

[0033] Construct a cost matrix to measure local time differences; Construct a cost matrix for the time series of the inertial measurement unit and the lidar: ; Where D is represented as a real matrix of size N×M, D i,j Inertial measurement unit timestamp t i IMU With the lidar timestamp t i LIDAR The absolute value of the time difference between them is as follows: ; The cost matrix reflects the time difference between the perception modalities at each moment and is used for subsequent matching path search.

[0034] Use dynamic time warping algorithm to search for the minimum cost path in the cost matrix; Based on the above cost matrix, the dynamic time warping (DTW) algorithm is used to find a minimum cost path P={(i1,j1),(i2,j2),...,(i k ,j k )}, the path satisfies the following constraints: The path starts at (1,1) and ends at (N,M); Path continuity: the steps between adjacent nodes are (1,0), (0,1) or (1,1); Minimize the cumulative cost: ; Through DTW path search, the optimal match between the lidar sampling points and the inertial measurement unit time nodes can be achieved, adapting to the nonlinear differences between different modal time series.

[0035] Output the nonlinear time mapping index between the inertial measurement unit and the lidar perception unit; According to the matching path P obtained by the DTW algorithm, the lidar timestamp sequence {t j LiDAR} is mapped to the main time axis of the inertial measurement unit to generate a set of nonlinear time mapping index tables: ; Each lidar data point corresponds to the optimal inertial measurement unit time node to achieve the unification of modal time reference.

[0036] The time index of the laser radar perception unit is adjusted according to the mapping index to achieve alignment with the main time axis of the inertial measurement unit. According to the mapping index table, the three-dimensional point cloud data sequence of the laser radar perception unit {P j LiDAR} Rearrange and interpolate to match the inertial measurement unit time axis sequence, and output the lidar feature data sequence consistent with the inertial measurement unit time series {P ij LiDAR}, as the input to the subsequent perceptual feature fusion modeling step.

[0037] Furthermore, the perceptual feature fusion modeling step includes: Input the time-aligned data of the inertial measurement unit, visual perception unit, and lidar perception unit into the time interpolation perception fusion network model; Build deep neural network structures for processing time series data, including multilayer perceptrons or Transformer networks; A temporal continuity constraint mechanism is introduced to complement the missing features in low-frequency sampling modes; Jointly generate multimodal fusion feature output results on a unified timeline.

[0038] The temporal interpolation perception fusion network model mainly includes a Transformer-based feature interpolation module and a modal confidence assessment sub-module. The interpolation module uses a standard multi-head self-attention mechanism to process the time series input of three types of modal data (inertial measurement unit, vision, and lidar) on a unified time axis. Each modality uses an independent encoder for feature extraction, and uses the attention mechanism in the fusion layer to integrate cross-modal information. The modal confidence assessment module calculates the confidence score based on the output residual, entropy index and feature variation of each modality, and selects and dynamically switches the main modality. The overall structure supports parallel reasoning and can dynamically adjust the information dominance according to the missing or distorted conditions of different modalities.

[0039] This step builds a time interpolation perception fusion network model to complete the following functions: Interpolation and structure completion of time series between modes: For modalities with low sampling frequency or sampling gaps, a time series interpolation mechanism is introduced. The interpolation structure can be constructed based on Transformer, gated recurrent neural network (GRU) or multi-layer perceptron (MLP). The interpolation model combines contextual time information with adjacent modal states to achieve reasonable estimation of missing features.

[0040] Generate unified fusion feature representation: The feature representations of different modalities are fused through a shared attention mechanism or parallel sub-networks to output a structured feature vector with unified dimensions, temporal consistency and modal complementarity. The output result serves as the input of the path control and decision-making module.

[0041] Introducing the temporal continuity constraint mechanism: The time domain regularization term is used to ensure the smoothness and interpretability of the interpolation results, avoiding the drastic fluctuations in perception results caused by mode switching or frame loss. During the training process, the prediction stability is improved by minimizing the feature differences between time series.

[0042] The generated unified time feature output will be passed as input to the perception delay compensation control signal generation step to construct the feedforward control quantity. The modal confidence identification result can also be fed back to the time alignment step to adaptively adjust the interpolation weights for low-confidence modes.

[0043] Furthermore, the perceptual feature fusion modeling step further includes a main modality identification process, which includes the following steps: Send the fused features to the modal analysis module; Statistically analyze the residual value, information entropy and related distribution characteristics of each perception modality; Construct a modal confidence evaluation function based on residual value, information entropy and related distribution characteristics; Dynamically select the perception modality with the highest current confidence as the main modality; When the confidence of the main modality is lower than the threshold, it switches to the suboptimal modality and updates the fusion network structure.

[0044] Specifically, for each perceptual modality i, its confidence function is defined as follows: ; Among them, En i (Entropy i ) represents the entropy value of the predicted output of the i-th mode, reflecting the uncertainty of the model, Re i (Residual i ) represents the residual between the modality perception feature and the fusion reference model, Dr i (DropRate i) represents the recent frame loss / unavailable ratio of the modality, which is used to measure its stability. f is a trainable nonlinear mapping function, which can also be approximated by a weighted linear combination model: ; Among them, α represents the entropy weight. Since entropy can effectively measure uncertainty, the default value of α is set to 0.4. β represents the residual weight. Since the residual can reflect the observation accuracy, the default value of β is set to 0.4. γ represents the frame loss weight. Since frame loss is a minor but important factor, the default value of γ is set to 0.2.

[0045] In each perception cycle of the system, the main mode M is identified t The steps are as follows: ; That is, the mode with the highest current confidence is selected as the dominant mode, and the threshold θ is set. If the confidence of the main mode ConfidenceM t <θ, the mode switching mechanism is triggered: ; That is, the mode with the highest confidence other than the current mode is selected. Preferably, θ is set in the range of [0.6, 0.75] and fine-tuned according to the overall signal-to-noise ratio of the system and the actual usage scenario. The higher θ is, the easier it is for the system to trigger mode switching (more conservative).

[0046] Further, if Figure 2-Figure 3 As shown, the step of generating the perceived delay compensation control signal includes: Calculate the immediate control signal based on the current fused perception features; The system first constructs a unified perception feature vector based on the time-aligned multimodal perception data. This feature vector is output by the time interpolation perception fusion network model and represents the flight state x at the current time t. t , the state is input to a conventional controller (such as a proportional-integral-derivative controller PID or a model predictive controller MPC) to generate an immediate control signal C(x t ).

[0047] Input the current state sequence into the feedforward prediction model to predict the state values ​​at multiple time points in the future; Considering that there is a non-negligible system delay from the generation of the control signal to the effect on the aircraft, the system needs to predict the state x after the delay window. t+Δ To this end, a feedforward prediction network with time series modeling capability, such as a recurrent neural network, is used for modeling and outputs the state estimation results for a period of time in the future. The output P of the prediction module is feedforward (x t+Δ ) represents the delay compensation feedforward control signal component.

[0048] Estimate the total system latency from when sensory data is collected until control commands take effect; The system needs to understand the perception-decision-execution delay δ in the current link t To perform real-time estimation, an extended Kalman filter (EKF) or a neural network estimator can be used to dynamically model the current delay value based on the error between the flight control expected output and the sensor feedback, combined with the current fusion perception state vector. The delay estimator structure takes the control error and the perception state as input and outputs the delay value δ t and is continuously updated in each control cycle through a feedback mechanism.

[0049] Calculate the delay compensation control signal by combining the current control quantity, the predicted state and the estimated delay value; The instantaneous control signal C(x t ), delay compensation feedforward control signal component P feedforward (x t+Δ ) and the output delay value δ t Perform synthesis to generate the final control signal after delay compensation: ; Among them, L lag (δ t ) is the hysteresis error compensation item, which is used to offset the execution error caused by the delay. The hysteresis compensation module can use the forward-backward method to combine the historical control results and the current prediction state to adjust the control command so that it still maintains the target control performance after the hysteresis interval.

[0050] Outputting the control signal to the flight control system to achieve path tracking control; The final control signal u t After being solved by the flight controller, attitude angle, velocity and acceleration instructions are generated for execution, which act on the UAV's motors or control surfaces and other actuators. The flight control system will return execution feedback in each control cycle. This feedback is used to further correct the prediction model, delay estimator and control gain to achieve system adaptive update.

[0051] Among them, Figure 2-Figure 3 It can be clearly seen that the control signal is improved by perceiving the delay compensation and applying the output of the prediction module P feedforward (x t+Δ ) represents the delay compensation feedforward control signal component, which can significantly improve the trajectory tracking accuracy.

[0052] Specifically, the horizontal axis in the figure represents the flight time, with a time range of 0-10 seconds, 20 sampling points per second (a total of 200 data points), and the horizontal axis shows the complete time history of the aircraft during high-speed flight. The time points correspond to specific actions during the flight process (such as acceleration, turning, deceleration, etc.). The vertical axis represents the deviation between the actual position of the aircraft and the expected position (true trajectory). A positive error indicates that the aircraft position is ahead of the expected position, and a negative error indicates that the aircraft position lags behind the expected position. The error value directly reflects the accuracy of the control system.

[0053] Figure 2 The black curve in the figure represents the ideal flight path, which serves as a reference. The most significant changes occur between 2 and 4 seconds (turning) and 6 and 8 seconds (acceleration). In the uncompensated path estimate (dark gray dashed line), the maximum positive error (lead) reaches 8 cm at 2.5 seconds, the maximum negative error (lag) reaches 12 cm at 4 seconds, and the error peaks at 10 cm at 7 seconds. The average error is 18.7 cm. This indicates that the error increases significantly in areas with large acceleration changes and exhibits a systematic offset (overall higher than the true trajectory).

[0054] Figure 3 After medium delay compensation, the estimated trajectory is highly consistent with the true trajectory, with an average error of only 3.2 cm, a maximum error of 6.8 cm (an 81.9% reduction compared to without compensation), and an error standard deviation of 1.5 cm (an 81.9% reduction compared to without compensation). The most significant performance improvements are seen in areas where the error is reduced from 15 cm to 2 cm during sharp turns (3-4 seconds) and from 12 cm to 3 cm during acceleration (7-8 seconds).

[0055] In summary, by accurately modeling the system delay value (δ t ) and applying delay-compensated feedforward control signal components successfully reduced path error by 82.9% and maximum error by 81.9%. Experimental data validated the effectiveness of this technology in high-speed dynamic scenarios (150-250 km / h). In particular, during critical maneuvers such as sharp turns and acceleration / deceleration, trajectory overlap increased to 96.7%, significantly improving the aircraft's control accuracy and safety.

[0056] Furthermore, the method for estimating the system delay includes the following steps: Calculate the error between the expected flight control output and the feedback signal. This error information typically includes: attitude error (difference in pitch, roll, and yaw angles), velocity error (difference between target velocity and feedback velocity), and position error (Euclidean distance between the predicted trajectory and the current position). Extract the current fused perception state vector, including the 3D position and attitude after fusion of the inertial measurement unit, vision, and lidar, the current velocity and acceleration, and the confidence index of each mode in the perception system (for weighting). This state input and the error together constitute the input variables for modeling system delay. An extended Kalman filter or a neural network-based function structure is used as a delay estimator. Optionally, an extended Kalman filter is selected and an extended Kalman filter is designed. The state variable includes the delay parameter, and the control output residual is used as the observation quantity to recursively calculate the delay estimate, which can achieve robust modeling of small-scale nonlinear delay changes. Input the control error and current state into the delay estimator to calculate the delay value δ t , the delay value δ t The core parameters of the subsequent control compensation execution module participate in the feedforward prediction compensation process of the control instructions, and update the system delay estimation results in each control cycle and feed them back to the perception delay compensation control signal generation step.

[0057] Furthermore, the construction of the feedforward prediction model includes: Collect the fusion perception state data of the current moment and several historical moments to form a time series; Build a recurrent neural network, gated recurrent unit, or Transformer prediction structure based on the sequence; Input the time series into the prediction network to establish the state transition relationship; Output the state estimation results within the specified time window in the future; The predicted state is used as a feedforward control variable to participate in the perception delay compensation control signal generation step.

[0058] Specifically, we first define the current time as t. To predict the state value at the future time t+Δ, we need to construct a time series input based on historical data: ; Where N represents the length of the time window. The size of N should balance the expressiveness of historical information with the computational burden of the model. If it is too small, the prediction ability will be weak, while if it is too large, redundancy will be introduced and the delay will increase. It is generally 5~20. t The fused perception state vector at time t contains the fused features of the inertial measurement unit, vision, and LiDAR. All states are aligned with the main time axis of the inertial measurement unit to ensure time consistency.

[0059] Considering the limited edge computing capabilities of the drone platform, the feedforward prediction module prioritizes the use of a lightweight gated recurrent unit (GRU) structure for state modeling. Its calculation expression is as follows: ; ; Among them, h t is the hidden state of GRU, the dimension can be set to 64~128, W o , b o is the output layer weight and bias, T represents the last moment after GRU expansion, which is equal to t, x t+Δ Represented as a predicted future state.

[0060] To enhance the feedforward compensation capability, the model can be designed to output state estimates for multiple future time steps at once: ; This structure enables more flexible delay compensation strategies (for example, dynamically selecting the optimal prediction time point). The output dimension is [Δ×d], where d is the dimension of a single state vector. In addition, the prediction model must undergo standardization and regularization control to avoid numerical explosion during prediction.

[0061] Furthermore, the method further comprises a model self-correction step, which comprises: Compare the flight control system output with the desired control quantity of the fusion perception system to obtain the control deviation; Taking the control deviation as input, a loss function is constructed to evaluate the error; Calculate the gradient of the loss function with respect to the parameters of the time interpolation perception fusion network model and the delay modeling module; Update model parameters according to the set learning rate; When the error exceeds the set threshold, the model is updated and a closed-loop optimization mechanism is formed.

[0062] Specifically, during the model self-correction process, the system measures the difference between the current fusion control output and the actual flight control feedback by constructing a loss function L, which is as follows: ; Where N represents the control instruction dimension (such as 3D attitude + thrust), is the predicted control signal output by the perception fusion module, u t is the actual flight control system feedback output. When L>ϵ (where ϵ ranges from [0.05 to 0.5]), the system will trigger the online update of the fusion model and delay modeling module parameters using the following gradient descent formula: ; Among them, η∈[10 -5 ,10 -2 ] is the learning rate, θ includes the parameter set of the fusion network model and the delay modeling module, ∇ θ L represents the gradient of the loss function, which is expressed as: ; This process can be triggered every 10 to 20 frames, ensuring that the operating load of the embedded platform is controllable and improving the closed-loop control stability of the overall system in highly dynamic task scenarios.

[0063] Furthermore, this method is deployed on an embedded flight control platform with edge AI inference capabilities. The platform has the following core features: support for multi-sensor parallel acquisition and high-frequency data processing, a tensor computing acceleration module (such as an NPU or GPU) to run deep learning models, support for shared memory and real-time communication mechanisms to coordinate data flow between modules, a flight control output interface that can directly control attitude, speed, and trajectory commands, and includes the following functional modules: Time alignment processing module: performs interpolation alignment and resampling between the inertial measurement unit and the visual perception unit, and completes the nonlinear time mapping of the lidar perception unit through the dynamic time warping algorithm to achieve a unified modal time base; Perception fusion reasoning module: Builds a time-interpolation perception fusion network model based on the Transformer or multi-layer perceptron structure, generates joint features and identifies the main modality, and supports dynamic switching of the main information source; Control compensation execution module: This module integrates delay modeling, feedforward prediction, and perception delay compensation control signal generation steps to output corrective control instructions that take system delay into account. Each module collaboratively processes unified time series and fused feature data through a shared memory mechanism to ensure the real-time and consistency of the overall system.

[0064] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for dynamic path control of unmanned aerial vehicles based on multi-source perception fusion, characterized in that: The following steps are involved: The multimodal perception data synchronous collection step, during the flight of the UAV, the inertial measurement unit, visual perception unit and lidar perception unit synchronously collect raw perception data; The perception data time alignment step uses the time series of the inertial measurement unit as the unified main time axis, interpolates and resamples the data collected by the visual perception unit to achieve time alignment, and then uses the dynamic time warping algorithm to map the timestamp of the lidar perception unit to the inertial measurement unit time axis; In the perception feature fusion modeling step, the time-aligned multimodal perception data is input into the time interpolation perception fusion network model, feature representations of the inertial measurement unit, visual perception unit, and lidar perception unit are extracted, and feature fusion is performed on a unified time axis. A sensing delay compensation control signal generation step estimates the system delay from sensing data acquisition to control command output, builds a feedforward prediction model based on the current state, predicts the future state, and then generates a delay compensation control signal; The dynamic path control execution step inputs the delay compensation control signal into the flight controller, calculates and outputs attitude, speed and acceleration instructions, drives the UAV to perform path control tasks, and performs parameter adaptive adjustment based on the flight control feedback information.

2. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1 is characterized in that: In the perception data time alignment step, the interpolation resampling of the data collected by the visual perception unit includes: Set the inertial measurement unit time series as the main time axis as the alignment reference for the visual image frames; According to the original timestamp of the visual image frame and the time point of the adjacent inertial measurement unit data, a time interpolation function is constructed to resample the visual image data; Choose linear interpolation, Lagrange interpolation or spline interpolation strategy according to sampling density; Generate a visual perception unit feature sequence corresponding to an inertial measurement unit time node; Outputs a visual perception unit feature data sequence that is consistent with the inertial measurement unit timestamp.

3. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1 is characterized in that: In the perception data time alignment step, the nonlinear time mapping of the data collected by the lidar perception unit includes: Extract the original timestamp sequence of the inertial measurement unit and lidar perception unit; Construct a cost matrix to measure local time differences; Use dynamic time warping algorithm to search for the minimum cost path in the cost matrix; Output the nonlinear time mapping index between the inertial measurement unit and the lidar perception unit; The time index of the lidar perception unit is adjusted according to the mapping index to achieve alignment with the main time axis of the inertial measurement unit.

4. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1, characterized in that: The perceptual feature fusion modeling step includes: Input the time-aligned data of the inertial measurement unit, visual perception unit, and lidar perception unit into the time interpolation perception fusion network model; Build deep neural network structures for processing time series data, including multilayer perceptrons or Transformer networks; A temporal continuity constraint mechanism is introduced to complement the missing features in low-frequency sampling modes; Jointly generate multimodal fusion feature output results on a unified timeline.

5. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1 is characterized in that: The perceptual feature fusion modeling step further includes a main modality identification process, which includes the following steps: Send the fused features to the modal analysis module; Statistically analyze the residual value, information entropy and related distribution characteristics of each perception modality; Construct a modal confidence evaluation function based on residual value, information entropy and related distribution characteristics; Dynamically select the perception modality with the highest current confidence as the main modality; When the confidence of the main modality is lower than the threshold, it switches to the suboptimal modality and updates the fusion network structure.

6. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1, characterized in that: The step of generating the perceived delay compensation control signal comprises: Calculate the immediate control signal based on the current fused perception features; Input the current state sequence into the feedforward prediction model to predict the state values ​​at multiple time points in the future; Estimate the total system latency from when sensory data is collected until control commands take effect; Calculate the delay compensation control signal by combining the current control quantity, the predicted state and the estimated delay value; The control signal is output to the flight control system to achieve path tracking control.

7. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1, characterized in that: The method for estimating system delay comprises the following steps: Calculate the error between the desired flight control output and the feedback signal; Extract the current moment fusion perception state vector; Using an extended Kalman filter or a neural network-based function structure as a delay estimator; Input the control error and the current state into the delay estimator to model the delay value; The system delay estimation result is updated in each control cycle and fed back to the perception delay compensation control signal generation step.

8. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 6 is characterized in that: The construction of the feedforward prediction model includes: Collect the fusion perception state data of the current moment and several historical moments to form a time series; Build a recurrent neural network, gated recurrent unit, or Transformer prediction structure based on the sequence; Input the time series into the prediction network to establish the state transition relationship; Output the state estimation results within the specified time window in the future; The predicted state is used as a feedforward control variable to participate in the perception delay compensation control signal generation step.

9. The method for dynamic path control of an unmanned aerial vehicle based on multi-source perception fusion according to claim 1, characterized in that: The method further comprises a model self-correction step, which comprises: Compare the flight control system output with the desired control quantity of the fusion perception system to obtain the control deviation; Taking the control deviation as input, a loss function is constructed to evaluate the error; Calculate the gradient of the loss function with respect to the parameters of the temporal interpolation-aware fusion network model and the delay estimator; Update model parameters according to the set learning rate; When the error exceeds the set threshold, the model is updated and a closed-loop optimization mechanism is formed.

Citation Information

Patent Citations

  • Unmanned aerial vehicle control delay compensation method based on predictive control

    CN110096064A

  • Unmanned aerial vehicle flight path optimization system and method based on artificial intelligence and Internet of Things

    CN120255548A

  • Multi-mode unmanned aerial vehicle perception data fusion analysis method

    CN120337132A

  • Speed-adaptive variable-overload half-rolling inversion control method for unmanned aerial vehicle

    CN120406548A

  • Method and apparatus for compensation for operator in a closed-loop control system

    US5777871A

Cited By

  • XR large-space outdoor positioning and automatic calibration method

    CN121430638A

  • A method for XR large-space outdoor positioning and automatic calibration

    CN121430638B

  • Smooth speed regulation control method and system of laser weeding robot based on prediction feedback

    CN121613964A

  • Unmanned aerial vehicle tracking system and method based on intelligent lamp pole network

    CN121899800A