Unmanned aerial vehicle self-optimization stabilization system and method based on multi-modal perception

By combining multimodal perception, intelligent decision-making, and a self-optimization layer, the problems of sensor errors and control lag in UAVs under complex environments are solved, enabling stable and intelligent flight of UAVs in diverse scenarios.

CN121900181APending Publication Date: 2026-04-21福建金创利信息科技发展股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
福建金创利信息科技发展股份有限公司
Filing Date
2026-01-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing UAV stabilization systems suffer from problems such as sensor error accumulation, vision system failure, and lagging control algorithms in complex environments, making it difficult to achieve environmental adaptability and intelligent decision-making.

Method used

A multimodal perception layer is used for spatiotemporal registration and tightly coupled state estimation of heterogeneous sensors. Combined with adaptive denoising and dynamic weight allocation of the intelligent decision layer, hierarchical attitude control, and flight data compression and federated learning updates by the self-optimization layer, a self-optimizing stable system is constructed.

Benefits of technology

It improves the accuracy and robustness of UAV state estimation in complex environments, enhances the dynamic adaptability and stability of control response, and realizes the continuous evolution and intelligence of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900181A_ABST
    Figure CN121900181A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle self-optimization stabilization system and method based on multi-mode perception, and relates to the field of unmanned aerial vehicle flight control. The system comprises a multi-mode sensing layer, an intelligent decision-making layer and a self-optimization layer, wherein the sensing layer synchronously realizes microsecond registration of heterogeneous sensors through FPGA hardware, and calculates attitude, disturbance torque and sensor states through a tight coupling model; the decision-making layer generates an optimal instruction by adopting self-adaptive denoising, environment identification and dynamic weight distribution and combining layered anti-interference control; and the self-optimization layer compresses data through incremental PCA, learns an aggregation model through federation, and realizes strategy iteration by relying on the security OTA. Interference caused by sensor mismatch, disturbance response lag and the like is reduced, disturbance suppression response time can be shortened, attitude stability of multiple scenes of the unmanned aerial vehicle is improved, and the method is suitable for requirements of scenes such as professional surveying and mapping, emergency rescue and film and television creation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) flight control, and more particularly to a self-optimizing stabilization system for UAVs based on multimodal perception. Background Technology

[0002] Current UAV stabilization systems face the practical need for improved environmental adaptability and intelligent decision-making in various application scenarios. Existing technologies mainly employ single or simple fusion architectures of inertial measurement units or visual sensors, which have certain limitations in complex operating environments: under strong wind disturbances, MEMS inertial sensors are susceptible to vibration coupling, resulting in cumulative drift and errors in system attitude calculation; in low-light environments, the effective detection range of traditional vision systems is relatively limited, affecting obstacle avoidance reliability. Mainstream fixed-parameter PID control algorithms have limited dynamic adaptability in time-varying environments and slow closed-loop response to sudden disturbances, which may cause short-term instability and oscillations in the fuselage.

[0003] The design of existing technologies is relatively linear in the "perception-fusion-decision" chain. A single sensor is difficult to achieve full frequency domain coverage, fixed weight fusion makes little use of prior environmental information, and traditional control algorithms are not good at predicting nonlinear disturbances.

[0004] Currently, while the industry has introduced deep learning for environment classification, most of these methods operate in an open-loop model with offline training and online applications, and the closed-loop adaptive optimization of control parameters still needs improvement. However, as the application of drones expands in scenarios such as professional surveying, emergency rescue, and film and television production, it is necessary to enhance the dimensions of environmental perception and adaptive decision-making capabilities, and to build a stable system that can be continuously optimized with the accumulation of data. Summary of the Invention

[0005] The purpose of this invention is to provide a self-optimizing stabilization system and method for unmanned aerial vehicles based on multimodal perception, which solves the technical problems of spatiotemporal mismatch of heterogeneous sensors, hysteresis response to disturbances, extreme visual failure, and system self-evolution security.

[0006] A self-optimizing stabilization system and method for unmanned aerial vehicles (UAVs) based on multimodal perception, characterized by comprising:

[0007] A multimodal sensing layer is used for spatiotemporal registration and tightly coupled state estimation of heterogeneous sensors.

[0008] The intelligent decision-making layer is used for environmental adaptive noise reduction, dynamic weight allocation, and hierarchical attitude control.

[0009] A self-optimizing layer is used for flight data compression, federated learning evolution, and secure OTA updates.

[0010] The multimodal sensing layer achieves microsecond-level spatiotemporal registration through an FPGA hardware synchronization module. This module employs a linear interpolation extrapolation algorithm to uniformly align the raw observation data from each sensor under its own clock reference to a globally synchronized timescale, controlling the interpolation error to within a second-order small quantity. The system pre-calibrates the clock deviation and transmission delay of each sensor, uses attitude quaternion spherical interpolation for high-frequency IMU data, and employs linear extrapolation for low-frequency visual and laser data, achieving seamless integration of cross-frequency band data fusion.

[0011] The perception layer constructs a tightly coupled state estimation model of inertial, vision, and lidar, with the relative pose between the UAV's body coordinate system and the world coordinate system as the core state variable. The system integrates visual feature point reprojection errors, IMU pre-integration residuals, and lidar point-to-plane geometric distance constraints, solving for the state increment through a weighted least squares framework. The information matrix is ​​composed of the inverse matrix of the noise covariance of each sensor, where the lidar observation covariance is dynamically adjusted according to the environmental recognition probability, adaptively reducing its confidence weight in feature-sparse scenarios to enhance the system's robustness to degraded environments.

[0012] The perception layer utilizes ground point cloud data collected by lidar during near-ground flight, extracts ground normal vectors through principal component analysis, and directly observes the pitch and roll angles of the UAV under conditions of no significant horizontal acceleration, serving as an independent attitude reference when the vision system fails, thus preventing long-term drift from pure inertial calculations.

[0013] The perception layer works in conjunction with a miniature anemometer and a temperature and humidity sensor to measure the three-dimensional wind speed and ambient air density in the aircraft's coordinate system in real time, and estimates the equivalent disturbance torque based on the UAV's aerodynamic parameter model. This disturbance torque is directly fed into the intelligent decision layer for feedforward compensation, shortening the disturbance suppression response time.

[0014] The perception layer includes sensor health assessment and fault isolation functions, and performs fault diagnosis by calculating the Mahalanobis distance of the observation residuals in real time. When the distance exceeds the chi-square distribution threshold, the system determines that the corresponding sensor is faulty, and then activates the robust filtering mode, adaptively increasing the observation noise covariance of the sensor and reducing its information weight, while activating redundant sensor channels to ensure the continuity and integrity of state estimation.

[0015] The intelligent decision-making layer includes an adaptive denoising preprocessing module. For high-frequency vibration coupling noise from the IMU, an improved wavelet thresholding denoising algorithm is employed. Its threshold function introduces an exponential transition region based on the traditional hard threshold, avoiding signal amplitude deviation. The number of wavelet decomposition layers is dynamically selected according to the vibration frequency band, the threshold is adaptively calculated using the Stein unbiased risk estimation criterion, and the noise variance is obtained by a robust median estimator, effectively resisting sudden outlier interference.

[0016] For LiDAR point clouds, preprocessing employs a strategy combining statistical outlier filtering and voxel mesh downsampling. Statistical outlier filtering removes distance anomalies and eliminates noise caused by sunlight interference or abnormal reflections. Voxel mesh downsampling compresses the point cloud size to 30% of its original size while preserving environmental geometric features, significantly reducing the computational load for subsequent processing.

[0017] For images taken by a black light camera in extremely low-light environments, the system employs a variance-stabilized transform to approximate Gaussianization of mixed noise, followed by improved nonlocal mean denoising. During similar block matching, a joint illumination-texture similarity metric is introduced, and the filtering parameters are adaptively adjusted based on the local noise level. After denoising, an inverse transform is performed to restore brightness, and adaptive histogram equalization is used to enhance contrast while avoiding noise amplification.

[0018] For binocular depth cameras, the system constructs a multi-feature fusion cost volume, combining Census transform, gradient magnitude difference, and semantic segmentation feature difference to improve matching robustness in low-texture or repetitive texture scenes. Cost aggregation employs an improved semi-global matching algorithm, performing dynamic programming in multiple directions and introducing confidence weights to reflect texture richness, allowing for disparity jumps. The final disparity map removes outliers through sub-pixel thinning and left-right consistency checks.

[0019] Visual feature extraction and tracking employ a Gaussian pyramid hierarchical strategy, extracting improved FAST corner points and BRIEF descriptors at each layer. During feature matching, the pose increment obtained from IMU pre-integration is used to perform motion-compensated projection on feature points. After projecting feature points from the previous frame to the current frame, matching is searched only within the projection neighborhood. The search radius is determined by the IMU uncertainty covariance, significantly reducing computational complexity. Successfully matched feature points have their 3D positions updated temporally using Kalman filtering, with higher confidence levels for longer tracking durations. Points with excessive reprojection errors or tracking loss exceeding a certain number of frames are identified as outliers and removed, ensuring long-term consistency of visual observations.

[0020] The intelligent decision-making layer includes an environment recognition and dynamic weight allocation module. This module employs a lightweight convolutional neural network, taking a tensor composed of time-series data from multiple sensors as input and outputting an environmental probability vector. Classification categories include strong wind, low light, near-obstacle, and normal. A dynamic weight generator calculates sensor weight allocation in real-time based on environmental probabilities, introducing a penalty based on the rate of environmental change to prevent weight jumps. This mechanism enables the system to increase the weights of the IMU and anemometer in strong wind scenarios and increase the weights of the LiDAR and black light camera to over 90% in low-light or near-obstacle scenarios, achieving adaptive scheduling of sensor resources.

[0021] The intelligent decision-making layer includes a hierarchical attitude control algorithm module, which is divided into three levels: bottom layer, middle layer and top layer.

[0022] The underlying layer employs a linear active disturbance rejection controller (ADRC) to establish a second-order motion model for each attitude channel of the UAV. A third-order extended state observer is used to estimate the total disturbance in real time, which includes model uncertainties and external interference. The observer bandwidth and controller bandwidth maintain a fixed proportional relationship, and the gain is tuned using the bandwidth method to avoid tedious calculations. The control law compensates for the disturbance estimate in real time, achieving a rapid response mechanism of disturbance observation and cancellation, suppressing most instantaneous disturbances.

[0023] The middle layer employs a deep reinforcement learning hyperparameter tuner, constructed based on a near-end policy optimization algorithm. The state space includes attitude error, error derivative, extended state observer perturbation estimates, ambient wind speed, and uncertainty indices, comprehensively reflecting system dynamics and environmental disturbances. The action space defines perturbations to the underlying controller parameters and outputs the policy distribution through a neural network. The reward function consists of attitude error penalties, control smoothness penalties, and stability rewards; training uses generalized dominance estimation to balance bias and variance. This hyperparameter tuner is trained offline on millions of simulation datasets and fine-tuned online after deployment, achieving dynamic adaptation of the underlying controller parameters to time-varying wind fields.

[0024] The top layer employs model predictive control, combining anemometer and visual data to predict wind field changes within the next few seconds and generate attitude feedforward commands. Its cost function includes attitude error in the prediction time domain, angular acceleration penalty, and terminal cost. The angular acceleration term is used to suppress rapid changes in the actuators to extend their lifespan, while the terminal cost ensures closed-loop stability. Wind field prediction uses a lightweight time-series model to input historical wind speeds and environmental characteristics, outputting future wind speed distributions. Model predictive control generates predictive compensation commands based on this, adjusting the attitude in advance to offset gust impacts and reducing the lag of traditional feedback control.

[0025] The self-optimization layer includes a data storage and feature extraction module. Local storage uses a flight event triggering mechanism to record key data segments, including attitude errors, control commands, and environmental parameters. Incremental principal component analysis is used to perform online dimensionality reduction on time-series features, recursively updating the sample mean and covariance matrix, extracting principal components whose cumulative variance contribution rate exceeds a preset threshold, and compressing the original high-dimensional data into low-dimensional feature vectors to reduce wireless transmission bandwidth requirements.

[0026] The self-optimization layer incorporates a federated learning evolution mechanism. The cloud server constructs a global optimization objective and weights and aggregates the local losses of each UAV, with the weights proportional to the data quality of each UAV to prevent low-quality data from polluting the global model. After accumulating a preset flight duration, the cloud aggregates the encrypted model gradients uploaded by each UAV, updates the global model, and then distributes it via over-the-air download technology, enabling continuous evolution of the control strategy and improved scenario generalization capabilities.

[0027] The self-optimization layer includes a secure OTA update strategy, employing an A / B partition dual backup mechanism. The A partition algorithm runs in flight mode, while the B partition silently receives updates. Before updating, the digital signature and hash value of the new model are verified to ensure the source is trustworthy, and the range of new parameters is constrained by Lyapunov stability theory. Online verification uses simplified criteria to ensure that the updated controller gain change and observer bandwidth do not exceed the stability margin limit. The update window period is selected during standby time. Switching conditions are jointly determined by stability criteria and timing logic; if the update fails, the system automatically rolls back to the previous version to ensure flight safety. Rollback is triggered when the attitude error exceeds the limit or the Lyapunov function diverges. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the architecture of a UAV self-optimization stabilization system based on multimodal perception provided in an embodiment of the present invention. Detailed Implementation

[0029] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0030] like Figure 1 As shown, embodiments of the present invention propose a self-optimizing stabilization system for unmanned aerial vehicles (UAVs) based on multimodal perception. The method includes the following steps:

[0031] Step 1: Multimodal sensing layer, used for spatiotemporal registration and tightly coupled state estimation of heterogeneous sensors. Microsecond-level clock alignment is achieved through FPGA hardware synchronization, an inertial-vision-LiDAR joint optimization framework is constructed, and gravity reference and disturbance torque observation are extracted.

[0032] Step 2: Intelligent decision layer, used for sensor data adaptive denoising, online environmental identification and dynamic weight allocation, and generation of hierarchical attitude control strategies, including bottom-level LADRC disturbance rejection control, middle-level DRL parameter tuning and top-level MPC feedforward compensation;

[0033] Step 3: Self-optimization layer, used for lightweight compression of flight data, global model evolution through cloud-based federated learning, and secure OTA incremental updates, enabling continuous iteration of control strategies and online verification at the edge.

[0034] In this embodiment of the invention, an inertial-vision-LiDAR tightly coupled joint optimization framework is constructed based on spatiotemporal registration of heterogeneous sensors. This framework, combined with double-buffered synchronization and manifold optimization, extracts multi-dimensional perception features. Compared to traditional single-sensor or loosely coupled fusion methods, this approach better reflects the complex dynamic environmental characteristics of UAVs during flight, such as strong winds, low light, and high-speed maneuvers. It accurately captures the coupling relationship between the aircraft's attitude and environmental disturbances, improving the robustness and accuracy of state estimation. A hierarchical attitude control strategy is generated based on environmental recognition and dynamic weight allocation. Combined with LADRC disturbance rejection and DRL adaptive parameter tuning, the control response more closely matches the dynamic changes of actual flight conditions, ensuring flight stability in extreme environments. The generated control strategy iterative update information is synchronized to the onboard control platform, enabling the UAV to autonomously adapt to diverse flight scenarios. This allows for continuous evolution of control performance without human intervention, effectively reducing the risk of flight loss of control in complex environments and improving the reliability and intelligence level of UAV operations.

[0035] In a preferred embodiment of the present invention, step 1 (multimodal sensing layer), based on a collaborative architecture of FPGA hardware synchronization and ARM core computation, realizes spatiotemporal registration and tightly coupled state estimation of heterogeneous sensors, simultaneously completing gravity reference extraction, environmental disturbance torque observation, and sensor health assessment, obtaining time-consistent and robust body attitude and environmental state data, which may include:

[0036] Step 101: Based on the overall architecture of the multimodal perception layer, clock deviation calibration is completed in the FPGA hardware synchronization module of the UAV-borne heterogeneous computing platform. The synchronization protocol adopted is the IEEE 1588 Precision Time Protocol (PTP). The synchronization objects cover all heterogeneous sensors such as IMU, LiDAR, binocular black light camera, and anemometer. Specifically, during the system power-on initialization phase, the main control MCU sends Sync messages to each sensor through the IEEE 1588 Precision Time Protocol (PTP), records the Delay_Req timestamp returned by the sensor, and calculates the clock deviation. With transmission delay The calibration process takes less than 1 second, the PTP synchronization message transmission frequency is 10 Hz, and the convergence error is less than [missing information]. The clock skew of the IMU sensor Synchronization is performed every 10 minutes to compensate for long-term drift caused by crystal oscillator temperature drift, ensuring that the convergence error of the data clock of each sensor is consistently less than [value missing]. .

[0037] Step 102: Based on the clock deviation calibration results completed in Step 101, a hard real-time data scheduling unit is configured inside the FPGA to achieve stable caching and parallel reading and writing of multi-sensor data. Specifically, this includes: configuring a ring FIFO buffer with a depth of 1024 samples inside the FPGA; ensuring the determinism of data reading through an interrupt mechanism; allocating independent FIFO channels for different sensor data types, with the IMU data FIFO channel width set to 32 bits (matching 1 kHz sampling rate and 16-bit precision data), the LiDAR point cloud data FIFO channel width set to 128 bits (adapting to 3D coordinates + reflection intensity data), and the binocular black light camera image data using a dedicated frame buffer FIFO with a depth extended to 2048 samples. An independent interrupt request line is configured for each FIFO, with the interrupt trigger threshold set to 75% of the FIFO depth to prevent data overflow; a double-buffered Ping-Pong structure is adopted, where buffer B is read when buffer A is written to. The switching of buffer blocks is controlled by the FPGA's internal state machine, and the switching logic is dynamically calibrated by the clock deviation calibration results to ensure that the switching delay is less than one FPGA clock cycle, thus achieving hard real-time data scheduling.

[0038] Step 103: Based on the multi-sensor data scheduled in Step 102, a differentiated interpolation algorithm is used to perform data extrapolation and alignment to obtain sensor data with a unified global time scale. Specifically, this includes: for IMU data with a sampling rate of 1 kHz, the attitude quaternion spherical linear interpolation (Slerp) formula is used to interpolate the sampling time... Data extrapolation to global synchronization timescale The interpolation error is a second-order small quantity. ,in The Slerp interpolation formula is:

[0039]

[0040] in , For 10 Hz LiDAR point clouds and 30 Hz visual frames, a linear extrapolation algorithm is used, with the extrapolation window limited to within 20 ms. This applies when the UAV is in a high-speed maneuvering scenario (e.g., angular velocity). When the extrapolation window is shortened to 10 ms, an angular acceleration compensation term is introduced to reduce extrapolation error.

[0041] Step 104: Based on the aligned multi-sensor data from Step 103, configure a tightly coupled state estimation sliding window optimization framework on the ARM Cortex-A78 core to provide basic parameter configuration for state estimation. Specifically, this includes defining the optimal system state vector as follows: ,in Pose represented by a unit quaternion (4-dimensional) For the body's angular velocity (3D), and The zero biases are gyroscope and accelerometer (3D each), and the sliding window contains 10 pose nodes, with a total dimension of Keyframe selection is based on viewpoint changes ( ) or translation distance ( ).

[0042] Step 105: Based on the sliding window framework configured in step 104, construct the multi-sensor joint residual vector and complete the Jacobian matrix calculation, specifically including: multi-sensor joint residual... Ceres-Solver automatic differentiation calculations are used for lidar point cloud sets. Points in First, the KD-Tree algorithm is used to perform nearest neighbor search, with a search radius of... The number of nearest neighbors is 20-30, then the nearest neighbor set is... Execute RANSAC algorithm (50 iterations, interior point threshold) Output plane normal vector With the center of mass Finally, based on the external parameters of the laser-radar system... (calibration accuracy) , ), calculate residuals Residual Jacobian matrix Real-time calculations are performed using automatic differentiation of bilateral manifolds.

[0043] Step 106: Based on the residual vector and Jacobian matrix obtained in Step 105, the weighted least squares algorithm is used to solve for the state estimation result, specifically including: Jacobian matrix. Dimensions ( (Total residuals), the Levenberg-Marquardt algorithm is used for iterative optimization (initial damping factor). Maximum number of iterations: 20; convergence threshold: ), information matrix Among them, the observation covariance of lidar , Dynamically adjusted based on environmental feature sparsity, for sparse feature scenes (such as open lawns) from Increase to (Weight reduced by 75%), feature-rich scenes (such as urban building complexes) remain unchanged. .

[0044] Step 107: Based on the aligned lidar data from Step 103, candidate ground point clouds are selected in near-ground flight mode. Specifically, this step involves: in near-ground flight mode (relative altitude...) Triggered by ( ) on lidar point cloud Perform height filtering to extract ( For candidate ground points, a new planar continuity test is added, requiring the height difference between adjacent points. This improves robustness.

[0045] Step 108: Based on the candidate ground point clouds selected in Step 107, extract and verify the ground normal vectors through principal component analysis. Specifically, this includes performing principal component analysis (PCA) on the candidate points and extracting the ground normal vector corresponding to the smallest eigenvalue. The verification condition is And the number of interior points Added 5 consecutive frames of normal vector angle changes Consistency checks are performed to ensure the reliability of the results.

[0046] Step 109: Based on the reliable ground normal vector obtained in step 108, calculate the gravity reference attitude observation value, specifically including: directly calculating the roll angle. With pitch angle The formula is

[0047]

[0048] This observation is in a visual failure scenario (nighttime illumination). As an independent attitude reference, the weighting coefficient The efficiency improved from 0.1 to 0.8 in a low-light environment and in a strong vibration environment (IMU noise). Simultaneously increase the weight to suppress inertial drift.

[0049] Step 110: Based on the anemometer data scheduled in step 102, complete the measurement and filtering of the environmental wind speed. Specifically, this includes: the miniature three-dimensional anemometer outputting the system wind speed at a sampling rate of 100 Hz. High-frequency noise is removed by a fifth-order Butterworth low-pass filter (cutoff frequency 5 Hz).

[0050] Step 111: Based on the filtered wind speed data and temperature and humidity sensor data obtained in step 110, correct the ambient air density. Specifically, this includes: adjusting the ambient air density according to the ambient temperature data collected by the temperature and humidity sensor. With humidity The air density is corrected using the ideal gas law, and the formula is as follows:

[0051]

[0052] in For water vapor partial pressure, Atmospheric pressure, altitude The barometer measures the time in real time. Otherwise, use standard atmospheric pressure. .

[0053] Step 112: Based on the wind speed data obtained in Step 110 and the corrected air density in Step 111, calculate the environmental disturbance torque. Specifically, this includes: calling the aerodynamic Jacobian lookup table (LUT) constructed from the wind tunnel experiment, with the system wind speed as the input. With normalized angular velocity ( Environmental disturbance torque can be queried online using bilinear interpolation. It is fed into the subsequent control layer at a frequency of 100 Hz.

[0054] Step 113: Based on the state estimation results obtained in step 106 and the original observation data of each sensor, calculate the sensor observation residuals and residual covariance, specifically including: for the first... One sensor, calculates the observation residual. and residual covariance ( (For state estimation error covariance), IMU residuals are smoothed using a 10-frame sliding window, and visual residuals are robustly estimated using RANSAC.

[0055] Step 114: Based on the residual covariance obtained in step 113, determine the sensor fault state using the Mahalanobis distance. This specifically includes: calculating the Mahalanobis distance. If 5 times in a row (Chi-square distribution 95% confidence threshold), and the condition is still met after a fault confirmation delay of 200 ms, triggering the fault flag.

[0056] Step 115: Based on the fault determination result in step 114, perform robust sensor fault switching and weight adjustment, specifically including: within 20 ms after the fault, adjust the sensor observation noise covariance. Increase the weight by 10 times and decrease the information weight, while activating redundant channels. In case of IMU failure, switch to visual gyroscope mode (angular velocity estimation accuracy approximately...). When visual failure occurs, the LiDAR weight increases to 0.9, and LiDAR failure degrades to inertial-visual integrated navigation, with state covariance... Automatically magnifies 3 times.

[0057] This embodiment combines the heterogeneity of UAV sensor data with the dynamics of the flight environment to construct a tightly coupled joint optimization framework of inertial-vision-lidar, which improves the lack of robustness caused by traditional attitude estimation relying on a single sensor or loose coupling fusion, and improves the accuracy and reliability of state estimation in complex environments.

[0058] In a preferred embodiment of the present invention, step 2 (intelligent decision layer), based on the state estimation and environmental data obtained in step 1, completes adaptive denoising preprocessing, environment identification and dynamic weight allocation, hierarchical attitude control and fault redundancy switching, to obtain the optimal attitude control command adapted to the current environment, which may include:

[0059] Step 201: Based on the aligned IMU data from step 103, a wavelet denoising algorithm is used to remove heterogeneous noise. Specifically, this includes: using the Daubechies-4 wavelet basis to decompose the data into layers. The threshold function is

[0060]

[0061] threshold Using Stein's unbiased risk estimation (SURE) in The interval golden section search was performed 20 times to determine the value, and the noise variance was determined by a robust median estimator. The propeller is obtained by setting the third layer of detail coefficients corresponding to the frequency (80-120Hz). (Preserve vibration characteristics), other layers are designed This is used to enhance noise reduction.

[0062] Step 202: Based on the aligned LiDAR data from step 103, noise reduction and compression are performed using statistical filtering and downsampling algorithms. Specifically, this includes: configuring the number of nearest neighbor points using statistical outlier filtering. Outlier determination coefficient To eliminate noise from sunlight interference, the voxel grid size was reduced by 5 cm, the point cloud size was compressed to 30% of the original size, and the voxel size was reduced to 2 cm during near-ground flight, while retaining fine terrain features.

[0063] Step 203: Based on the binocular black light camera data aligned in step 103, perform noise reduction and enhancement processing for low-light environments, specifically including: illumination... At that time, the Anscombe transform is performed first. Then apply the improved NLM filter (search window) Pixels, similar block windows Filter parameters (Adaptive adjustment), and finally enhanced by CLAHE (mesh). , The image SNR was improved by 15-20 dB, and the feature matching success rate increased from 12% to 68%.

[0064] Step 204: Based on the enhanced visual image from step 203, construct a multi-feature fusion cost volume to optimize binocular disparity. This specifically includes: constructing a multi-feature fusion cost volume.

[0065] (Low-light scenes) (Weight increased to 0.7), SGM penalty item , Parallax search radius Near field 96, far field 64, after left-right consistency check, mark the unconfidence area, repeating texture scene (such as white wall) will The weight is increased to 0.4, and semantic priors are used to eliminate ambiguity.

[0066] Step 205: Based on the multimodal data preprocessed in steps 201-204, flight environment classification and recognition are achieved using a lightweight CNN. Specifically, this includes: employing a MobileNetV4-ConvSmall network, inputting... The tensor (3 channels of RGB image + 2 channels of IMU time-spectral image + 4 channels of LiDAR depth image) has 2.3 M parameters and is optimized using TensorRTFP16 to reduce processing time. The IMU time-spectral image is generated by Short Time Fourier Transform (STFT) (window length 256, overlap rate 50%, frequency resolution 10 Hz). The environmental classification threshold is strong wind. Low light Near-field obstacles .

[0067] Step 206: Based on the environment classification results of step 205, generate multi-sensor dynamic weights using fuzzy logic, specifically including: dynamic weights. Fuzzy weight matrix Adjusted offline based on expert experience (e.g.) hour, and All are set at 0.4. Reduced to 0.1; hour, , Reduced to 0.05), environmental change rate penalty item To prevent weight jumps.

[0068] Step 207: Based on the dynamic weights generated in step 206, smooth weight switching is achieved through filtering. Specifically, the weight output is applied to state estimation after being filtered by a first-order low-pass filter (cutoff frequency 0.5 Hz). A 0.5 s gradual transition is introduced during switching to ensure smooth scheduling.

[0069] Step 208: Based on the environmental disturbance torque and attitude error obtained in Step 1, configure the underlying LADRC disturbance rejection controller, specifically including: taking the pitch channel as an example, establishing a second-order motion model. in, For the total disturbance, the control efficiency coefficient The bandwidth of the third-order ESO was calibrated through hovering balancing experiments. Controller initial bandwidth Corresponding gain , , Control Law .

[0070] Step 209: Based on the LADRC control effect and attitude error from step 208, configure the mid-level DRL adaptive parameter tuner, specifically including: PPO algorithm state space. Includes 18 dimensions (3-dimensional attitude error, 3-dimensional error derivative, 3-dimensional ESO perturbation estimation, 3-dimensional wind speed, and 6-dimensional uncertainty), and 4-dimensional motion space (perturbation). ), motion limit reward function

[0071]

[0072] ( , Online fine-tuning of the learning rate in the policy network Regular local fine-tuning is completed through step 303.

[0073] Step 210, based on the environmental prediction and attitude target from step 205, configure the top-level MPC feedforward compensation controller, specifically including: prediction time domain Seconds, the distance from the walk Seconds (200 steps in total), cost function

[0074]

[0075] ( , The wind field prediction uses a lightweight LSTM (64 hidden units, 2 layers). It takes the wind speed sequence of the past 2 seconds (20 sampling points) as input and outputs the mean and variance of the wind speed for the next 2 seconds. MPC adjusts the target attitude in advance based on this.

[0076] This embodiment combines the heterogeneous noise characteristics of multiple sensors with the requirements for identifying complex flight environments to construct an adaptive preprocessing and hierarchical control framework, which improves the limitations of traditional control strategies in terms of weak generalization ability and difficulty in adapting to diverse environments, and enhances the anti-interference ability and dynamic adaptability of UAV attitude control.

[0077] In a preferred embodiment of the present invention, step 3 (self-optimization layer), based on the control commands and flight status data obtained in step 2, completes data storage and feature extraction, federated learning model evolution, and secure OTA updates to achieve continuous optimization and stable iteration of the control strategy, and may include:

[0078] Step 301: Based on the attitude error obtained in Step 1 and the control commands in Step 2, flight data is recorded during system operation. Specifically, this step is executed synchronously in the flight background, triggered by the attitude error. Control command saturation ( (range) or sudden changes in wind speed At that time, a 30-second data segment is recorded to eMMC storage.

[0079] Step 302: Based on the flight data recorded in step 301, feature extraction and dimensionality reduction are performed using the incremental PCA algorithm. Specifically, this includes: recursively updating the sample mean and covariance matrix every 100 accumulated samples, and extracting the cumulative variance contribution rate. Principal components, compressing data volume.

[0080] Step 303: Based on the dimensionality-reduced feature data from Step 302, complete the local fine-tuning training of the DRL policy network. Specifically, this includes: triggering incremental learning after a single UAV has accumulated 100 hours of flight time, extracting feature vectors related to the "environment recognition-weight allocation" mapping relationship from historical flight data, fine-tuning the parameters of the last two layers of the DRL policy network based on the recorded data, and setting the learning rate to [missing information]. A batch size of 256 was used, with 10 training rounds. This improved the system's control adaptability in dynamic environments.

[0081] Step 304: Based on the locally trained gradients from step 303, upload them to the cloud using a privacy protection mechanism. This specifically includes adding Gaussian noise to the gradients before uploading. ( To achieve differential privacy, Paillier homomorphic encryption is used to encrypt gradients, allowing aggregation to be performed in the cloud without decryption.

[0082] Step 305: Based on the encrypted gradients uploaded by multiple drones, a global optimization model is obtained through a cloud aggregation algorithm. Specifically, this includes: the server using the FedAvg algorithm to aggregate gradients and aggregate weights. Local data quality score Proportional to this, the weight of high-quality drones is increased by 30%, the global model update cycle is 7 days, and an OTA-based breakpoint resume mechanism is adopted.

[0083] Step 306: Based on the global optimization model distributed from the cloud, prepare for the update using the A / B partitioning mechanism. Specifically, the system is divided into two 32 GB partitions. Partition A runs in flight mode, while partition B silently receives the update. The update package is written to partition B after being encrypted with AES-256 and verified with RSA-2048 signature.

[0084] Step 307, based on the update package from step 306, performs static stability verification during the ground-based standby period. Specifically, this includes: preloading and verifying the new model parameters for partition B when the UAV is detected to be in a ground-locked state and the motors are not rotating; and calculating the derivatives of the Lyapunov functions offline. Check the change in controller gain And the observer bandwidth If verification fails, the system will automatically roll back immediately without performing a switch. Complete status logs, 100 ms before and after preloading, will be uploaded to the cloud.

[0085] This embodiment combines flight data value screening with the privacy protection requirements of federated learning to build a self-optimizing iterative framework, which improves the limitations of traditional control strategies that are difficult to continuously evolve and have weak adaptability to new scenarios, thereby improving the long-term stability and intelligent iteration efficiency of UAV control performance.

Claims

1. A self-optimizing stabilization system for unmanned aerial vehicles (UAVs) based on multimodal perception, characterized in that, include: A multimodal sensing layer is used for hardware synchronization, spatiotemporal registration, and tightly coupled state estimation of heterogeneous sensors. The intelligent decision-making layer is used for adaptive noise reduction preprocessing of multi-sensor data, flight environment identification, dynamic weight allocation, and hierarchical attitude control. The self-optimization layer is used for compressed storage of flight data, federated learning-driven evolution of control strategies, and secure OTA updates. The multimodal perception layer outputs the aircraft attitude, environmental disturbance torque, and sensor health status information to the intelligent decision layer. The intelligent decision layer generates the optimal attitude control command, and the self-optimization layer continuously optimizes the control strategy of the intelligent decision layer based on flight data.

2. The system according to claim 1, characterized in that, The multimodal sensing layer includes: The hardware synchronization module uses a precise time protocol to calibrate the clock deviation and compensate for the transmission delay of the IMU, LiDAR, black light binocular depth camera and anemometer. The tightly coupled state estimation module constructs a state vector containing attitude quaternions, angular velocity, and sensor zero bias. It integrates visual reprojection error, IMU pre-integration residual, and point-to-plane geometric constraints of the lidar. The state increment is solved through a weighted least squares framework. The information matrix dynamically adjusts the lidar observation covariance according to the sparsity of environmental features.

3. The system according to claim 2, characterized in that, The multimodal sensing layer also includes: The gravity reference observation module filters the ground point cloud of lidar and extracts the normal vector in near-ground flight mode, and calculates the pitch angle and roll angle as an independent attitude reference when vision fails. The disturbance torque calculation module estimates the environmental disturbance torque based on the three-dimensional wind speed and air density measured by the anemometer through an aerodynamic parameter model, and feeds it into the intelligent decision layer for feedforward compensation.

4. The system according to claim 1, characterized in that, The intelligent decision-making layer includes: The adaptive denoising preprocessing module employs an improved wavelet threshold denoising algorithm for IMU data, introducing an exponential transition region into the threshold function and adaptively calculating it through unbiased risk estimation; it uses statistical outlier filtering and voxel downsampling for LiDAR point clouds; and it uses variance-stabilized transformation and improved nonlocal mean filtering for low-light images, with the filtering parameters adaptively adjusted according to the local noise level.

5. The system according to claim 4, characterized in that, The intelligent decision-making layer also includes: The environment recognition and dynamic weight allocation module uses a lightweight convolutional neural network to input multi-sensor time-series data and output a flight environment probability vector. The dynamic weight generator calculates the weight of each sensor in real time according to the environment probability, introduces an environment change rate penalty term to prevent weight jumps, and achieves adaptive scheduling of sensor resources after smoothing filtering.

6. The system according to claim 5, characterized in that, The intelligent decision-making layer also includes: The hierarchical attitude control module includes a bottom-level linear active disturbance rejection controller, a middle-level deep reinforcement learning parameter tuner, and a top-level model prediction controller. The underlying controller establishes a second-order motion model for each attitude channel, uses an extended state observer to estimate the total disturbance in real time, and achieves disturbance compensation by tuning the gain using the bandwidth method. The mid-level parameter tuner is based on the near-end policy optimization algorithm, with attitude error, disturbance estimation and environmental parameters as the state space, and outputs the perturbation of the parameters of the bottom-level controller. The top-level controller generates attitude feedforward commands based on wind field prediction, and the cost function includes attitude error, angular acceleration penalty and terminal cost in the prediction time domain.

7. The system according to claim 1, characterized in that, The self-optimization layer includes: The federated learning evolution module uses a cloud server to weight and aggregate the encrypted model gradients uploaded by each drone. The weights are proportional to the local data quality score. After updating the global model, the data is distributed via over-the-air download technology to achieve continuous evolution of the control strategy and scenario generalization.

8. The system according to claim 7, characterized in that, The self-optimization layer also includes: The secure OTA update module adopts an A / B partition dual backup mechanism, with the A partition algorithm running in flight mode and the B partition silently receiving updates; Before the update, the source credibility is verified by digital signature and hash value, and the range of new parameters is constrained by Lyapunov stability theory to ensure that the change in controller gain and observer bandwidth after the update does not exceed the upper limit of stability margin.

9. The system according to claim 8, characterized in that, The secure OTA update module performs partition switching during standby periods. If the attitude error exceeds the limit or the Lyapunov function diverges after the switch, the system will automatically roll back to the previous version within a set time.

10. A self-optimizing stabilization method for unmanned aerial vehicles based on multimodal perception, implemented using the system described in any one of claims 1-9, characterized in that, include: Step 1: Multimodal sensing is achieved through hardware synchronization and tight coupling optimization, and the body attitude, environmental disturbance torque and sensor health status are output. Step 2: Based on the environmental recognition results, perform adaptive denoising, dynamic weight allocation, and hierarchical anti-disturbance control to generate the optimal attitude control command; Step 3: Record flight data and compress features. Through federated learning and evolutionary control strategies, the system achieves self-optimization through safe OTA updates and online stability verification, and automatically rolls back in case of anomalies.