GPS-free unmanned aerial vehicle autonomous navigation method based on multi-modal perception and reinforcement learning

By employing a collaborative approach combining cross-modal perception and reinforcement learning, the problem of insufficient perception of transparent or reflective materials by UAVs in GPS-free environments was solved, enabling reliable obstacle identification and safe detour, thus improving the robustness and safety of navigation.

CN121346798AInactive Publication Date: 2026-01-16HUNAN UNIV OF SCI & ENG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511429121.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-08
Publication Date
2026-01-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing UAV autonomous navigation technologies lack sufficient perception of transparent or reflective materials in environments without global satellite positioning signals. Traditional single-channel grid occupancy is prone to misjudgment, safety control lacks local plane normal differentiation, and active perception and strategy optimization are insufficient, resulting in obstacle avoidance failure, slow convergence in uncertain areas, and high risk exposure.

Method used

A cross-modal perception and reinforcement learning collaborative approach is adopted. By calculating inconsistency and uncertainty through a cross-modal fusion network, a dual-channel occupancy model and anisotropic safety control are constructed. The local plane normal and normal uncertainty are estimated, upper limits for normal and tangential velocities are set, active verification and Bayesian updates are performed, and safe reinforcement learning is realized.

Benefits of technology

It enhances the ability to reliably identify and safely bypass transparent or reflective obstacles, quickly eliminates uncertainties, improves navigation robustness and safety, and adapts to complex environments without GPS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121346798A_ABST
    Figure CN121346798A_ABST
Patent Text Reader

Abstract

The invention discloses a GPS-free unmanned aerial vehicle autonomous navigation method based on multi-modal perception and reinforcement learning, and aims to solve the problem of obstacle avoidance failure caused by insufficient perceptibility of a transparent or reflective material. The method comprises the following steps: constructing a dual-channel occupancy graph containing an entity occupancy probability and a permeable layer probability, estimating a local plane normal in a suspected area and establishing an anisotropic safety barrier aligned with the normal, setting normal and tangential speed upper limits, combining safety reinforcement learning active verification and sensor parameter self-adaption, and cooperating with Bayesian updating and local path planning to obtain an anisotropic safety barrier. Reliable identification and safe detour of transparent or reflective obstacles are realized, the collision risk is reduced, and the robustness and stability of autonomous navigation in an environment without global satellite positioning signals are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) navigation, and more particularly to a GPS-free UAV autonomous navigation method based on multimodal perception and reinforcement learning. Background Technology

[0002] In the field of autonomous navigation for unmanned aerial vehicles (UAVs) in environments without global satellite positioning signals, existing technologies generally use visual and inertial odometry or lidar odometry for positioning, combined with grid maps or semantic maps for local path planning. To improve environmental understanding, multi-sensor fusion is gradually becoming the mainstream. Common combinations include cameras, lidar, inertial measurement units, millimeter-wave radar and polarimetric imaging, and deep learning is introduced for cross-modal feature fusion and uncertainty estimation.

[0003] For the detection of transparent or reflective materials, the study explored methods such as polarization imaging, radar echo characteristics, and multi-view geometry. However, the systematic integration of these methods into the navigation closed loop remains limited. In terms of risk constraints and safety control, cost planning and velocity limiting based on occupancy grids are mostly adopted. Some works have attempted to use reinforcement learning for exploration, but the combination of safety constraints and active perception is insufficient.

[0004] However, existing technologies still have shortcomings:

[0005] 1. Insufficient perceptibility of transparent or reflective materials: The contradictory information between multimodal modes is not quantitatively characterized and utilized. Traditional single-channel grid occupancy is prone to misjudging transparent or mirrored areas as passable or open, leading to obstacle avoidance failure.

[0006] 2. Safety controls are mostly isotropic uniform safety margins and speed limits: differentiated normal and tangential safety constraints are not established in conjunction with local plane normals, which can easily lead to excessively aggressive normal directions or excessively conservative tangential directions in areas that are suspected to be transparent or reflective.

[0007] 3. Insufficient proactive perception and strategy optimization: Reinforcement learning focuses more on trajectory optimization and lacks coordination with sensor parameter adaptation. It does not trigger verification and closed-loop Bayesian updates based on local uncertainties, resulting in slow convergence in uncertain regions and high risk exposure.

[0008] Therefore, there is an urgent need for a navigation method that can quantitatively utilize cross-modal inconsistencies and fusion uncertainties under multimodal conditions, apply anisotropic safety constraints to transparent or reflective materials, and collaborate with active verification through safety reinforcement learning. Summary of the Invention

[0009] One objective of this invention is to propose an autonomous navigation method for unmanned aerial vehicles (UAVs) without GPS signals, based on multimodal perception and reinforcement learning collaboration. Addressing the shortcomings of existing technologies, such as insufficient perception of transparent or reflective materials, limited occupancy representation, and lack of active verification leading to obstacle avoidance failures, this invention proposes a dual-channel occupancy modeling and anisotropic safety control approach centered on cross-modal inconsistency and fusion uncertainty. This involves using a cross-modal fusion network to output inconsistency and uncertainty, recursively updating entity occupancy probabilities and permeable layer probabilities; estimating local plane normals and normal uncertainties in suspected areas, constructing an anisotropic safety barrier aligned with the normal, setting upper limits for normal and tangential velocities, and generating a set of possible actions; actively verifying actions using safety reinforcement learning under trigger conditions, combining sensor parameter adaptation and a safety constraint layer to select the action to execute; and performing Bayesian updates based on new observations after execution to complete local path planning. This invention offers the technical advantages of reliable identification and safe detour around transparent or reflective obstacles, closed-loop suppression and accelerated resolution of uncertainties, and improved navigation robustness and safety in environments without GPS signals.

[0010] An autonomous navigation method for GPS-free unmanned aerial vehicles based on multimodal perception and reinforcement learning according to an embodiment of the present invention is characterized by comprising the following steps:

[0011] S1. Collect multimodal raw observation data from at least two sensors, perform time synchronization and spatial calibration, estimate the UAV pose using visual inertial odometry, construct a local grid map with the UAV as a reference under the UAV pose, and output time-aligned multimodal observation data, UAV pose and local grid map.

[0012] S2. Using time-aligned multimodal observation data, UAV pose, and local grid map as input, a deep neural network that fuses cross-modal features is used to calculate and output the cross-modal inconsistency and fusion uncertainty.

[0013] S3. Take multimodal observation data with cross-modal inconsistency, fusion uncertainty, and time alignment, UAV pose and local grid map as input, construct a dual-channel occupancy map on the local grid map and calculate the entity occupancy probability and permeability probability grid by grid and output it.

[0014] S4. Take multimodal observation data, UAV pose and local grid map as input, including entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty and time alignment, UAV pose and local grid map, and estimate and output the local plane normal and normal uncertainty based on multimodal observation data in the suspected transparent or reflective area where the permeability probability and cross-modal inconsistency meet the preset conditions.

[0015] S5. Using local plane normal, normal uncertainty, entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose, and local grid map as input, construct an anisotropic safety barrier aligned with the local plane normal. Calculate the normal safety margin and tangential safety margin at grid locations within suspected transparent or reflective areas. Set upper limits for normal velocity and tangential velocity according to the anisotropic safety barrier to obtain a set of possible actions consistent with the anisotropic safety barrier. Output the anisotropic safety barrier parameters, upper limits for normal velocity, upper limits for tangential velocity, and set of possible actions.

[0016] S6. Taking anisotropic safety barrier parameters, upper limit of normal velocity, upper limit of tangential velocity, set of possible actions, local plane normal, normal uncertainty, entity occupancy probability, layer penetration probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose and local grid map as input, call reinforcement learning policy network for active verification, generate a set of candidate actions containing flight actions and sensor parameter configuration based on the input and output it;

[0017] S7. The candidate action set, anisotropic safety barrier parameters, normal velocity limit, tangential velocity limit and feasible action set are taken as input. The safety constraint layer constrains and makes the candidate actions feasible according to the anisotropic safety barrier and the velocity limit. The action to be executed is selected in the feasible action set and the execution sensor parameter configuration is matched. The action to be executed and the execution sensor parameter configuration are output.

[0018] S8. Taking the executed action, the executed sensor parameter configuration, the entity occupancy probability, the permeability probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multimodal observation data, the UAV pose, and the local grid map as input, the UAV is controlled to acquire new observations according to the executed action and the executed sensor parameter configuration. Based on the new observations and combined with the cross-modal inconsistency and the fusion uncertainty, the entity occupancy probability and the permeability probability are updated using Bayesian method. Based on the updated entity occupancy probability and the permeability probability, local path planning is performed on the local grid map, and navigation control commands are generated and output.

[0019] Optionally, step S1 specifically includes:

[0020] Collect multimodal raw observation data from at least two of the following sensors: camera, lidar, inertial measurement unit, millimeter-wave radar, and polarimetric camera; establish a unified time reference and align the timestamps of each sensor; estimate and compensate for time offset and time delay to obtain time-aligned multimodal observation data.

[0021] Spatial calibration is performed on each sensor to obtain the intrinsic parameters of each sensor and the extrinsic parameters relative to the coordinate system with the UAV as the reference. The time-aligned multimodal observation data is then transformed to the coordinate system with the UAV as the reference.

[0022] A visual-inertial odometry system with a tightly coupled camera and inertial measurement unit is used to fuse the transformed observations to output the UAV pose.

[0023] Using the UAV pose as a reference, the spatially calibrated, time-aligned multimodal observation data is projected and rasterized to construct a local raster map with the UAV as a reference.

[0024] Output time-aligned multimodal observation data, UAV pose, and local grid map.

[0025] Optionally, step S2 specifically includes:

[0026] The time-aligned multimodal observation data, UAV pose, and local grid map are used as inputs. Based on the UAV pose, the time-aligned multimodal observation data are projected onto the grid coordinate system of the local grid map, and a feature representation is generated for each sensor in each grid. The feature representation includes at least indicators reflecting geometric relationships and appearance information.

[0027] Using the grid-level feature representation and the local grid map as input, a deep neural network for cross-modal feature fusion is invoked to evaluate the observation consistency of different sensors within the same grid and estimate the reliability of the fusion result. Cross-modal inconsistency and fusion uncertainty are calculated and output on each grid of the local grid map.

[0028] Optionally, step S3 specifically includes:

[0029] Using multimodal observation data with cross-modal inconsistency, fusion uncertainty, and time alignment, UAV pose, and local grid map as input, a dual-channel occupancy state is established in each grid of the local grid map, namely entity occupancy probability and layer penetration probability. With prior probabilities as initial values, the time-aligned multimodal observation data is projected onto the corresponding grid under the UAV pose to form observation evidence. The entity occupancy probability and the layer penetration probability are recursively updated according to the probability update rule. Cross-modal inconsistency is used to adjust the update direction and weight of the two probabilities, so that the update weight of the layer penetration probability increases with the increase of cross-modal inconsistency and the update weight of the entity occupancy probability decreases with the increase of cross-modal inconsistency. Fusion uncertainty is used to adjust the update gain so that the update gain of the two probabilities decreases when the fusion uncertainty increases, so as to suppress unreliable observations.

[0030] After updating the entire raster, output the entity occupancy probability and the permeability probability.

[0031] Optionally, step S4 specifically includes:

[0032] The system takes multimodal observation data, UAV pose, and local grid map as input, including entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty, and time alignment. Based on the preset conditions that the permeability probability and cross-modal inconsistency meet, it selects suspected transparent or reflective areas from the local grid map as areas to be estimated.

[0033] Features for plane estimation are extracted from time-aligned multimodal observation data within the region to be estimated. These features include at least one of the following: spatial points and their echo properties provided by lidar or millimeter-wave radar, edges and intensity gradients provided by a camera, and motion consistency measures provided by a vision and inertial measurement unit.

[0034] A local plane model is established in each region to be estimated based on the aforementioned features, and a robust fit is performed to obtain the local plane normal.

[0035] The normal uncertainty is calculated based on the fitting residual, the number of supporting observations, the cross-modal inconsistency, and the fusion uncertainty, so that the normal uncertainty increases accordingly when the fusion uncertainty increases.

[0036] Output local plane normal and normal uncertainty.

[0037] Optionally, step S5 specifically includes:

[0038] Using local plane normal, normal uncertainty, entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose, and local grid map as input, suspected transparent or reflective areas are selected in the local grid map according to the permeability probability and cross-modal inconsistency satisfying preset conditions, and an anisotropic safety barrier aligned with the local plane normal is constructed.

[0039] The anisotropic safety barrier defines the normal direction with the local plane normal and the tangential direction with the direction perpendicular to the local plane normal. For each grid in the suspected transparent or reflective area, the normal safety margin and the tangential safety margin are calculated respectively. The normal safety margin is determined according to the weighted relationship of fusion uncertainty, entity occupancy probability and permeability probability, so that the normal safety margin increases when the fusion uncertainty increases, increases when the entity occupancy probability increases, and decreases when the permeability probability increases. The tangential safety margin is determined according to the monotonic relationship of the permeability probability, so that the tangential safety margin increases when the permeability probability increases.

[0040] The speed of the UAV is decomposed based on the anisotropic safety barrier, and upper limits for normal speed and tangential speed are set respectively. The upper limit for normal speed decreases with the increase of fusion uncertainty and entity occupancy probability and increases with the increase of layer penetration probability. The upper limit for tangential speed increases with the increase of layer penetration probability.

[0041] By combining the dynamic constraints of the UAV, the anisotropic safety barrier, the normal safety margin, the tangential safety margin, the upper limit of normal velocity, and the upper limit of tangential velocity, a set of possible actions that satisfy the anisotropic safety barrier is generated, and the anisotropic safety barrier parameters, the upper limit of normal velocity, the upper limit of tangential velocity, and the set of possible actions are output.

[0042] Optionally, step S6 specifically includes:

[0043] The system takes anisotropic safety barrier parameters, upper limit of normal velocity, upper limit of tangential velocity, set of possible actions, local plane normal, normal uncertainty, entity occupancy probability, layer penetration probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose, and local grid map as input. It determines whether to trigger active verification based on the comparison result of cross-modal inconsistency or fusion uncertainty ahead of the planned trajectory with a preset threshold. The preset threshold can be adaptively adjusted according to fusion uncertainty and mission risk.

[0044] When active verification is triggered, the local state, consisting of entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty, local plane normal, normal uncertainty, and upper limits of normal and tangential velocities, is used as the input to the reinforcement learning policy network. The reinforcement learning policy network is a trained policy network that can be trained or updated before or during operation. Its goal is to improve the confidence in entity occupancy probability and permeability probability and reduce local uncertainty under the conditions of satisfying anisotropic safety barriers and velocity upper limit constraints. The output includes a set of candidate actions containing flight actions and sensor parameter configurations. The flight actions are used to generate small changes in local position or attitude, and the sensor parameter configurations are used to adjust the sampling characteristics of the sensors to enhance the observation of suspected transparent or reflective areas.

[0045] Optionally, step S7 specifically includes:

[0046] The candidate action set, anisotropic safety barrier parameters, upper limit of normal velocity, upper limit of tangential velocity, and set of feasible actions are taken as input, and the safety constraint layer is invoked to perform constraint judgment and feasibility processing on the candidate action set.

[0047] The constraint judgment includes detecting the displacement and velocity decomposition of each candidate action based on the anisotropic safety barrier parameters and the local grid map, determining whether its distance in the normal direction is not less than the corresponding normal safety margin, and determining whether its motion in the tangential direction is within the allowable range of the corresponding tangential safety margin. At the same time, the normal velocity component and tangential velocity component of each candidate action are limited according to the upper limit of normal velocity and the upper limit of tangential velocity, so that the normal velocity component does not exceed the upper limit of normal velocity and the tangential velocity component does not exceed the upper limit of tangential velocity.

[0048] The feasibility processing includes eliminating candidate actions that violate constraints or mapping within the set of feasible actions based on the principle of minimum modification to obtain alternative actions that satisfy the constraints, and performing consistent constraint checks and limiting on the scheduling of sensor parameters included in the candidate actions to ensure consistency with the sensor parameter constraints in the set of feasible actions.

[0049] After completing the constraint judgment and feasibility processing of all candidate actions, the feasible actions that satisfy the anisotropic safety barrier and the speed limit are selected as candidates. The action to be executed is selected and matched with the corresponding sensor parameter configuration, and the execution action and execution sensor parameter configuration are output.

[0050] Optionally, step S8 specifically includes:

[0051] The system takes the following as inputs: the execution action, the execution sensor parameter configuration, the entity occupancy probability, the permeability probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multimodal observation data, the UAV pose, and the local grid map. It controls the UAV to execute the execution action and drives the sensors to collect new observations according to the execution sensor parameter configuration.

[0052] The newly added observations are time-synchronized and projected onto the local grid map based on the UAV pose to form observational evidence for updating;

[0053] Using the entity occupancy probability and the permeable layer probability as priors, the observation evidence is weighted in combination with the cross-modal inconsistency and the fusion uncertainty. The entity occupancy probability and the permeable layer probability are then updated using Bayesian updates according to the probability update rules to obtain the updated entity occupancy probability and the updated permeable layer probability.

[0054] Based on the updated entity occupancy probability and the updated permeability probability, local path planning is performed on the local grid map. The planning cost function simultaneously penalizes the reduction of the entity occupancy probability and the permeability probability to prioritize areas with high permeability probability and low entity occupancy probability and implement detours for areas with high entity occupancy probability, thereby generating and outputting navigation control commands.

[0055] The beneficial effects of this invention are:

[0056] (1) Improve the ability to perceive and distinguish transparent or reflective obstacles: By using dual-channel occupancy modeling guided by cross-modal inconsistency and fusion uncertainty, the entity occupancy probability and the permeable layer probability are expressed separately, and the update direction and gain are adaptively adjusted, which significantly reduces the risk of misjudging transparent or mirror areas as passable or empty, and improves the obstacle avoidance success rate.

[0057] (2) Safe and efficient motion control: Estimate the local plane normal and normal uncertainty in the suspected area, construct an anisotropic safety barrier aligned with the normal, and set the upper limit of normal velocity and the upper limit of tangential velocity respectively, so that the normal is more conservative and the tangential is more flexible, reducing unnecessary detours and sudden stops while ensuring safety margin, and taking into account both safety and efficiency.

[0058] (3) Closed-loop navigation that quickly resolves uncertainties: Based on reinforcement learning active verification triggered by uncertainty, combined with sensor parameter adaptation and safety constraint layer to select actionable actions, Bayesian update is performed using new observations, so that local uncertainties can be quickly converged, improving the stability of local path planning and the robustness of overall navigation, and adapting to complex environments without global satellite positioning signals. Attached Figure Description

[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0060] Figure 1 This is a flowchart of an autonomous navigation method for GPS-free unmanned aerial vehicles based on multimodal perception and reinforcement learning proposed in this invention. Detailed Implementation

[0061] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0062] refer to Figure 1 A GPS-free UAV autonomous navigation method based on multimodal perception and reinforcement learning, characterized by the following steps:

[0063] S1. Collect multimodal raw observation data from at least two sensors, perform time synchronization and spatial calibration, estimate the UAV pose using visual inertial odometry, construct a local grid map with the UAV as a reference under the UAV pose, and output time-aligned multimodal observation data, UAV pose and local grid map.

[0064] S2. Using time-aligned multimodal observation data, UAV pose, and local grid map as input, a deep neural network that fuses cross-modal features is used to calculate and output the cross-modal inconsistency and fusion uncertainty.

[0065] S3. Take multimodal observation data with cross-modal inconsistency, fusion uncertainty, and time alignment, UAV pose and local grid map as input, construct a dual-channel occupancy map on the local grid map and calculate the entity occupancy probability and permeability probability grid by grid and output it.

[0066] S4. Take multimodal observation data, UAV pose and local grid map as input, including entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty and time alignment, UAV pose and local grid map, and estimate and output the local plane normal and normal uncertainty based on multimodal observation data in the suspected transparent or reflective area where the permeability probability and cross-modal inconsistency meet the preset conditions.

[0067] S5. Using local plane normal, normal uncertainty, entity occupancy probability, permeability probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose, and local grid map as input, construct an anisotropic safety barrier aligned with the local plane normal. Calculate the normal safety margin and tangential safety margin at grid locations within suspected transparent or reflective areas. Set upper limits for normal velocity and tangential velocity according to the anisotropic safety barrier to obtain a set of possible actions consistent with the anisotropic safety barrier. Output the anisotropic safety barrier parameters, upper limits for normal velocity, upper limits for tangential velocity, and set of possible actions.

[0068] S6. Taking anisotropic safety barrier parameters, upper limit of normal velocity, upper limit of tangential velocity, set of possible actions, local plane normal, normal uncertainty, entity occupancy probability, layer penetration probability, cross-modal inconsistency, fusion uncertainty, time-aligned multimodal observation data, UAV pose and local grid map as input, call reinforcement learning policy network for active verification, generate a set of candidate actions containing flight actions and sensor parameter configuration based on the input and output it;

[0069] S7. The candidate action set, anisotropic safety barrier parameters, normal velocity limit, tangential velocity limit and feasible action set are taken as input. The safety constraint layer constrains and makes the candidate actions feasible according to the anisotropic safety barrier and the velocity limit. The action to be executed is selected in the feasible action set and the execution sensor parameter configuration is matched. The action to be executed and the execution sensor parameter configuration are output.

[0070] S8. Taking the executed action, the executed sensor parameter configuration, the entity occupancy probability, the permeability probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multimodal observation data, the UAV pose, and the local grid map as input, the UAV is controlled to acquire new observations according to the executed action and the executed sensor parameter configuration. Based on the new observations and combined with the cross-modal inconsistency and the fusion uncertainty, the entity occupancy probability and the permeability probability are updated using Bayesian method. Based on the updated entity occupancy probability and the permeability probability, local path planning is performed on the local grid map, and navigation control commands are generated and output.

[0071] In this specific embodiment, S1 specifically refers to:

[0072] System selection of sensor set Then, at least two types (e.g., camera, lidar, inertial measurement unit, millimeter-wave radar, polarimetric camera) are selected for fusion. Using the UAV's body coordinate system as a reference, a unified time reference is first established, and the timestamps of each sensor are corrected and aligned: For the k-th observation of the i-th sensor, the timestamp is corrected as follows:

[0073]

[0074] Where t i (k) represents the original timestamp of the i-th sensor at index k. Represents the synchronization timestamp relative to a unified reference clock, Δt i τ represents the static time offset of the i-th sensor relative to the reference clock. i Let represent the equivalent time delay introduced by the data acquisition and transmission of the i-th sensor, k represent the index of the observation sequence of this sensor, and i represent the sensor number in the sensor set. This represents the set of sensors participating in the fusion, and N represents the number of sensors in that set.

[0075] Subsequently, spatial calibration is performed to obtain the extrinsic parameters of each sensor relative to the body coordinate system, and the observations of each sensor are unified to the body coordinate system. Specifically, the rigid body transformation from the sensor coordinate system to the body coordinate system adopts the following relationship:

[0076]

[0077] Where X i Represents a spatial point X in homogeneous form in the coordinate system of the i-th sensor. b Represents the homogeneous coordinates of the same spatial point in the body coordinate system. This represents the three-dimensional pose transformation matrix that maps the coordinate system of the i-th sensor to the body coordinate system;

[0078] After completing time synchronization and spatial calibration, a tightly coupled visual-inertial odometry system (VIO) combining the camera and inertial measurement unit is used to fuse the synchronized and transformed observations, outputting the UAV's pose in the local world coordinate system. in This represents the pose transformation at time t, which maps the body coordinate system to the world coordinate system; t represents the continuous time variable; w represents the world coordinate system label; and b represents the body coordinate system label.

[0079] Using this pose as a reference, the calibrated time-aligned observations are projected and rasterized to construct a local raster map. Its raster index is determined by the following relationship:

[0080]

[0081] Where u represents the row and column index vector of the 2D raster, x represents the 2D coordinates of the point projected onto the map plane, x0 represents the 2D coordinates of the local raster map origin on the plane, and δ represents the raster resolution. This represents the floor operator for each component of a vector. Represents a local raster map with reference to the drone at time t;

[0082] The final output is time-aligned multimodal observation data. UAV pose With local raster map in This represents the set of multimodal observations at time t after undergoing a unified time base and spatial transformation. The symbol represents the original multimodal observation set, and the symbol ~ represents the notation after applying time and space corrections to the set.

[0083] In this specific embodiment, S2 specifically refers to:

[0084] The system takes time-aligned multimodal observations, UAV pose, and local grid map as input. It first projects the observations onto the local grid coordinate system based on the UAV pose and extrinsic parameters of each sensor, and then generates sensor feature representations for each grid cell. For this purpose, a projection function is defined as follows:

[0085] g = π i,t (o i,j );

[0086] Where g represents the raster index of the hit, π i,t (·) represents the projection mapping of the grid observed by the i-th sensor at time t, o i,j Let represent the j-th original observation from the i-th sensor, i represent the sensor index, j represent the observation sequence number, t represent the current processing time, and (·) represent the placeholder for the function's independent variable;

[0087] Aggregating homogeneous observations falling within the same grid to obtain the feature vector for each sensor in each grid, using:

[0088] φ g,i =pool({h(o i,j )∣π i,t (o i,j )=g});

[0089] Where φ g,i The i-th sensor represents the feature vector in grid g; pool(·) represents the pooling operator used to statistically aggregate elements within the set; h(·) represents the observation coding function that maps a single observation to a fixed-length descriptor; {·} represents the set of elements that satisfy the conditions; and | represents the condition qualifier within the set.

[0090] Subsequently, raster-level cross-modal features are constructed and input into a cross-modal feature fusion network for consistency evaluation and credibility estimation, denoted as:

[0091]

[0092] Where Φ g The concatenation vector represents the cross-modal feature concatenation of grid g; concat(·) represents the channel-wise concatenation operator. The set of sensor indices participating in the fusion, d g The cross-modal inconsistency is used to measure the degree of conflict in observations by different sensors on grid g, σ g The fusion uncertainty is used to measure the confidence level of the fusion output, f. θ (·) represents a deep neural network mapping with learnable weights θ, c g θ represents contextual features from a local raster map, such as location encoding and neighborhood structure statistics, and θ represents the set of network parameters.

[0093] In implementation, the observation code h(·) can simultaneously include both geometric and appearance metrics to cover distinguishing clues for transparent or reflective materials, such as depth and parallax statistical echoes or reflection intensity polarization-related parameters, edge and gradient direction optical flow, and motion consistency measures, etc., and the fusion network f θ (·) A deep architecture incorporating cross-modal attention and uncertainty branches can be adopted, and pre-trained and fine-tuned in-orbit using supervised or self-supervised methods, ultimately outputting {d} in each grid cell. g ,σ g This is used for the dual-channel occupancy update in step S3.

[0094] In this specific embodiment, S3 specifically refers to:

[0095] The system takes cross-modal inconsistency and fusion uncertainty, time-aligned multimodal observations, UAV pose, and local grid maps as inputs, and operates on a local grid set. For each grid cell, a state with two channels, "entity occupancy" and "permeable layer," is established and expressed recursively using logarithmic probability. By adjusting the update weights of the two channels based on cross-modal inconsistency with each new observation and suppressing update gain based on fusion uncertainty, the permeable layer receives a larger update under high inconsistency, while entity occupancy remains cautious under high uncertainty. Initialization uses:

[0096] With recursion using L c (g,t+1)=L c (g,t)+γ(σ g )w c (d g )l c (e g (t) and finally mapped to probability.

[0097] in The set of indices representing local graticules is used to enumerate graticules on the map; g represents the index of a single graticule within it, used to locate the update unit; c∈{occ,trans} represents the state channel, where occ is the entity occupancy channel and trans is the permeable layer channel; L c (g,t) represents the logarithmic probability state of channel c of raster g at time t, used for recursive calculation. The prior probability of channel c of grid g is used for initialization; ln(·) represents the natural logarithm operator used to construct the log odds; t represents the discrete-time index used to identify the update step; γ(σ) represents the prior probability of channel c of grid g ... g )=exp(-βσ g ) represents the update gain adjusted by fusion uncertainty, where σ g The fusion uncertainty of the raster g is used to measure the confidence level of the fusion output, while β>0 represents the gain scaling factor, w occ (d g )=1-αd g with w trans (d g )=αd g The update weights of the two channels are represented by d. g The cross-modal inconsistency of grid g is used to measure the degree of observational conflict between different sensors on that grid, while α∈[0,1] represents the weight slope coefficient, l c (e g (t) represents the log-likelihood increment from the observation evidence of grid g at time t to channel c, used to apply the evidence to the channel, e. g(t) represents the evidence vector at time t projected from time-aligned multimodal observations and aggregated onto grid g to support recursive updates; exp(·) represents the exponential function used to define the gain; P c (g,t+1) represents the probability of channel c of grid g at time t+1, used to output the dual-channel occupancy map.

[0098] In this specific embodiment, S4 specifically refers to:

[0099] The system uses probability and metrics, combined with time-aligned multimodal observations, UAV pose, and local grid maps, to perform local plane normal estimation and normal uncertainty assessment. Firstly, it performs local grid set estimation... Based on the criteria of transparency and inconsistency, suspected transparent or reflective areas are filtered to construct a set:

[0100]

[0101] in A grid set representing a suspected region The set of indices for a local raster, g represents a single raster index within it, and P... trans (g) represents the permeability probability of grid g, d g The cross-modal inconsistency of grid g is represented by τ. t The threshold for determining the probability of penetration, τ d The threshold for determining cross-modal inconsistency;

[0102] Then for each Neighborhood aggregation is used for multimodal feature estimation in the plane and unified into a three-dimensional geometric support sample set. Here x j The j-th supporting sample represents a 3D point in the UAV reference coordinate system, which can be composed of spatial points from lidar or millimeter-wave radar, constraint points from camera back projection, or geometrically consistent points obtained by fusing visual and inertial constraints.

[0103] Establish a weighted plane model and impose a unit norm constraint on the normal vectors, then define the objective function:

[0104]

[0105] Where J(n,b) represents the weighted sum of squared residuals with respect to the plane parameters, n represents the local plane normal vector to be estimated, b represents the plane bias to be estimated, and w j Let represent the weight of the j-th supporting sample, ∑ represent the operator for summing over set indices, and (·) represent the weight of the j-th supporting sample. T The vector transpose operator is represented by ||·||2, which represents the L2 norm.

[0106] To demonstrate the effectiveness of different sensor reliability in suppressing fusion uncertainty, the following approach was adopted:

[0107] Set weights;

[0108] in Indicates the sensor type s from the sample source j The credibility coefficient of the decision Indicates the grid to which the sample belongs, g j The fusion uncertainty, β>0 represents the scaling factor for uncertainty suppression, and exp(·) represents the exponential function;

[0109] Solving the problem of minimizing J(n,b) yields... and in Indicates the estimated local plane normal. This represents the estimated plane offset;

[0110] To quantify the uncertainty in the normal direction, we first define the residual. It uses its second-order statistic to characterize the local fit variance, and then combines the number of supports, cross-modal inconsistency, and fusion uncertainty to give a scalar uncertainty measure:

[0111]

[0112] Where u represents the normal uncertainty scalar, Representing residual scale, Indicates the number of supported samples, It represents the expression by {σg j The mean of the fusion uncertainty obtained is... Indicates by {dg j The mean value of cross-modal inconsistency and λ obtained are calculated. σ >0 and λ d >0 represents the amplification factor of fusion uncertainty and cross-modal inconsistency on normal uncertainty, respectively. This represents the square root operator, and the final output is... This is used in conjunction with u to construct anisotropic safety barriers and velocity constraints aligned with the normal in subsequent steps.

[0113] In this specific embodiment, S5 specifically includes:

[0114] Based on local plane normals, probabilities, and uncertainties, the system constructs anisotropic safety barriers aligned with the normals within grids that satisfy potential transparency or reflection conditions, and constrains feasible actions accordingly. First, the normal and tangential safety margins for each grid g are defined as follows:

[0115] S n (g)=s0+aσ σ g +a o P occ (g)-a t P trans (g) and S t (g)=t0+b t P trans (g);

[0116] Where S n (g) represents the safety distance requirement in the local plane normal direction, S t (g) represents the safety distance requirement in the tangential direction orthogonal to the normal; g represents the local grid index used to locate a specific position; s0 and t0 represent the baseline safety terms in the normal and tangential directions, respectively; a σ >0 indicates the weighting factor of fusion uncertainty on the normal safety margin, σ g The fusion uncertainty of raster g is used to measure the reliability of the fusion result, a o >0 indicates the weighting factor of entity occupancy probability on normal safety margin, P occ (g) represents the entity occupancy probability of grid g, a t >0 indicates the weight reduction factor of the permeability probability on the normal safety margin, P trans (g) represents the permeability probability of grid g, b t >0 represents the weighting factor of the permeability probability on the tangential safety margin, making σ g With P occ (g) S increases n (g) increases while P trans (g) S increases n (g) decreases and S t (g) Increase;

[0117] Then, using the unit normal vector Set a speed limit for the decomposition direction:

[0118]

[0119] in and These represent the upper limits of the normal and tangential velocities at grid point g, respectively. This indicates that the unit normal obtained from estimating the suspected region is used to define the direction, v. n0 >0 and v t0 >0 represents the upper limit of the reference in both directions, λ σ >0 indicates the suppression coefficient of fusion uncertainty on the upper limit of normal velocity, λ o >0 indicates the suppression coefficient of entity occupancy probability on the upper limit of normal velocity, λ t>0 indicates the coefficient that increases the penetration probability on the upper limit of the normal velocity, μ t >0 indicates the factor that increases the tangential velocity upper limit by the probability of penetration; exp(·) indicates that the exponential function ensures the upper limit is positive, making the normal upper limit increase with σ. g With P occ (g) increases and decreases with P trans (g) increases with P and the upper limit of the tangential direction increases with P trans (g) Increase and improve;

[0120] To perform directional decomposition and feasibility determination of velocity commands, the normal and tangential components of the velocity vector v are defined as follows:

[0121] and

[0122] Where v represents the instantaneous velocity command in the machine system or local world system, v n (v) represents the scalar component of velocity in the normal direction, v t (v) represents the modulus component of the velocity in the tangential direction. The dot product of normal and velocity is represented by |·|, the absolute value operator is represented by |·|2, and the Euclidean second norm is represented by (·). T Represents the vector transpose operator;

[0123] Based on this, a grid-level set of feasible actions consistent with anisotropic safety barriers is given:

[0124]

[0125] in This represents the set of velocity actions that satisfy the upper limits of normal and tangential velocities, where {·} represents the set symbol, and | represents the condition qualifier within the set. Ultimately, it will... As anisotropic safety barrier parameter and with and Output the candidate and execution action constraints for subsequent steps.

[0126] In this specific embodiment, S6 specifically refers to:

[0127] The system uses anisotropic safety barrier parameters, upper limits for normal and tangential velocities, and a set of possible actions, along with entity occupancy probabilities, layer penetration probabilities, and local plane normal and normal uncertainties, to constitute the prior constraints and local state information of the reinforcement learning policy network. Risk is assessed within a local window ahead of the planned trajectory to trigger active verification. The triggering rule is as follows:

[0128]

[0129] Where χ gThe trigger value of grid g is used to comprehensively characterize local risk, d g The cross-modal inconsistency of grid g is used to measure the degree of conflict between observations from different sensors, σ g The fusion uncertainty of raster g is used to measure the confidence level of the fusion output; τ represents the adaptive trigger threshold used to determine whether to proceed with active verification; τ0 represents the baseline term of the threshold; λ R >0 indicates the influence coefficient of task risk on the threshold, R represents the task risk scalar used to reflect the task's safety requirements, and λ σ >0 indicates the influence coefficient of uncertainty on the threshold. The average value of the fusion uncertainty in the neighborhood ahead of the planned trajectory is used to characterize the overall level of uncertainty;

[0130] When the triggering condition is met, a local state vector is constructed, containing elements such as the probability of occupancy and permeability, cross-modal inconsistency and fusion uncertainty, local plane normal and normal uncertainty, normal and tangential safety margins and corresponding velocity limits, and map context. This vector is then input into the policy network. During the action sampling phase, the policy uses masks or projections to ensure that the output naturally falls into the set of actionable actions given in step S5. Simultaneously, sensor parameters are linked to enhance the observation density and quality of potentially transparent or reflective areas. Candidate generation employs the following method:

[0131] (a k ,η k ) = π θ (s g ,ξ k ), k = 1, ..., K;

[0132] Where a k The k-th flight action is used to produce a small change in local position or attitude, η. k This represents the parameter configuration vector for the k-th sensor, used to adjust the sensor's sampling characteristics, π. θ (·) represents the policy network mapping with learnable weights θ, where θ represents the set of policy network parameters, and s g The local state vector of grid g is used to drive the policy to generate candidates, ξ. k The kth exploration noise is used to increase candidate diversity, k represents the candidate index, and K represents the number of candidates;

[0133] To reflect the optimization objective of "improving confidence and reducing uncertainty while meeting safety constraints," immediate reward shaping can be used:

[0134] r t =-w σ σ′ g -w u u′+w t P′trans (g)-w o P′ occ (g);

[0135] Where r t The immediate reward at time t is used to guide policy updates, w σ >0 and w u >0 represents the penalty weights for fusion uncertainty and normal uncertainty, respectively. t >0 indicates a higher reward weight for increasing the probability of penetration, w o >0 indicates a reward weight for a decrease in the probability of entity occupancy, σ′ g The expression represents the fusion uncertainty in grid g after executing the candidate and acquiring new observations, u′ represents the normal uncertainty after execution, and P′ represents the uncertainty in the normal direction after execution. trans (g) represents the probability of permeability after execution, P′ occ (g) represents the entity occupancy probability after execution, and t represents the discrete-time index used to identify the policy update step;

[0136] The final output includes a candidate set of flight actions and sensor parameter configurations, which is then passed to the subsequent safety constraint layer for feasibility screening and execution matching.

[0137] In this specific embodiment, S7 specifically refers to:

[0138] The system inputs the candidate action set, along with the anisotropic safety barrier and the velocity limit, into the safety constraint layer. It then decomposes, performs constraint checks, feasibility assessments, and selects the optimal action. The candidate set is denoted as:

[0139]

[0140] in Represents the candidate set located in grid g, a k =(v k ,Δp k ,η k ) represents the k-th candidate velocity command and displacement increment and sensor parameter configuration, K represents the number of candidates, and g represents the local grid index;

[0141] Based on the unit normal vector, the velocity and displacement are decomposed into normal and tangential components, using:

[0142]

[0143] in Represents the local planar unit normal, v n (·) and v t (·) represent the magnitudes of the velocity in the normal and tangential directions, respectively, and d. n (·) and d t(·) represents the displacement in the normal and tangential directions, respectively. T The vector transpose is used for the inner product, |·| represents the absolute value operator, and ||·|2 represents the Euclidean norm 2.

[0144] Then, based on the speed limit and safety margin, a set of constraints is given:

[0145]

[0146] in Represents the set of feasible solutions. and Representing the upper limits of the normal and tangential velocities of grid g, respectively, and S n (g) and S t (g) represent the normal and tangential safety margins of grid g, respectively. This indicates the upper and lower bounds of the allowed component levels in the sensor parameter vector;

[0147] For not satisfied Candidate execution minimum modification feasibility mapping:

[0148]

[0149] in Indicates the feasible alternatives of the k-th candidate. Represents a set The least squares projection, y represents the projection variable, and ||·|2 represents the Euclidean second norm used to measure the magnitude of the change;

[0150] After completing constraint checks and feasibility assessments for all candidates, a feasible set is obtained. And select the action k to be performed based on the comprehensive score. * =argmax k ψ k With pairing configuration

[0151] in Represents the set of possible actions, ψ k Represents the overall score of the k-th candidate, k * Indicates the selected index, a * Indicates the final flight maneuver performed, η * This indicates the final sensor parameter configuration to be executed and used for subsequent execution and verification.

[0152] In this specific embodiment, S8 specifically refers to:

[0153] The system executes action a * With sensor parameter configuration η *The drone is driven to execute and collect new observations. After time synchronization and pose projection, the new observations are aggregated into observation evidence in a local grid. g And combined with cross-modal inconsistency d g With fusion uncertainty σ g The evidence is weighted, and then a weighted Bayesian log-probability algorithm is used to complete the recursive update from prior to posterior in the two-channel approach. Core writing:

[0154] L c (g,t + ) = L c (g,t)+γ g l c (e g );

[0155] Where L c (g,t) represents the logarithmic priori a priori value of channel c of grid g at time t. c (g,t + ) represents the log-probability posterior after including new observations, c∈{occ,trans} represents the entity occupancy and permeability channels respectively, g represents the local raster index, and t represents the discrete time step. Indicated by evidence e g Log-likelihood increment to channel c, γ g =exp(-βσ g ) represents the update gain suppressed by fusion uncertainty, β>0 represents the gain scaling factor, exp(·) represents the exponential function, and a * Represents the final executed flight motion vector, η * This represents the final sensor parameter vector executed.

[0156] To use probabilistic representation of posterior states for planning, the following approach is adopted:

[0157]

[0158] Where P c (g,t + ) indicates that the raster g at time t + The posterior probability of channel c, This is the explicit form of the probabilistic mapping;

[0159] Subsequently, path planning is performed on the local raster map, simultaneously penalizing entity occupancy and insufficient permeability. The planning cost and optimal path are then written as follows:

[0160]

[0161] Where c g Represents the planning cost of grid g, P occ (g,t +) and P trans (g,t + ) represent the updated entity occupancy probability and the permeability probability, respectively, and λ o >0 and λ t >0 indicates the weights of the two costs, π represents a feasible path consisting of adjacent grid cells, Π represents the set of feasible paths connecting the current pose and the local target, and π * The optimal path with the minimum total cost is represented by J(π), the path cost is represented by J(π), the summation operator along the path is represented by ∑, the operator that takes the independent variable that minimizes the objective is represented by arg min, ∈ represents the set membership relation, and 1 represents a constant used to construct a penalty for the probability of penetrability.

[0162] Finally, the optimal path is traced and mapped to generate navigation control, which is then issued and executed. (Written as follows:)

[0163] u t =κ(π) * );

[0164] Where u t Let κ(·) represent the navigation control command vector at time t, κ(·) represent the trajectory tracking mapping that converts the discrete path into an executable control sequence, and (·) represent the placeholder for the independent variable of the function.

[0165] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A GPS-free autonomous navigation method for unmanned aerial vehicles based on multi-modal perception and reinforcement learning, characterized in that, The method comprises the following steps: S1, collecting multi-modal original observation data of at least two sensors, performing time synchronization and space calibration, estimating the pose of the unmanned aerial vehicle using a visual-inertial odometer, and constructing a local grid map with the unmanned aerial vehicle as a reference under the pose of the unmanned aerial vehicle, and outputting time-aligned multi-modal observation data, the pose of the unmanned aerial vehicle and the local grid map; S2, taking the time-aligned multi-modal observation data, the pose of the unmanned aerial vehicle and the local grid map as inputs, calculating the cross-modal inconsistency and fusion uncertainty using a deep neural network for cross-modal feature fusion and outputting the same; S3, taking the cross-modal inconsistency, fusion uncertainty, time-aligned multi-modal observation data, pose of the unmanned aerial vehicle and local grid map as inputs, constructing a double-channel occupancy map on the local grid map and calculating entity occupancy probability and transparent layer probability grid by grid and outputting the same; S4, taking the entity occupancy probability, transparent layer probability, cross-modal inconsistency, fusion uncertainty, time-aligned multi-modal observation data, pose of the unmanned aerial vehicle and local grid map as inputs, estimating the local plane normal and normal uncertainty based on the multi-modal observation data in the suspected transparent or reflective region where the transparent layer probability and the cross-modal inconsistency meet the preset conditions and outputting the same; S5, taking the local plane normal, normal uncertainty, entity occupancy probability, transparent layer probability, cross-modal inconsistency, fusion uncertainty, time-aligned multi-modal observation data, pose of the unmanned aerial vehicle and local grid map as inputs, constructing an anisotropic safety barrier aligned with the local plane normal, calculating the normal safety margin and the tangential safety margin at the grid in the suspected transparent or reflective region, setting the upper limit of the normal velocity and the upper limit of the tangential velocity according to the anisotropic safety barrier respectively, obtaining an actionable set consistent with the anisotropic safety barrier, and outputting the anisotropic safety barrier parameters, the upper limit of the normal velocity, the upper limit of the tangential velocity and the actionable set; S6, taking the anisotropic safety barrier parameters, the upper limit of the normal velocity, the upper limit of the tangential velocity, the actionable set, the local plane normal, the normal uncertainty, the entity occupancy probability, the transparent layer probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multi-modal observation data, the pose of the unmanned aerial vehicle and the local grid map as inputs, calling a reinforcement learning policy network for active verification, generating a candidate action set containing flight actions and sensor parameter configurations based on the inputs and outputting the same; S7, taking the candidate action set, the anisotropic safety barrier parameters, the upper limit of the normal velocity, the upper limit of the tangential velocity and the actionable set as inputs, performing constraint and feasibility on the candidate actions according to the anisotropic safety barrier and the upper limit of the velocity through a safety constraint layer, selecting an execution action in the actionable set and pairing an execution sensor parameter configuration, and outputting the execution action and the execution sensor parameter configuration. S8, taking the execution action, the execution sensor parameter configuration, the entity occupancy probability, the transparent layer probability, the cross-modal inconsistency degree, the fusion uncertainty, the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map as inputs, controlling the unmanned aerial vehicle to obtain new observations according to the execution action and according to the execution sensor parameter configuration, performing Bayesian update on the entity occupancy probability and the transparent layer probability based on the new observations and in combination with the cross-modal inconsistency degree and the fusion uncertainty, performing local path planning on the local grid map according to the updated entity occupancy probability and the transparent layer probability, and generating and outputting navigation control instructions.

2. The GPS-free UAV autonomous navigation method based on multi-modal perception and reinforcement learning according to claim 1, characterized in that, S1 is specifically: collecting multi-modal raw observation data from at least two sensors among a camera, a laser radar, an inertial measurement unit, a millimeter wave radar and a polarization camera, establishing a unified time reference and aligning the time stamps of each sensor, estimating and compensating time offset and time delay to obtain time-aligned multi-modal observation data; spatially calibrating each sensor to obtain intrinsic parameters of each sensor and extrinsic parameters relative to a coordinate system with the unmanned aerial vehicle as a reference, and transforming the time-aligned multi-modal observation data to the coordinate system with the unmanned aerial vehicle as a reference; fusing the transformed observation data by using a tightly coupled visual-inertial odometer of a camera and an inertial measurement unit to output the unmanned aerial vehicle pose; projecting and gridding the spatially calibrated time-aligned multi-modal observation data with the unmanned aerial vehicle pose as a reference to construct a local grid map with the unmanned aerial vehicle as a reference; outputting the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map. 3.The GPS-free UAV autonomous navigation method based on multi-modal perception and reinforcement learning of claim 1, wherein, S2 is specifically: taking the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map as inputs, projecting the time-aligned multi-modal observation data to a grid coordinate system of the local grid map according to the unmanned aerial vehicle pose and generating feature representations for each sensor in each grid, the feature representations at least including indicators reflecting geometric relationships and appearance information; taking the grid-level feature representations and the local grid map as inputs, calling a deep neural network for cross-modal feature fusion to evaluate the observation consistency of different sensors in the same grid and estimate the credibility of the fusion result, and calculating and outputting the cross-modal inconsistency degree and the fusion uncertainty in each grid of the local grid map.

4. The GPS-free UAV autonomous navigation method based on multi-modal perception and reinforcement learning according to claim 1, characterized in that, S3 is specifically: The cross-modal inconsistency, fusion uncertainty, time-aligned multi-modal observation data, unmanned aerial vehicle pose and local grid map are taken as inputs, a double-channel occupancy state is established in each grid of the local grid map, and an entity occupancy probability and a transparent layer probability are respectively taken as initial values, the time-aligned multi-modal observation data is projected to a corresponding grid under the unmanned aerial vehicle pose to form observation evidence, and the entity occupancy probability and the transparent layer probability are recursively updated according to a probability updating rule, wherein the cross-modal inconsistency is used for adjusting the update direction and weight of the two types of probabilities, so that the update weight of the transparent layer probability increases with the increase of the cross-modal inconsistency, and the update weight of the entity occupancy probability decreases with the increase of the cross-modal inconsistency, and the fusion uncertainty is used for adjusting the update gain, so that the update gain of the two types of probabilities decreases when the fusion uncertainty increases, so as to realize the suppression of unreliable observation. After the update of the full grid is completed, the entity occupancy probability and the transparent layer probability are output.

5. The method of claim 1, wherein, S4 specifically is: The entity occupancy probability, the transparent layer probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map are taken as inputs, and a suspected transparent or reflective area is selected as an estimated area from the local grid map according to the fact that the transparent layer probability and the cross-modal inconsistency meet a preset condition; Features for plane estimation are extracted in the estimated area based on the time-aligned multi-modal observation data, and the features at least include one of the following: spatial points and echo attributes provided by a laser radar or a millimeter wave radar, edges and intensity gradients provided by a camera, and motion consistency measures provided by vision and an inertial measurement unit; A local plane model is established in each estimated area based on the features, and a local plane normal is obtained through robust fitting; A normal uncertainty is calculated according to a fitting residual, a support observation number, a cross-modal inconsistency and a fusion uncertainty, so that the normal uncertainty increases when the fusion uncertainty increases; The local plane normal and the normal uncertainty are output.

6. The method of claim 1, wherein, S5 specifically is: The local plane normal, the normal uncertainty, the entity occupancy probability, the transparent layer probability, the cross-modal inconsistency, the fusion uncertainty, the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map are taken as inputs, a suspected transparent or reflective area is selected from the local grid map according to the fact that the transparent layer probability and the cross-modal inconsistency meet a preset condition, and an anisotropic safety barrier aligned with the local plane normal is constructed. The anisotropic safety barrier defines a normal direction with the local plane normal and a tangent direction perpendicular to the normal direction, and calculates a normal safety margin and a tangent safety margin for each grid of the suspected transparent or reflective area respectively, wherein the normal safety margin is determined according to a weighted relationship of the fusion uncertainty, the entity occupancy probability and the transparent layer probability, so that the normal safety margin increases as the fusion uncertainty increases, the entity occupancy probability increases, and the transparent layer probability decreases, and the tangent safety margin is determined according to a monotonic relationship of the transparent layer probability, so that the tangent safety margin increases as the transparent layer probability increases. The speed of the unmanned aerial vehicle is decomposed according to the anisotropic safety barrier, and a normal speed upper limit and a tangent speed upper limit are respectively set, so that the normal speed upper limit decreases as the fusion uncertainty and the entity occupancy probability increase, and increases as the transparent layer probability increases, and the tangent speed upper limit increases as the transparent layer probability increases. The anisotropic safety barrier, the normal safety margin, the tangent safety margin, the normal speed upper limit and the tangent speed upper limit are combined with the dynamic constraint of the unmanned aerial vehicle to generate an actionable set satisfying the anisotropic safety barrier, and the anisotropic safety barrier parameters, the normal speed upper limit, the tangent speed upper limit and the actionable set are output.

7. The method of claim 1, wherein, S6 specifically is: The anisotropic safety barrier parameters, the normal speed upper limit, the tangent speed upper limit, the actionable set, the local plane normal, the normal uncertainty, the entity occupancy probability, the transparent layer probability, the cross-modal inconsistency degree, the fusion uncertainty, the time-aligned multi-modal observation data, the unmanned aerial vehicle pose and the local grid map are taken as inputs, and whether to trigger active verification is determined according to a comparison result of the cross-modal inconsistency degree or the fusion uncertainty in front of the planned trajectory with a preset threshold, and the preset threshold can be adaptively adjusted according to the fusion uncertainty and the task risk. When the active verification is triggered, a local state composed of the entity occupancy probability, the transparent layer probability, the cross-modal inconsistency degree, the fusion uncertainty, the local plane normal, the normal uncertainty, the normal speed upper limit and the tangent speed upper limit is taken as an input of a reinforcement learning strategy network, the reinforcement learning strategy network is a trained strategy network and can be trained or updated before or during operation, so as to improve the confidence of the entity occupancy probability and the transparent layer probability and reduce the local uncertainty under the condition of satisfying the anisotropic safety barrier and the speed upper limit constraint, and output a candidate action set containing flight actions and sensor parameter configurations, the flight actions produce small changes in local position or attitude, and the sensor parameter configurations are used to adjust the sampling characteristics of the sensor to enhance the observation of the suspected transparent or reflective area. 8.The GPS-free autonomous navigation method for UAV based on multi-modal perception and reinforcement learning of claim 1, wherein, S7 specifically is: The candidate action set, the anisotropic safety barrier parameters, the normal speed upper limit, the tangent speed upper limit and the actionable set are taken as inputs, and the safety constraint layer is called to constrain the candidate action set and perform feasibility processing. The constraint judgment comprises detecting displacement and velocity decomposition of each candidate action according to anisotropic safety barrier parameters and a local grid map, judging whether the distance in the normal direction is not less than the corresponding normal safety margin and whether the motion in the tangential direction is within the corresponding tangential safety margin, and limiting the normal velocity component and the tangential velocity component of each candidate action according to the upper limit of the normal velocity and the upper limit of the tangential velocity, so that the normal velocity component does not exceed the upper limit of the normal velocity and the tangential velocity component does not exceed the upper limit of the tangential velocity; The feasibility processing comprises eliminating candidate actions that violate constraints or mapping them in the set of actionable actions based on the principle of minimum change to obtain alternative actions that satisfy constraints, and performing consistent constraint checking and limiting on sensor parameter scheduling contained in the candidate actions to keep consistent with sensor parameter constraints in the set of actionable actions; After completing the constraint judgment and feasibility processing of all candidate actions, the actionable actions that satisfy the anisotropic safety barrier and the upper limit of the velocity are selected as candidates, from which an execution action is selected for execution and paired with a corresponding sensor parameter configuration, and the execution action and the execution sensor parameter configuration are output. 9.The GPS-free autonomous navigation method for UAV based on multi-modal perception and reinforcement learning of claim 1, wherein, S8 specifically comprises: The execution action, the execution sensor parameter configuration, the entity occupancy probability, the penetrable layer probability, the cross-modal inconsistency degree, the fusion uncertainty, the time-aligned multi-modal observation data, the UAV pose and the local grid map are taken as inputs, the UAV is controlled to execute the execution action and drive the sensor to collect new observations according to the execution sensor parameter configuration; The new observations are time-synchronized and projected onto the local grid map based on the UAV pose to form observation evidence for updating; The entity occupancy probability and the penetrable layer probability are taken as priors, the observation evidence is weighted in combination with the cross-modal inconsistency degree and the fusion uncertainty, and the entity occupancy probability and the penetrable layer probability are updated according to the probability updating rule to obtain updated entity occupancy probability and updated penetrable layer probability; Local path planning is performed on the local grid map according to the updated entity occupancy probability and the updated penetrable layer probability, wherein a planning cost function simultaneously penalizes the decrease of the entity occupancy probability and the penetrable layer probability, so as to preferentially select regions with high penetrable layer probability and low entity occupancy probability and implement detouring for regions with high entity occupancy probability, and navigation control instructions are generated and output.

Citation Information

Cited By

  • Unmanned aerial vehicle autonomous return flight control method in interference environment based on multi-source information

    CN121879390A