A radar-optoelectronic fusion obstacle avoidance method for unmanned aerial vehicle navigation
By using a tangential escape force method that dynamically adjusts the repulsive field gain coefficient and confidence gradient, combined with a dynamic weight fusion model and an improved artificial potential field method, the problems of lag and insufficient accuracy in UAV obstacle avoidance response are solved, and highly reliable autonomous obstacle avoidance is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN ZHONGXING AVIATION TECH CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-05-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing UAV obstacle avoidance technologies, the simple fusion of radar and photoelectric sensors cannot fully leverage their complementary advantages, resulting in delayed obstacle avoidance response, insufficient accuracy, and weak tracking ability for dynamic obstacles in complex environments, making it difficult to meet the requirements for highly reliable autonomous obstacle avoidance.
A radar-electro-optical fusion obstacle avoidance method is adopted. By dynamically adjusting the repulsive field gain coefficient and the tangential escape force of the confidence gradient, the tangential force perpendicular to the repulsive force is generated by utilizing the spatial distribution gradient of the confidence, guiding the UAV to circle around to the high confidence region. The local obstacle avoidance path is planned by combining a dynamic weight fusion model and an improved artificial potential field method.
It improves the reliability of autonomous obstacle avoidance in drone navigation, ensuring that it avoids blurry or dangerous areas in complex environments, and improves the accuracy and safety of obstacle avoidance. It mimics the behavior of drones when flying in fog or at night, avoiding drones from flying straight into blurry objects in front of them, achieving a safer direction and solving the obstacle avoidance problem.
Smart Images

Figure CN122110115A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of radio navigation and control technology, and specifically relates to a radar-electro-optical fusion obstacle avoidance method for UAV navigation. Background Technology
[0002] As an intelligent piece of equipment integrating reconnaissance, mapping, communication, and inspection functions, unmanned aerial vehicles (UAVs) have been widely used in various fields, including civilian, industrial, and military applications. During autonomous navigation, obstacle avoidance capability is the core key to ensuring flight safety and completing operational tasks, directly determining the UAV's adaptability and operational reliability in complex environments.
[0003] Currently, obstacle avoidance in drones mainly relies on single sensors or simple sensor overlay schemes. Among these, pure radar obstacle avoidance and pure electro-optical obstacle avoidance are the most widely used. Pure radar obstacle avoidance (such as millimeter-wave radar and lidar) has all-weather operation capability, unaffected by environmental factors such as light, rain, snow, and fog. It can accurately acquire the distance, speed, and other motion parameters of obstacles. However, it has a weakness in recognizing the outline and type of obstacles, making it difficult to distinguish small obstacles (such as low-altitude cables and tree branches) from the environmental background, which can easily lead to misjudgments or omissions. Furthermore, radar point cloud data lacks semantic information, failing to provide richer environmental references for obstacle avoidance decisions. Pure electro-optical obstacle avoidance (such as RGB cameras and infrared cameras) can clearly acquire the outline, texture, and type of obstacles, accurately identifying obstacle types. However, it is greatly affected by ambient lighting conditions. In strong light, backlight, low light, or severe weather (rain, snow, fog), the imaging quality drops sharply, the reliability of obstacle avoidance decreases significantly, and it may even fail completely.
[0004] While some existing technologies employ obstacle avoidance fusion methods combining radar and photoelectric sensors, most utilize a simple post-fusion approach. This involves each sensor processing its own data before the data is superimposed at the decision layer, failing to achieve deep data-level fusion. This results in spatiotemporal asynchrony, data redundancy, and fixed weight allocation. Such simple fusion schemes cannot fully leverage the complementary advantages of radar and photoelectric sensors. In complex dynamic environments (such as urban buildings, mountainous areas, and complex low-altitude conditions), they still suffer from delayed obstacle avoidance response, insufficient obstacle avoidance accuracy, and weak tracking capabilities for dynamic obstacles, making it difficult to meet the high-precision, high-reliability autonomous obstacle avoidance requirements of UAVs. Summary of the Invention
[0005] This invention provides a radar-electro-optical fusion obstacle avoidance method for UAV navigation, which addresses the technical problem of low reliability in autonomous obstacle avoidance of UAV navigation in the prior art. The repulsive field gain coefficient is dynamically adjusted according to the confidence level of the fused data. A tangential escape force method based on the confidence gradient is adopted. By utilizing the spatial distribution gradient of confidence, a tangential force perpendicular to the repulsive force is generated to guide the UAV to a high confidence area, effectively improving the reliability of autonomous obstacle avoidance of UAV navigation.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0007] A radar-electro-optical fusion obstacle avoidance method for UAV navigation includes the following steps:
[0008] Step 1: Acquire point cloud data from millimeter-wave radar and image data from visible light camera, and perform spatiotemporal synchronization calibration on the two types of data;
[0009] Step 2: Construct a dynamic weighted fusion model to calculate ambient light intensity, target reflectivity, and relative motion speed characteristics in real time;
[0010] Step 3: Based on the dynamic weight fusion model, the weighted point cloud spatial coordinates are mapped to the image pixel plane to generate a 3D dense obstacle map with semantic labels;
[0011] Step 4: Based on the 3D dense obstacle map and the current flight attitude of the UAV, a local obstacle avoidance path is planned using an improved artificial potential field method. The improved artificial potential field method uses the repulsive field gain coefficient to be dynamically adjusted with the confidence level of the fused data.
[0012] Optionally, in step 1, to address the issues of sparse radar point clouds lacking semantics and dense images lacking precise depth and velocity information, a cross-modal semantic-geometric entanglement operator is used for computation, constructing a dynamic fusion engine that uses radar geometry as the anchor point, image semantics as the content, Doppler velocity as the priority, and edge alignment as the credibility criterion.
[0013] Optionally, in step 1, a continuous-time implicit manifold reconstruction method is adopted. By abandoning the concept of discrete frames, radar point clouds and image pixels are regarded as sampling points on a continuous-time manifold. The neural implicit field is used to learn the continuous motion trajectory of the scene during the start and end times. Thus, radar points at any time can be synthesized to the camera's precise exposure time to simulate the continuous motion process of the object during this time, rather than a jump update.
[0014] Optionally, in step 2, when strong backlight interference is detected in the ambient light, the confidence weight of the millimeter-wave radar data is automatically increased, and the point cloud data is used to generate a preliminary obstacle outline.
[0015] When sufficient ambient light is detected and the target texture is clear, the confidence weight of the visible light data is automatically increased, and a deep learning network is used to extract the target semantic category and refine the edges.
[0016] Furthermore, since the features are not independent and have strong coupling and mutual exclusion relationships, the dynamic weight fusion model uses a dynamic mutual exclusion tensor for calculation. It uses a 3×3 symmetric matrix to model the nonlinear coupling and conflict relationship between the light intensity, reflectivity and relative motion speed of the target features, so as to characterize the cross-coupling, mutual exclusion and enhancement of features.
[0017] Optionally, in step 3, a semantically aware nonlinear projection mapping is used to map the weighted 3D coordinates onto the image plane and introduce a semantic offset correction term to maintain the semantic structure during nonlinear projection.
[0018] Optionally, in step 3, since the resolution of traditional voxel maps is fixed, flat areas waste memory and complex edge details are insufficient. The adaptive resolution method guided by information entropy makes the map resolution a function that varies with spatial location and is determined by local geometric complexity or entropy. Through the nonlinear scaling mechanism driven by local geometric entropy, the intelligent adaptive allocation of 3D map resolution is realized. While ensuring the accuracy of key areas, redundant data is greatly compressed, which is an efficient mapping method for resource-constrained platforms.
[0019] Optionally, in step 4, the spatial location, distribution density and geometric contour of obstacles in the environment are obtained through a three-dimensional dense obstacle map, and a three-dimensional environment model containing the obstacle repulsion field is constructed. At the same time, the current flight attitude and position information of the UAV are collected in real time, and the UAV is regarded as a moving point in three-dimensional space.
[0020] The improvement made by the artificial potential field method is:
[0021] A dynamic gravitational field is introduced, with the UAV target point as the gravitational source. The magnitude of the gravitational force is adaptively adjusted according to the distance between the UAV and the target point and the stability of the flight attitude, so as to avoid tracking errors caused by sudden attitude changes.
[0022] An adaptive repulsive field is constructed. Based on the density and distance of obstacles in a 3D dense obstacle map, the repulsive coefficient and the range of action are dynamically adjusted. The repulsive force is enhanced for close-range, high-density obstacles, while the influence of sparse obstacles is weakened, thus preventing local optima and oscillations.
[0023] Furthermore, by dynamically adjusting the repulsive field gain coefficient according to the confidence level of the fused data, and using the tangential escape force method based on the confidence gradient, a tangential force perpendicular to the repulsive force is generated by utilizing the spatial distribution gradient of the confidence level, guiding the UAV to circle around to the high confidence level region.
[0024] The beneficial effects of this invention are:
[0025] This invention utilizes the dynamic adjustment of the repulsive field gain coefficient with the confidence level of the fused data, employing a tangential escape force method based on the confidence gradient. By leveraging the spatial distribution gradient of confidence, a tangential force perpendicular to the repulsive force is generated, guiding the drone to navigate towards a high-confidence region. This tangential escape force method based on the confidence gradient is used to improve the problem of drones getting trapped in local minima (such as U-traps or deadlocks in narrow passages) in the Artificial Potential Field (APF) method. By introducing information about the spatial distribution of confidence, the drone is guided to actively navigate towards a clearer and safer direction in areas of ambiguous perception or danger.
[0026] When a drone is near an obstacle and its visibility is poor (low confidence), it will move in the direction of better visibility (confidence gradient direction), thus avoiding dangerous and blurry areas. This mimics the behavior of drones flying in fog or at night; instead of heading straight for blurry objects, it changes its flight direction and heads towards areas with better visibility. This effectively improves the reliability of autonomous obstacle avoidance in drone navigation. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a schematic diagram of the workflow of the present invention;
[0029] Figure 2 This is a visual illustration of obstacle avoidance for the UAV according to the present invention. Detailed Implementation
[0030] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0031] Example 1;
[0032] like Figure 1 As shown, this embodiment provides a radar-electro-optical fusion obstacle avoidance method for UAV navigation, including the following steps:
[0033] Step 1: Acquire point cloud data from millimeter-wave radar and image data from visible light camera, and perform spatiotemporal synchronization calibration on the two types of data;
[0034] Millimeter-wave radar generates sparse point cloud data containing target range, azimuth, elevation, and radial velocity by transmitting millimeter waves and receiving the echoes. Its advantages are that it is not affected by light, rain, or fog, and it can directly measure velocity.
[0035] Visible light cameras acquire high-resolution image data, providing rich texture, color, and edge information. However, they are affected by lighting conditions (such as darkness or backlighting), and cannot directly acquire depth information in monocular situations, but can directly acquire depth information in multi-view situations.
[0036] For time synchronization, since the radar and camera have different sampling frequencies (e.g., radar 20Hz, camera 30Hz), it is necessary to use hardware triggering (PTP protocol) or software interpolation algorithm to stamp the data of both with a unified timestamp to ensure that the environmental state at the same moment is being processed, and to avoid motion blur or position deviation caused by time difference during high-speed flight.
[0037] For spatial calibration, the extrinsic parameter matrix (rotation and translation relationship) between the radar coordinate system and the camera coordinate system is calculated using a joint calibration board or a self-calibration algorithm. This allows the radar's three-dimensional points to be accurately projected onto the two-dimensional plane of the image, or to map image features to the radar's spatial coordinates, achieving pixel-level fusion.
[0038] Step 2: Construct a dynamic weighted fusion model to calculate ambient light intensity, target reflectivity, and relative motion speed characteristics in real time;
[0039] Among them, when strong backlighting is detected to interfere with the ambient light (which may cause the camera to overexpose or become blind), the confidence weight of the millimeter-wave radar data is automatically increased, and the point cloud data is used to generate a preliminary obstacle outline.
[0040] Furthermore, when ambient lighting is affected by strong backlighting, the camera weight is automatically reduced to improve radar confidence. Point cloud clustering algorithms (such as DBSCAN or Euclidean clustering) are used to quickly generate preliminary geometric contours and distance information of obstacles, ensuring that they are visible but not collided with, thus guaranteeing safety.
[0041] When sufficient ambient light is detected (the camera can clearly capture details without blurring) and the target texture is clear (the surface texture, edges, and details of the object to be photographed are clear), the confidence weight of the visible light data is automatically increased. A deep learning network is then used to extract the target semantic category and refine the edges. Edge refinement is a technique in digital image processing that converts wide edges into single-pixel wide curves, primarily used in computer vision. It can be divided into methods based on mathematical morphology and methods based on binary edge images: the former achieves refinement through thinning operations or hit-and-miss transformations, while the latter uses layer-by-layer edge point removal or an improved Prewitt operator to process grayscale images.
[0042] Furthermore, using deep learning networks (such as YOLO or Mask R-CNN) to extract the semantic category of the target (i.e., people, birds, power lines, and buildings) and refine the edges helps drones make smarter decisions (e.g., gently avoiding birds and braking suddenly against walls).
[0043] Step 3: Based on the dynamic weight fusion model, the weighted point cloud spatial coordinates are mapped to the image pixel plane to generate a 3D dense obstacle map with semantic labels;
[0044] Among them, the weighted radar point cloud spatial coordinates Precise mapping to the image pixel plane Visual depth estimation can also be used to supplement sparse areas of radar point clouds. Visual depth estimation is a core technology in the field of computer vision, which aims to infer the distance information of each pixel in a scene from the camera from two-dimensional images (single or multiple images), thereby giving machines a three-dimensional spatial perception ability similar to humans. It is used in the fields of autonomous driving, drone navigation, augmented reality (AR), virtual reality (VR) and 3D reconstruction.
[0045] Semantic labels are the assignment of category labels (such as obstacles like "high-voltage lines", "trees" and "buildings") identified by visual networks to corresponding 3D point cloud clusters.
[0046] Dense 3D obstacle maps are generated because radar point clouds are usually sparse. By fusing visual depth cues or utilizing temporal accumulation, dense 3D point clouds or voxel maps are generated to fill the holes on the obstacle surface, ultimately producing a 3D semantic map.
[0047] Step 4: Based on the 3D dense obstacle map and the current flight attitude of the UAV, a local obstacle avoidance path is planned using an improved artificial potential field method. The improved artificial potential field method uses the repulsive field gain coefficient to be dynamically adjusted with the confidence level of the fused data.
[0048] Flight attitude takes into account dynamic constraints such as the drone's current speed, acceleration, and maximum turning radius.
[0049] In a 3D dense obstacle map with semantic labels, different strategies are adopted for different semantic targets (e.g., the repulsive force is small for sparse vegetation that is "crossable", and extremely large for buildings that are "impassable").
[0050] Due to the limitations of the traditional artificial potential field (APF) method, the repulsive field gain coefficient is usually fixed. When facing sensor noise or targets with low confidence, it is prone to local minima (causing the drone to get stuck) or oscillations.
[0051] The improvement made by dynamically adjusting the repulsive field based on the confidence level is the dynamic adjustment of the repulsive field gain coefficient, which introduces the confidence level of the fused data as a variable into the repulsive field formula.
[0052] When the dynamic weight fusion model has a high confidence level in a certain obstacle (e.g., clearly identifying a wall), the repulsive force gain is increased to ensure that the drone maintains a safe distance and resolutely avoids it.
[0053] When the dynamic weight fusion model has low confidence in a certain obstacle (e.g., radar noise or visually blurred areas), the repulsive gain should be appropriately reduced or the gravity smoothing term should be increased to prevent the UAV from experiencing severe shaking or incorrect turning due to false obstacles.
[0054] Finally, a collision-free, smooth local obstacle avoidance path that conforms to dynamic characteristics from the current position to the target point is output and sent to the control system for execution in real time.
[0055] Example 2;
[0056] Based on Example 1, in step 1, for the obtained millimeter-wave radar point cloud data and visible light camera image data, to address the problems of sparse radar point clouds lacking semantics and dense images lacking precise depth and velocity information, a cross-modal semantic-geometric entanglement operator is used for calculation:
[0057] ;
[0058] in, The input features for the current layer are the first-order features. The fusion feature representation of a layer is obtained by passing through the previous layer or The multidimensional tensor after layer processing represents the current cognitive state of the network regarding the scene, including the fused visual and radar information; The input features for the next layer are the first layer. The fusion feature representation of a layer is obtained by passing through the current layer or The multidimensional tensor after layer processing represents the cognitive state of the subsequent network regarding the scene, including the fused visual and radar information;
[0059] This is a geometric-semantic attention mechanism, representing cross-modal conditional attention. It uses the geometric structure of radar waves to query the semantic content of an image and uses motion speed to adjust the intensity of the query. Radar point cloud feature embedding (including location) , and (reflection intensity and Doppler velocity); It is an image feature map (from a CNN or ViT encoder, containing semantic information about texture, color, and edges); The Doppler velocity vector (scalar or vector, representing the radial motion of the target) provided to the radar.
[0060] The gating mechanism, or adaptive mask, is used to control the contribution ratio of the attention output and prevent invalid or noisy information from contaminating the fusion result. The attention mechanism is only allowed to function when the radar and image are consistent in edge structure and the target has significant motion, thus avoiding artifacts caused by forced fusion in flat, textureless areas (such as the sky or walls).
[0061] The Hadamard product (element-wise multiplication) multiplies the attention output with the gating signal to achieve selective enhancement, which is equivalent to a smart switch that activates the fused features only in key regions (edges + motion).
[0062] Residual connections are a classic Transformer structure that ensures stable information flow. Even if the attention module fails, the original features can still be passed to the next layer.
[0063] Representation layer normalization normalizes the feature dimensions of each sample, accelerating training convergence and improving generalization ability. It is especially important in multimodal fusion because radar and image feature scales differ greatly.
[0064] This formula is the update unit in the multimodal fusion neural network of millimeter-wave radar and visible light camera. It integrates three dimensions: geometric structure, semantic information, and motion physical quantities, and achieves adaptive weighting through a gating mechanism. A dynamic fusion engine is constructed with radar geometry as the anchor point, image semantics as the content, Doppler velocity as the priority, and edge alignment as the confidence criterion, enabling the perception system to selectively pay attention to key targets in complex environments.
[0065] Spatiotemporal synchronization calibration is performed on both types of data using a continuous-time implicit manifold reconstruction method. This method discards the concept of discrete frames, treating radar point clouds and image pixels as sampling points on a continuous-time manifold. A neural implicit field is used to learn the continuous motion trajectory of the scene during its start and end times, allowing radar points at any given moment to be synthesized to the camera's precise exposure time. Specifically:
[0066] ;
[0067] in, These are the corrected spatial coordinates, representing the original radar points. At its collection time The location, after time travel, at the target time The position (usually the center of the camera exposure) indicates the location of the object; it does not represent a simple translation, but a new position calculated based on the actual trajectory of the object during that time period.
[0068] These are the original radar point coordinates, which are the [number]th [unit]. Each radar point at the time of data acquisition Three-dimensional spatial coordinates These are the sensor's raw observations, without any time compensation.
[0069] The potential continuous velocity field is an implicit function parameterized by a neural network, with the current spatial position as the input. and time The output is the instantaneous velocity vector of that point at that moment; where, The learnable parameters of a neural network are the set of all neural network weights (convolutional kernels, fully connected layer matrices, and attention weights). These learnable parameters encode prior knowledge of the motion trajectory and are learned through a large amount of training data (radar, image, and ground truth trajectories).
[0070] Let be the time integral operator, representing the time of data collection from the radar point. To the target time Integrating over a continuous time interval simulates the continuous motion of an object during that time interval, rather than updating in a jump manner.
[0071] This formula addresses the spatiotemporal synchronization problem between millimeter-wave radar and visible light cameras in scenarios involving high-speed motion, non-uniform acceleration, and microsecond-level time misalignment. It abandons the traditional discrete thinking based on frame alignment or constant velocity assumptions, instead treating sensor data as sampling points on a continuous spatiotemporal manifold, and achieving accurate interpolation and alignment at any given time by learning a latent continuous velocity field.
[0072] Example 3;
[0073] Based on Example 1, in step 2, since the features are not independent and have strong coupling and mutual exclusion relationships, the dynamic weight fusion model is calculated using a Dynamic Mutual Exclusion Tensor. A 3×3 symmetric matrix is used to model the illumination intensity of the target features. Target reflectivity Relative velocity The nonlinear coupling and conflict relationships between features are used to characterize the cross-coupling, mutual exclusion, and enhancement of features, specifically:
[0074] ;
[0075] in, This is a feature mutual exclusion tensor; all diagonal elements are 1, indicating that each feature has a single-position confidence or basic weight contribution in the absence of interference. It is the starting point of the weights. It is suppressed or enhanced by off-diagonal terms. Even without external interference, each feature still has basic confidence, reflecting the self-consistency baseline.
[0076] The off-diagonal elements, located at [0,1] and [1,0], are cross-coupling coefficients, describing the negative synergistic effect (i.e., "mutual exclusion") that occurs when two features appear simultaneously. A negative sign indicates "suppression". It is the glare coupling coefficient, which can control the intensity of the mutual repulsion effect; The ratio of illumination to reflectance indicates the risk of glare, when the illumination intensity is... High and target surface reflectivity High temperatures (such as sunlight on car bodies or glass curtain walls) can easily cause specular reflection, overexposure, and light spot saturation, leading to visual sensor failure.
[0077] It also represents off-diagonal elements, located at [1,2] and [2,1], and is also a mutual exclusion or coupling coefficient between features. It describes the negative synergistic effect (i.e., "mutual exclusion") that occurs when two features appear simultaneously, and is represented by a negative sign to indicate "suppression". It is the dynamic fuzzy coupling coefficient, and also controls the strength of the mutual exclusion effect; The speed multiplied by the reflectivity indicates dynamic blur plus texture loss.
[0078] High-speed movement This can lead to image blurring, and target recognition algorithms that rely on reflectivity (texture, color, and edges) will experience a sharp decline in performance under these circumstances.
[0079] The 0 element is located in [0,2] and [2,0], indicating that illumination and velocity are not directly mutually exclusive, i.e., coupled. In the model, there is no direct physical conflict or co-deterioration mechanism between illumination intensity and motion velocity. Setting it to 0 indicates that these two features act independently and do not interfere with each other.
[0080] formula It is an environmental conflict sensor that actively suppresses feature combinations that would deteriorate under physical conditions and lead to sensing failure by using negative coupling terms. This improves the robustness and security of the multimodal fusion system, ensuring that the fusion system does not blindly trust a single sensor or feature in extreme scenarios. When two features appear simultaneously and cause physical conflict (e.g., strong light + high reflection causing glare), they are not simply added together, but rather weaken each other. When two features appear simultaneously without direct physical conflict (light intensity × speed), light intensity and motion speed are coupled together.
[0081] Example 4;
[0082] Based on Example 1, in step 3, the weighted point cloud spatial coordinates are mapped to the image pixel plane. Semantic-Aware Non-linear Projection is used to map the weighted 3D coordinates to the image plane, and a semantic offset correction term is introduced to maintain the semantic structure during non-linear projection. Specifically:
[0083] ;
[0084] in, The output scalar or vector squared represents the semantic energy or activation intensity after final projection, and is considered as the input semantic offset correction term. The degree of existence or confidence in the semantic space; The modulo-square operation is to take the square of the Euclidean norm of a complex vector, which collapses the superposition state into an observable classical quantity.
[0085] To summate, is to... The basis vectors (semantic prototypes or conceptual bases) are linearly combined, and the superposition principle of quantum states is used to put an input into a superposition of multiple semantic states at the same time, until it collapses into a definite value when measured or the modulus is squared.
[0086] The amplitude coefficient is a real number function output by the neural network, controlling the first... The weights or contribution ratios of each basis vector in the superposition correspond to the probability amplitudes of each eigenstate in the quantum state; the larger the amplitude, the more relevant the semantic concept.
[0087] The phase factor represents assigning a phase offset to each basis vector to adjust its interference relationship with other basis vectors.
[0088] When two components are in phase (0° out of phase), it is constructive interference, which enhances the output; when two components are out of phase (π out of phase), it is destructive interference, which weakens or even cancels the output. This is a key mechanism for capturing contextual ambiguity. Because the phases are different in different contexts, the final semantics point in different directions.
[0089] It is the base of the natural logarithm; The imaginary unit; The phase angle function is a real-valued function, and the input semantic offset correction term is used. The output is a real angle (usually in radians). (The input is...) The dynamically determined phase offset controls the first The rotation angle of each basis vector in the complex plane. Because the same word may correspond to different phases in different contexts, the interference effect with other word concepts will vary.
[0090] For the first The semantic prototype of a basis vector is a fixed or learnable vector that represents a semantic concept or atomic feature, such as the basic semantics of "animal", "finance", and "nature". It is equivalent to the eigenstate or basis vector in quantum mechanics. All complex semantics are superpositions of these basis vectors.
[0091] This formula is a semantic synthesis mechanism based on complex superposition and phase interference, a wave phenomenon modulated by context. Using the mathematical framework of quantum mechanics, semantic understanding is described as a dynamic wave process: input data excites multiple potential semantic concepts (basis vectors), which propagate, superimpose and interfere (enhance or cancel) in complex space like waves, and finally collapse into a definite semantic intensity value through observation or modulo squaring.
[0092] In the generation of 3D dense obstacle maps, traditional voxel maps suffer from fixed resolution, leading to wasted memory in flat areas and insufficient detail in complex edges. An information-entropy-guided adaptive resolution approach is adopted to adjust the map resolution. It becomes a function that varies with spatial location, determined by local geometric complexity or entropy, specifically:
[0093] ;
[0094] in, Indicates position The adaptive voxel resolution (unit: meters / voxel side length) determines the actual physical size represented by each voxel when building a map in a spatial location;
[0095] The base resolution is the preset minimum resolution (i.e., the finest granularity), the limit of accuracy determined by hardware or algorithm capabilities, and also the benchmark scale of the entire adaptive system.
[0096] The scaling strength factor controls the maximum expandable ratio. The numbers are 4, 8, and 16;
[0097] It is an sigmoid function, also called the Logistic function, which calculates information entropy. The value is "compressed" and mapped to a fixed range (0,1), forming a smooth switch. In the complete Logistic function, by taking the reciprocal, "high information entropy" is cleverly mapped to "high activation value (close to 1)" and "low information entropy" is mapped to "low activation value (close to 0)", thus realizing an intelligent and differentiable switch function.
[0098] For the natural logarithm As the base of the exponential function, its value depends on the sign and magnitude of the exponent, acting as a reverse regulator. Based on the smooth change of information entropy, it responds with a small value for high information entropy (complex regions) and a large value for low information entropy (simple regions).
[0099] It is the base of the natural logarithm (equal to 2.718); This represents the kurtosis factor or sensitivity gain, used to control the steepness of the transition band of the Sigmoid function. The larger the value, the more sensitive it is to changes in geometric entropy, and the more abrupt the resolution switching becomes. The smaller the value, the smoother the transition. The resolution changes gradually, similar to the temperature parameter in a neural network, which affects the sharpness of the decision boundary.
[0100] The time is denoted as , indicating that the geometric entropy value corresponding to the center point of the Sigmoid function is determined according to different times. Local geometric entropy is an indicator used to measure local geometric uncertainty or structural complexity.
[0101] Through a nonlinear scaling mechanism driven by local geometric entropy, intelligent adaptive allocation of 3D map resolution is achieved. While ensuring accuracy in key areas, redundant data is significantly compressed, making it an efficient mapping solution for resource-constrained platforms (such as embedded robots and drones). By precisely injecting limited computational energy into areas with the highest entropy, maximum knowledge gain is achieved. A nonlinear balance point between error and cost is found, describing a self-organizing system that spontaneously evolves a map with minimum energy consumption (minimum cost) while satisfying obstacle avoidance safety (error constraints).
[0102] This formula is a non-linear scaler based on the Sigmoid function, which increases the base resolution. According to local geometric entropy The size is adjusted inversely: when the geometric complexity is high, the resolution becomes smaller (more detailed); when the geometric complexity is low, the resolution becomes larger (coarser). This appropriate resource allocation uses high resolution in areas with rich detail and low resolution in flat or open areas to save memory and computation, achieving map construction with the lowest energy consumption.
[0103] Example 5;
[0104] Based on Example 1, in step 4, a three-dimensional dense obstacle map is used as the basis for environmental perception, and real-time flight attitude information of the UAV is fused to perform local obstacle avoidance path planning using an improved artificial potential field method.
[0105] First, the spatial location, distribution density, and geometric contour of obstacles in the environment are obtained through a three-dimensional dense obstacle map, and a three-dimensional environment model containing the obstacle repulsion field is constructed. At the same time, the current flight attitude (pitch angle, roll angle, and yaw angle) and position information of the UAV are collected in real time, and the UAV is regarded as a moving point mass in three-dimensional space.
[0106] An improvement upon the traditional artificial potential field method is:
[0107] A dynamic gravitational field is introduced, with the UAV target point as the gravitational source. The magnitude of the gravitational force is adaptively adjusted according to the distance between the UAV and the target point and the stability of the flight attitude, so as to avoid tracking errors caused by sudden attitude changes.
[0108] An adaptive repulsive field is constructed. Based on the density and distance of obstacles in a 3D dense obstacle map, the repulsive coefficient and the range of action are dynamically adjusted. The repulsive force is enhanced for close-range, high-density obstacles, while the influence of sparse obstacles is weakened, thus preventing local optima and oscillations.
[0109] By adding attitude constraints, the flight attitude angle limit, maximum angular velocity and turning radius of the UAV are incorporated into the potential field function, so that the generated path not only meets the obstacle avoidance requirements, but also conforms to the UAV dynamics and attitude stability constraints.
[0110] By solving the improved resultant force of gravity and repulsion in real time, the desired motion direction and speed command of the UAV at the current moment are obtained, and a smooth, safe and flyable local obstacle avoidance path is generated iteratively, realizing online and real-time obstacle avoidance of the UAV in a complex three-dimensional obstacle environment.
[0111] Furthermore, to dynamically adjust the repulsive field gain coefficient with the confidence level of the fused data, a confidence-gradient tangential escape method is adopted. This method utilizes the spatial distribution gradient of confidence (i.e., sensors are clearer in certain directions) to generate a tangential force perpendicular to the repulsive force, guiding the UAV to circle around to a high-confidence region. Specifically:
[0112] ;
[0113] in, The tangential escape force vector is the final calculated additional force vector. Its direction is perpendicular to the current repulsive force direction (i.e., the line connecting the obstacle to the drone). Its magnitude is determined by multiple factors. Instead of directly pushing the drone away from the obstacle, it pushes the drone to avoid the low-confidence area.
[0114] Escape force intensity coefficient is a scalar constant used to control the overall magnitude of escape force and to adjust the strength of the willingness to bypass the uncertainty region.
[0115] This is the location uncertainty weighting term, reflecting the degree of unreliability of the current location. It is a location The fusion confidence level (0~1) indicates that the closer the data is to 1, the more reliable the sensor data is, and the closer it is to 0, the less reliable the sensor data is.
[0116] The confidence gradient vector indicates the direction in which the confidence increases the fastest, and is used to tell the drone "which direction to go", so that the sensor data will become more reliable.
[0117] The unit vector pointing from the obstacle to the drone is the radial reference direction, which serves as an input to the cross product to generate a tangential force perpendicular to the radial direction.
[0118] The cross product operation, or two-dimensional pseudo-cross product, is a key operation used to generate the tangential direction. In two-dimensional space, the cross product usually refers to rotating a vector by 90° and then dot producting it with another vector, or directly constructing a perpendicular vector. Indicates will Rotate 90° to obtain the tangential direction, then project it onto... Orthogonalization creates an escape direction perpendicular to the radial direction.
[0119] This is a distance attenuation factor to ensure that the escape force only takes effect when the target is near an obstacle. It is the base of the natural logarithm; The distance from the drone to the nearest obstacle; This is the attenuation scale parameter (similar to the standard deviation of a Gaussian kernel). It avoids applying unnecessary lateral forces in open areas and is only activated when obstacle avoidance and escape are truly needed.
[0120] The tangential escape force method based on confidence gradient is used to improve the problem of UAVs getting trapped in local minima (such as U-shaped traps and deadlocks in narrow passages) in the Artificial Potential Field (APF) method. This formula guides the UAV to actively detour to a clearer and safer direction in areas with ambiguous perception or danger by incorporating information about the spatial distribution of confidence.
[0121] like Figure 2 As shown, when the drone is near an obstacle and feels that it cannot see clearly (low confidence), it will move in the direction where it can see more clearly (the direction of confidence gradient) to avoid dangerous and blurry areas. This mimics the behavior of drones when flying in fog or at night. Instead of flying straight towards blurry objects, it changes its flight direction and heads towards areas with better visibility.
[0122] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope described in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A radar-electro-optical fusion obstacle avoidance method for UAV navigation, characterized in that, Includes the following steps: Step 1: Acquire point cloud data from millimeter-wave radar and image data from visible light camera, and perform spatiotemporal synchronization calibration on the two types of data; Step 2: Construct a dynamic weighted fusion model to calculate ambient light intensity, target reflectivity, and relative motion speed characteristics in real time; Step 3: Based on the dynamic weight fusion model, the weighted point cloud spatial coordinates are mapped to the image pixel plane to generate a 3D dense obstacle map with semantic labels; Step 4: Based on the 3D dense obstacle map and the current flight attitude of the UAV, a local obstacle avoidance path is planned using an improved artificial potential field method. The improved artificial potential field method uses the repulsive field gain coefficient to be dynamically adjusted with the confidence level of the fused data.
2. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 1, to address the issues of sparse radar point clouds lacking semantics and dense images lacking precise depth and velocity information, a cross-modal semantic-geometric entanglement operator is employed for computation: ; in, The input features for the current layer are the first-order features. The fusion feature representation of the layer contains fused visual and radar information; The input features for the next layer are the first layer. The fusion feature representation of the layer contains fused visual and radar information; This is a geometric-semantic attention mechanism, representing cross-modal conditional attention, which uses the geometric structure of radar waves to query the semantic content of an image; Embedding radar point cloud features; Image feature map; The Doppler velocity vector provided for the radar; For gating mechanism, an adaptive mask is used to control the contribution ratio of attention output and prevent invalid or noisy information from contaminating the fusion result; The Hadamard product, or element-wise multiplication, multiplies the attention output with the gating signal to achieve selective enhancement. For residual connections, it is a classic Transformer structure that ensures stable information flow; Representation layer normalization normalizes the feature dimensions of each sample, accelerating training convergence and improving generalization ability.
3. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 1, a continuous-time implicit manifold reconstruction method is adopted. By abandoning the concept of discrete frames, radar point clouds and image pixels are regarded as sampling points on a continuous-time manifold. The neural implicit field is used to learn the continuous motion trajectory of the scene during the start and end times, so that radar points at any time can be synthesized to the camera's precise exposure time. Specifically: ; in, These are the corrected spatial coordinates, representing the original radar points. At its collection time The location, after time travel, at the target time Location; These are the original radar point coordinates, which are the first... Each radar point at the time of data collection Three-dimensional spatial coordinates These are the sensor's raw observations, without any time compensation; The potential continuous velocity field is an implicit function parameterized by a neural network, with the current spatial position as the input. and time The output is the instantaneous velocity vector of that point at that moment. These are the learnable parameters of the neural network, which is the set of all neural network weights; Let be the time integral operator, representing the time of data collection from the radar point. To the target time Integrating over a continuous time interval simulates the continuous motion of an object during that time interval, rather than updating in a jump manner.
4. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 2, when strong backlight interference is detected in the ambient light, the confidence weight of the millimeter-wave radar data is automatically increased, and the point cloud data is used to generate a preliminary obstacle outline. When sufficient ambient light is detected and the target texture is clear, the confidence weight of the visible light data is automatically increased, and a deep learning network is used to extract the target semantic category and refine the edges.
5. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 4, characterized in that, The dynamic weight fusion model is calculated using a dynamic mutual exclusion tensor and employs a 3×3 symmetric matrix to model the illumination intensity of target features. Target reflectivity Relative velocity The nonlinear coupling and conflict relationships between features are used to characterize the cross-coupling, mutual exclusion, and enhancement of features, specifically: ; in, For feature mutual exclusion tensors; all diagonal elements are 1, indicating that each feature has a single-position confidence or basic weight contribution in the interference-free state, which is the starting point of the weights, and is suppressed or enhanced through off-diagonal terms; The off-diagonal elements, located at [0,1] and [1,0], represent the mutual exclusion or coupling coefficients between features. They describe the negative synergistic effect produced when two features appear simultaneously, and are indicated by a negative sign to suppress this effect. It is the glare coupling coefficient, which controls the intensity of the mutual repulsion effect; It also represents off-diagonal elements, located at [1,2] and [2,1], which are also mutual exclusion or coupling coefficients between features. They describe the negative synergistic effect produced when two features appear simultaneously, and are suppressed by a negative sign. It is the dynamic fuzzy coupling coefficient, which controls the strength of the mutual exclusion effect; The 0 element is located in [0,2] and [2,0], indicating that there is no direct mutual exclusion or coupling between illumination and velocity. In the model, there is no direct physical conflict or co-deterioration mechanism between illumination intensity and motion velocity.
6. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 3, a semantically aware nonlinear projection mapping is used to map the weighted 3D coordinates onto the image plane, and a semantic offset correction term is introduced to maintain the semantic structure during nonlinear projection. Specifically: ; in, The output scalar or vector squared represents the semantic energy or activation intensity after final projection, and is considered as the input semantic offset correction term. The degree of existence or confidence in the semantic space; The modulo-square operation is to take the square of the Euclidean norm of a complex vector, which collapses the superposition state into an observable classical quantity. To summate, is to... Linear combination of basis vectors; The amplitude coefficient is a real number function, output by the neural network, which controls the first... The weights or contribution ratios of each basis vector in the superposition; The phase factor represents assigning a phase offset to each basis vector to adjust its interference relationship with other basis vectors. It is the base of the natural logarithm; The imaginary unit; The phase angle function is a real-valued function, and the input semantic offset correction term is... The output is a real number angle. For the first The semantic prototype of a basis vector is a fixed or learnable vector that represents a basic semantic concept or atomic feature.
7. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 3, because traditional voxel maps have a fixed resolution, flat areas waste memory and complex edge details are insufficient. An information entropy-guided adaptive resolution approach is adopted to make the map resolution a function of spatial location, determined by local geometric complexity or entropy. Specifically: ; in, Indicates position The adaptive voxel resolution at a given location determines the actual physical size represented by each voxel when building a map in a spatial location; The base resolution is the preset minimum resolution, the limit of accuracy determined by hardware or algorithm capabilities, and also the benchmark scale of the entire adaptive system. The scaling strength factor controls the maximum expandable factor. It is an sigmoid function, also called the Logistic function, which calculates information entropy. The values are compressed and mapped to a fixed range (0,1), forming a smooth switch; For the natural logarithm As the base of the exponential function, its value depends on the sign and magnitude of the exponent, thus acting as a reverse regulator. It is the base of the natural logarithm; This represents the kurtosis factor or sensitivity gain, used to control the steepness of the transition band of the Sigmoid function; The value at time represents the geometric entropy value corresponding to the center point of the Sigmoid function, which is determined based on different times. Local geometric entropy is an indicator used to measure local geometric uncertainty or structural complexity.
8. The radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 1, characterized in that, In step 4, the spatial location, distribution density and geometric contour of obstacles in the environment are obtained through a three-dimensional dense obstacle map, and a three-dimensional environment model containing the obstacle repulsion field is constructed. At the same time, the current flight attitude and position information of the UAV are collected in real time, and the UAV is regarded as a moving point in three-dimensional space. The improvement made by the artificial potential field method is: A dynamic gravitational field is introduced, with the UAV target point as the gravitational source. The magnitude of the gravitational force is adaptively adjusted according to the distance between the UAV and the target point and the stability of the flight attitude, so as to avoid tracking errors caused by sudden attitude changes. An adaptive repulsive field is constructed. Based on the density and distance of obstacles in a 3D dense obstacle map, the repulsive coefficient and the range of action are dynamically adjusted. The repulsive force is enhanced for close-range, high-density obstacles, while the influence of sparse obstacles is weakened, thus preventing local optima and oscillations.
9. A radar-electro-optical fusion obstacle avoidance method for UAV navigation according to claim 8, characterized in that, By dynamically adjusting the repulsive field gain coefficient according to the confidence level of the fused data, and employing a tangential escape force method based on the confidence gradient, a tangential force perpendicular to the repulsive force is generated using the spatial distribution gradient of the confidence level to guide the UAV to circle around to a high-confidence region. Specifically: ; in, The tangential escape force vector is the final calculated additional force vector. Its direction is perpendicular to the current repulsive force direction, and its magnitude is determined by multiple factors. Instead of directly pushing the drone away from the obstacle, it pushes the drone to avoid the low confidence area. The escape force intensity coefficient is a scalar constant used to control the overall magnitude of the escape force and to adjust the strength of the intention to bypass the uncertainty region. This is the location uncertainty weighting term, reflecting the degree of unreliability of the current location. It is a location Fusion confidence at the location; This is the confidence gradient vector, indicating the direction in which the confidence increases the fastest; The unit vector pointing from the obstacle to the drone is the radial reference direction, which is used as an input to the cross product to generate a tangential force perpendicular to the radial direction. Cross product, or two-dimensional pseudo-cross product, is a key operation used to generate the tangential direction. In two-dimensional space, cross product refers to rotating one vector by 90° and then dot producting it with another vector, or directly constructing a perpendicular vector. Indicates will Rotate 90° to obtain the tangential direction, then project it onto... Orthogonalization creates an escape direction perpendicular to the radial direction; This is a distance attenuation factor to ensure that the escape force only takes effect when the target is near an obstacle. It is the base of the natural logarithm; The distance from the drone to the nearest obstacle; The attenuation scale parameter is used to avoid applying unnecessary lateral forces in open areas and to activate only when obstacle avoidance and escape are truly needed.