Unmanned aerial vehicle autonomous obstacle avoidance method based on deep learning and binocular vision

Through the UAV obstacle avoidance method based on deep learning and binocular vision, the problem of unreasonable obstacle avoidance of traditional obstacle avoidance methods in complex environments is solved, the UAV can achieve efficient obstacle avoidance in dynamic environments, and the reliability and adaptability of obstacle avoidance are improved.

CN120722951AActive Publication Date: 2025-09-30SHANGHAI BOLI INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511203510.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-30
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Traditional drone obstacle avoidance methods have limitations in complex environments, making it difficult to meet real-time obstacle avoidance needs. They also lack the ability to adapt to dynamic changes in the environment, resulting in unreasonable obstacle avoidance paths and even causing collision risks.

Method used

An autonomous obstacle avoidance method based on deep learning and binocular vision is adopted. Stereoscopic visual features are acquired through a binocular vision acquisition terminal. A deep neural network model and a three-dimensional space reconstruction module are combined to generate a dense depth map and the initial position coordinates of obstacles. The dynamic obstacle analysis engine calculates the dynamic threat assessment index of obstacles, and uses a spatiotemporal trajectory prediction model to infer the future movement path of obstacles, generate a three-dimensional obstacle avoidance heading correction vector, and finally convert it into a flight control instruction set.

Benefits of technology

It improves the reliability and adaptability of UAV obstacle avoidance in complex dynamic environments, ensures the rationality and consistency of obstacle avoidance paths, realizes a closed loop from environmental perception to motion control, and improves the real-time and accuracy of obstacle avoidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120722951A_ABST
    Figure CN120722951A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned aerial vehicle obstacle avoidance, and discloses an unmanned aerial vehicle autonomous obstacle avoidance method based on deep learning and binocular vision. According to the method, original data streams of a left view and a right view are acquired through a binocular vision acquisition terminal, and stereoscopic vision feature representation is extracted through parallel convolutional coding branches of a deep neural network model; inputting into a three-dimensional space reconstruction module to generate a dense depth map and an obstacle initial position coordinate set; the dynamic obstacle analysis engine calculates a dynamic threat evaluation index by combining real-time flight attitude parameters of the unmanned aerial vehicle, and the space-time trajectory prediction model calculates future moving path probability distribution of the obstacle according to the dynamic threat evaluation index; and fusing the distribution with preset navigation path planning data to generate a three-dimensional obstacle avoidance course correction vector, converting the three-dimensional obstacle avoidance course correction vector into a flight control instruction set, and downloading the flight control instruction set to an execution module. According to the method, the reliability and adaptability of autonomous obstacle avoidance of the unmanned aerial vehicle in a complex dynamic environment are improved, and a powerful guarantee is provided for safe navigation of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) obstacle avoidance, and in particular to an autonomous obstacle avoidance method for UAVs based on deep learning and binocular vision. Background Art

[0002] As drone applications continue to expand, from low-altitude reconnaissance and logistics to agricultural plant protection, the demand for autonomous obstacle avoidance capabilities is becoming increasingly stringent. Traditional drone obstacle avoidance methods often rely on single sensors, such as ultrasound, lidar, or monocular vision, which often face limitations in complex environments. Ultrasonic sensors are limited by their detection range and angle, making them prone to missed detections in open spaces. While lidar can obtain precise depth information, its equipment cost is high, and data stability decreases in strong sunlight or rain or snow. Monocular vision estimates depth through motion parallax, but its accuracy decreases significantly in static scenes or when moving slowly, making it difficult to meet the requirements of real-time obstacle avoidance. Existing binocular vision-based methods can directly calculate depth by calculating parallax, but are limited by the efficiency of traditional feature matching algorithms. When dealing with dynamic obstacles, trajectory prediction errors are often caused by missing or mismatched feature points. Furthermore, most obstacle avoidance systems only consider the current position of an obstacle, without incorporating its motion trends and the drone's own navigation parameters. This can easily lead to unreasonable obstacle avoidance paths in complex environments, even creating collision risks. Traditional obstacle avoidance decisions mostly rely on preset rules and lack the ability to adapt to dynamic changes in the environment. In scenarios with multiple obstacles crossing each other, it is difficult to quickly generate the optimal obstacle avoidance strategy, which limits the application scope of drones in complex dynamic environments. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for autonomous obstacle avoidance of a UAV based on deep learning and binocular vision to solve the problems raised in the above background technology.

[0004] To achieve the above objectives, the present invention provides a method for autonomous obstacle avoidance of a UAV based on deep learning and binocular vision, the method comprising: The binocular vision acquisition terminal synchronously obtains the left view and right view raw data streams in the flight environment; The deep neural network model synchronously processes the left view and right view raw data streams, and extracts stereo visual feature representations from the binocular images through parallel convolutional coding branches; Inputting the stereoscopic visual feature representation into a three-dimensional space reconstruction module to generate a dense depth map of the flight environment and a set of initial position coordinates of obstacles; The dynamic obstacle analysis engine receives the dense depth map and the obstacle initial position coordinate set, and calculates the obstacle dynamic threat assessment index in combination with the real-time flight attitude parameters of the UAV; Based on the dynamic threat assessment indicators of obstacles, the probability distribution of the obstacle's future movement path is calculated through the spatiotemporal trajectory prediction model; The probability distribution of the future movement path of the obstacle is integrated with the preset navigation path planning data of the UAV to generate a three-dimensional obstacle avoidance heading correction vector; The three-dimensional obstacle avoidance heading correction vector is converted into a flight control instruction set and transmitted to the UAV power system execution module.

[0005] Preferably, the left view and right view original data streams contain RGB color channel information and infrared thermal imaging data, and the real-time flight attitude parameters of the UAV include pitch angle, roll angle, yaw angle, altitude and ground speed vector.

[0006] Preferably, the extracting stereoscopic visual feature representation in the binocular image by using parallel convolutional coding branches includes: Perform multi-scale feature pyramid convolution operations on the left view original data stream to generate a multi-level feature tensor sequence for the left view; Perform feature enhancement processing of the cross-channel attention mechanism on the right view original data stream to generate a right view optimized feature tensor sequence; The multi-level feature tensor sequence of the left view and the optimized feature tensor sequence of the right view are aligned with each other through cross-view feature matching to form a stereo vision feature representation matrix that integrates binocular disparity.

[0007] Preferably, the dynamic obstacle analysis engine receives the dense depth map and the obstacle initial position coordinate set, and calculates the obstacle dynamic threat assessment index in combination with the real-time flight attitude parameters of the UAV, including: Analyze the spatial distribution density of the initial position coordinate set of obstacles in the dense depth map; The angle between the ground speed vector direction in the real-time flight attitude parameters of the associated UAV and the normal vector of the obstacle surface; According to the spatial distribution density of obstacles and the angle of surface normal vectors, combined with the preset collision risk level mapping table, the obstacle dynamic threat assessment index is generated.

[0008] Preferably, the process of generating a dense depth map of the flight environment by the three-dimensional space reconstruction module includes: Decompose the stereo vision feature representation matrix into disparity feature channels and texture feature channels; Perform sub-pixel interpolation calculation on the disparity feature channel through the spatial rasterization algorithm to construct the initial depth probability distribution field; The edge gradient information in the texture feature channel is fused to perform noise filtering optimization on the initial depth probability distribution field and output a dense depth map.

[0009] Preferably, the stereoscopic visual feature representation is input into a three-dimensional space reconstruction module to generate a dense depth map of the flight environment, including: Extract the obstacle's historical movement trajectory fragment sequence from the dense depth map; A gated recurrent unit network is used to perform temporal modeling on the sequence of obstacle historical movement trajectory fragments to generate a latent variable representation of the obstacle's motion state. The latent variable representation of the obstacle's motion state is input into the conditional random field model to predict the probability distribution heat map of the obstacle's future movement path.

[0010] Preferably, the integration of the probability distribution of the future movement path of the obstacle and the preset navigation path planning data of the UAV includes: Convert the probability distribution heat map of the obstacle's future movement path into a three-dimensional space occupancy grid model; Mark the key waypoints of the path in the preset navigation path planning data of the UAV; The waypoint avoidance cost function value is calculated based on the spatial overlap area of ​​the three-dimensional space occupancy grid model and the set of key waypoints on the path.

[0011] Preferably, generating a three-dimensional obstacle avoidance heading correction vector includes: Select the set of alternative headings with the lowest cost according to the result of sorting the waypoint avoidance cost function values; Combined with the altitude constraint in the real-time flight attitude parameters of the UAV, the vertical feasibility of the alternative heading set is verified; The optimal three-dimensional obstacle avoidance heading correction vector is selected from the verified alternative heading set through the Bayesian decision model.

[0012] Preferably, converting the three-dimensional obstacle avoidance heading correction vector into a flight control instruction set includes: Decompose the three-dimensional obstacle avoidance heading correction vector into pitch control quantity, roll control quantity and yaw control quantity; According to the UAV power system dynamics model, the pitch control variable is converted into the rotor speed adjustment instruction; Convert the roll control quantity into the rudder deflection angle instruction; The yaw control quantity is converted into the tail thruster thrust vector instruction and combined to form a flight control instruction set.

[0013] Preferably, the downloading to the UAV power system execution module includes: Transmitting flight control instruction sets to the rotor speed controller, control surface servo controller, and tail thruster controller via the flight control bus; The rotor speed controller adjusts the motor drive current according to the rotor speed adjustment command; The rudder servo controller drives the servo actuator according to the rudder deflection angle instruction; The tail thruster controller adjusts the thruster nozzle direction and fuel supply according to the thrust vector command.

[0014] Compared with the prior art, the present invention has the following beneficial effects: The binocular vision acquisition terminal synchronously captures the raw data streams of the left and right views, providing accurate stereoscopic visual input for subsequent depth calculations and ensuring the integrity of environmental perception. The parallel convolutional coding branch of the deep neural network model efficiently extracts stereoscopic visual feature representations from binocular images. Compared to traditional feature matching methods, this reduces feature point loss or mismatches, improving the stability and robustness of feature extraction. The 3D spatial reconstruction module generates a dense depth map and a set of initial obstacle position coordinates based on extracted stereoscopic visual features. This not only provides information about the 3D structure of the environment but also accurately locates the initial positions of obstacles, laying the foundation for subsequent dynamic analysis. The dynamic obstacle analysis engine combines the drone's real-time flight attitude parameters to calculate dynamic obstacle threat assessment indicators. This combines the obstacle's motion state with the drone's own navigation status, making the threat assessment more accurate for the actual flight scenario and avoiding the one-sidedness of assessments based solely on obstacle location. The spatiotemporal trajectory prediction model calculates the probability distribution of an obstacle's future movement path based on dynamic threat assessment indicators. This model can predict the obstacle's movement trend in advance, allowing the drone sufficient reaction time. This probability distribution is then combined with the drone's preset navigation path planning data to generate a three-dimensional obstacle avoidance heading correction vector. This ensures that the obstacle avoidance path accounts for the dynamic changes of obstacles while conforming to the drone's preset navigation objectives, enhancing the rationality and consistency of obstacle avoidance decisions. The three-dimensional obstacle avoidance heading correction vector is converted into a flight control instruction set and transmitted to the execution module, completing a closed loop from environmental perception to motion control, ensuring the real-time and accurate obstacle avoidance actions. Overall, this method enables more comprehensive environmental perception, more accurate threat assessment, and more rational path planning in complex and dynamic environments, improving the reliability and adaptability of autonomous obstacle avoidance for drones and expanding their application possibilities in a variety of complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a timing diagram of the autonomous obstacle avoidance method for a UAV based on deep learning and binocular vision according to the present invention; Figure 2 Flowchart for stereo vision feature extraction; Figure 3 Flowchart for dense depth map generation; Figure 4 Flowchart for path fusion and avoidance cost calculation; Figure 5 Flowchart for flight control instruction set conversion. DETAILED DESCRIPTION

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0017] See also Figure 1 The present invention provides a method for autonomous obstacle avoidance of a UAV based on deep learning and binocular vision, the method comprising: The binocular vision acquisition terminal installed on the UAV body can synchronously capture the left view original data stream and the right view original data stream in the flight environment.

[0018] A deep neural network model is constructed that synchronously receives the acquired left and right view raw data streams. The model includes parallel left and right view convolutional encoding branches, which process the corresponding data streams, extracting and fusing them to form a stereoscopic feature representation representing the environment's geometry.

[0019] The output stereoscopic feature representation is fed into the 3D space reconstruction module. This module processes the stereoscopic feature representation, generates a dense 3D depth map of the flight environment, and identifies and outputs the initial position coordinates of obstacles in the environment based on this depth map.

[0020] The dynamic obstacle analysis engine receives the output of the dense depth map and the initial obstacle position coordinates. It also connects to the flight attitude parameters transmitted in real time by the drone's flight control system and comprehensively calculates the dynamic threat assessment index for each obstacle.

[0021] The spatiotemporal trajectory prediction model receives the output of the obstacle dynamic threat assessment index. Based on current and historical obstacle information, the model estimates the possible movement paths of obstacles within the future time window and their corresponding probability distribution.

[0022] The output of the obstacle's future movement path probability distribution is integrated with the navigation path planning data preset by the UAV mission system. Using a path conflict resolution algorithm, a three-dimensional obstacle avoidance heading correction vector is calculated to guide the UAV to avoid obstacles.

[0023] The flight control command converter receives the generated three-dimensional obstacle avoidance heading correction vector, parses it, and converts it into a specific flight control command set that the UAV power system can recognize and execute. This command set is transmitted to the UAV power system execution module via the airborne communication link.

[0024] Example 1: The binocular vision acquisition terminal comprises two independent but strictly synchronized image acquisition units, deployed on the left and right sides of the front of the drone. Each unit utilizes a global shutter CMOS image sensor with high dynamic range to adapt to complex lighting environments. The left image sensor generates the left view raw data stream, while the right image sensor generates the right view raw data stream. The two sensors are connected via a hardware synchronization signal line to ensure that the exposure start time and exposure duration are perfectly aligned, eliminating parallax calculation errors caused by time desynchronization. The sensor lenses use a fixed-focal-length wide-angle lens, providing sufficient horizontal and vertical field of view to effectively capture forward and lateral obstacles. The image sensor output data streams contain standard RGB three-channel color information, with each channel having a specific bit depth to preserve rich color details. Furthermore, each image sensor module integrates an uncooled microbolometer array that independently captures the scene's infrared radiation information and converts it into infrared thermal imaging data points aligned with the spatial resolution of the visible light image. These infrared data points are output synchronously with the RGB data streams to form the left and right view raw data streams containing multispectral information. The hardware interface of the binocular vision acquisition terminal transmits the raw data stream to the subsequent processing module through a high-speed serial bus.

[0025] The drone's real-time flight attitude parameters are continuously provided by an onboard multi-sensor fusion system. The core of this system consists of a high-precision micro-electromechanical system (MEMS) inertial measurement unit (IMU) and a multi-frequency, multi-mode GNSS receiver. The IMU, equipped with a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer, continuously measures angular velocity and linear acceleration in the aircraft's coordinate system. The GNSS receiver continuously receives signals from multiple satellite constellations to calculate the drone's absolute position. The sensor fusion algorithm in the flight control system processes the raw data from the IMU and GNSS receiver in real time. Based on a Kalman filter framework, this algorithm combines the advantages of both sensors: the high-frequency dynamic response characteristics of the IMU and the absolute positioning accuracy and low-frequency stability of the GNSS. The fusion solution continuously outputs a set of accurate real-time UAV flight attitude parameters. These parameters include the pitch angle, which describes the drone's rotation about its X-axis and reflects the tilt of the nose relative to the horizontal; the roll angle, which describes the drone's rotation about its Y-axis and reflects the height difference between the left and right sides of the drone relative to the horizontal; and the yaw angle, which describes the drone's rotation about its vertical Z-axis and indicates the horizontal orientation of the nose. Altitude parameters are derived by fusing barometric altitude data measured by a barometer with ellipsoidal height data provided by the Global Navigation Satellite System (GNSS), eliminating the effects of temperature drift and pressure variations that can occur when relying solely on the barometer. The ground velocity vector is derived by combining Doppler-shifted velocity measurements from the GNSS with radial velocity data from an onboard micro-Doppler radar. The Doppler radar is particularly helpful in providing velocity estimates during brief satellite signal outages. The fusion algorithm ultimately outputs a ground velocity vector consisting of three velocity components (northward velocity, eastward velocity, and celestial velocity) and a resultant motion direction. This vector describes the drone's instantaneous motion relative to the ground. All of these flight attitude parameters—pitch, roll, yaw, fused altitude, and the ground speed vector, which includes three-dimensional velocity components and directions—are transmitted to the dynamic obstacle analysis engine at a constant high frequency (e.g., 100 Hz) via a dedicated high-speed, low-latency data bus. The bus protocol design ensures data integrity and accurate timestamps, enabling subsequent modules to obtain attitude information that is strictly time-aligned with the visual data stream.

[0026] In actual operation of the binocular vision acquisition terminal, the optical characteristics of the two image sensors, including lens distortion coefficients, focal length, principal point coordinates, and baseline distance between the sensors, are precisely calibrated before shipment. These calibration parameters are stored in non-volatile memory and used for online geometric correction within the image processing pipeline to eliminate the impact of lens distortion on stereo matching accuracy. This calibration process ensures epipolar constraints between the left and right views, providing a geometric foundation for subsequent stereo feature extraction. Image sensor parameters such as exposure time and gain are dynamically adjusted by the flight control system based on ambient lighting conditions, implemented by a dedicated automatic exposure control algorithm. The goal is to obtain image data with sufficient contrast and signal-to-noise ratio in various lighting scenarios. In low-light environments, the importance of infrared thermal imaging data points becomes even greater, providing scene structure information that is independent of visible light. Data stream transmission utilizes an efficient image compression and packetization protocol, optimizing bus bandwidth usage while ensuring that critical feature information is not lost. The physical structure of the binocular vision acquisition terminal incorporates vibration resistance and temperature drift compensation to adapt to the vibration environment and temperature fluctuations experienced during UAV flight and maintain optical alignment stability. The attitude parameter fusion system continuously monitors the health and signal quality of each sensor. Upon detecting a sensor failure or signal degradation, it automatically switches fusion strategies or provides state estimation in a degraded mode, maintaining reliable perception of the UAV's attitude by the flight control system. The timestamps of the attitude parameters and the binocular image data stream are strictly synchronized via the system clock, typically using a precise hardware time synchronization protocol. This allows the dynamic obstacle analysis engine to accurately correlate the attitude parameters at a specific moment with the visual data acquired at that moment.

[0027] Example 2: See Figure 2The deep neural network model uses a dual-branch parallel architecture to process the raw data stream transmitted by the binocular vision acquisition terminal. The left-view convolutional encoding branch is based on a residual network architecture and consists of five consecutive processing stages. Each stage consists of multiple convolutional layers, batch normalization layers, and activation function layers, and includes cross-layer connections. The input left-view raw data stream first undergoes preprocessing, including pixel value normalization and spatial resizing. The first stage performs a standard convolution operation on the resized left view, outputting an initial feature map. Each subsequent stage performs a feature extraction process with decreasing spatial resolution based on the output of the previous stage. The second stage downsamples the feature map to extract a feature representation with reduced spatial information but increased channel dimensionality. The third stage further reduces the spatial resolution of the feature map while increasing the number of channels to capture more abstract image features. The fourth stage maintains a similar processing logic, outputting a feature tensor with medium resolution. The fifth stage, as the final processing layer, outputs a deep feature representation with the lowest spatial resolution but the largest number of channels. The outputs of the five stages together constitute a multi-level feature tensor sequence for the left view, where the feature maps of each level have different spatial scales and semantic abstraction levels, corresponding to local details, regional structures, and global context information in the image, respectively.

[0028] The right-view convolutional encoding branch also contains a basic convolutional architecture, but with a cross-channel attention module embedded after each convolutional stage. This module's processing flow is as follows: A global average pooling operation is performed on the input feature map to generate a compressed channel description vector. This vector is input into a two-layer fully connected network with a nonlinear activation function. The first fully connected layer reduces the channel dimension, while the second layer restores the original number of channels. After processing with a sigmoid activation function, a per-channel weight coefficient vector is generated. These weight coefficients are multiplied with the original input feature map channel by channel, enhancing the response of important channels and suppressing the influence of less important channels. The weighted feature map is then processed by subsequent convolutional layers. Similar to the left-view branch, the right-view branch also consists of five processing stages, each of which outputs a feature map enhanced by channel attention, ultimately forming a sequence of optimized feature tensors for the right view. Each processing stage in the left and right branches maintains consistent spatial resolution, facilitating subsequent cross-view information fusion.

[0029] The cross-view feature matching and alignment process uses a hierarchical strategy, sequentially processing feature maps of corresponding resolutions from the multi-level feature tensor sequence of the left view and the refined feature tensor sequence of the right view. For each pair of left and right feature maps of the same spatial resolution, a pixel-level correlation calculation is performed. Specifically, at each spatial location in the left feature map, the dot product similarity between its feature vector and the feature vectors of all possible matching locations on the corresponding epipolar line of the right feature map is calculated. This operation is performed within the entire search range of the right epipolar line, generating initial correlation volumes. Feature maps at different levels are processed using different disparity search ranges: deep, low-resolution feature maps handle larger disparity ranges, while shallow, high-resolution feature maps handle finer disparity adjustments. These initial correlation volumes are input to a cost aggregation network. This network consists of multiple 3D convolutional layers that filter and smooth the correlation volumes in both spatial and disparity dimensions, aggregating contextual information and suppressing noise and inconsistencies. The output of the cost aggregation is processed by a disparity regression layer to generate an initial disparity estimate. Subsequently, a cascaded upsampling module progressively upsamples the low-resolution disparity map to a higher resolution, simultaneously integrating contextual information from both high- and low-level feature maps. Finally, the disparity information of all levels is fused to form a stereoscopic visual feature representation matrix containing rich geometric and contextual information. This matrix encodes the three-dimensional structure information of the scene in the form of a dense feature map.

[0030] The dynamic obstacle analysis engine receives the dense depth map and initial obstacle coordinates output by the 3D spatial reconstruction module. The engine first analyzes the spatial distribution characteristics of the initial obstacle coordinates. The algorithm, centered around each obstacle coordinate point, calculates the number of other obstacle coordinates within a predetermined radius. A Gaussian kernel function is used to weight the density of points within the neighborhood, with points closer to the center receiving higher weights and those farther away receiving lower weights. The statistical distribution of the density values ​​for all obstacle neighborhoods is calculated to identify areas of high density. The engine also receives real-time drone flight attitude parameters, particularly the ground velocity vector, which includes three-dimensional velocity components and direction of motion. Based on the dense depth map, the engine uses a surface normal estimation algorithm to calculate the direction of the surface normal for each obstacle. Within each pixel neighborhood in the depth map, the algorithm estimates the local surface orientation through least-squares plane fitting, resulting in a surface normal vector. The engine then calculates the spatial angle between the drone's ground velocity vector and the obstacle surface normal vector. This angle is calculated using the vector dot product formula, resulting in a cosine value. A preset collision risk level mapping table is stored in the engine's non-volatile memory. This mapping table is a two-dimensional lookup table, with one dimension representing the intervals of the spatial density of obstacles and the other representing the intervals of the cosine of the angle between the ground velocity vector and the surface normal. Each cell stores a corresponding quantized threat level. The engine locates the corresponding cell in the mapping table based on the spatial density and cosine of the angle calculated for the current obstacle point. If the calculated value falls exactly at the center of the cell, the threat level for that cell is directly read; if the calculated value falls between cells, the threat level values ​​of adjacent cells are bilinearly interpolated. The engine independently calculates and outputs a scalar value for each obstacle point, namely the obstacle dynamic threat assessment index. This index is transmitted to subsequent modules along with the corresponding obstacle's location identifier and timestamp. The engine's internal timing synchronization mechanism ensures strict timestamp alignment of the input depth map data and attitude parameter data, ensuring that all calculations are based on the same perception state at the same moment.

[0031] Example 3: See Figure 3The 3D spatial reconstruction module receives a stereo feature representation matrix from a deep neural network model, which contains encoded scene geometry and texture information. The module first processes the input features through a set of separable convolutional layers. These convolutional layers contain 1x1 convolution kernels, which decouple mixed features. One convolutional path specifically extracts disparity-related geometric features, forming a disparity feature channel; the other path focuses on the image's structural texture information, forming a texture feature channel. The disparity feature channel is processed using a spatial rasterization algorithm: the continuous disparity range is divided into fine, discrete grid cells, each representing a specific disparity range. The algorithm constructs a probability distribution model in the disparity dimension. For each pixel position on the image plane, the probability of belonging to a different disparity grid cell is calculated. This calculation utilizes a feature similarity metric within a local window and incorporates spatial constraints between adjacent pixels. After the initial disparity probability distribution is generated, it is processed using sub-pixel interpolation techniques. Bicubic interpolation is applied to the disparity probability distribution field, smoothing the probability distribution curve and improving the spatial resolution of the depth estimate, resulting in an initial depth probability distribution field. The distribution field reflects the preliminary probability that each point in the scene is at different depth values.

[0032] The processing of the texture feature channel is performed independently: the directional gradient operator is applied to scan the entire feature map, and the rate of change of intensity in the horizontal and vertical directions of each pixel position is calculated. The gradient components in these two directions are combined to calculate the gradient amplitude of each pixel point and generate a gradient amplitude map. The gradient amplitude map reveals the position information of the edges of objects and texture-rich areas in the scene. Based on the gradient amplitude map, an adaptive filter kernel is constructed. The center weight of the filter kernel is inversely proportional to the gradient amplitude of the current position. This means that in the edge area (high gradient amplitude), the filter kernel weight is small to protect the edge sharpness; in the flat area (low gradient amplitude), the filter kernel weight is large to enhance the smoothing effect. This adaptive filter kernel is applied to the initial depth probability distribution field and a convolution operation is performed. This operation can be expressed in mathematical form:

[0033] in, Represents the depth probability value of the depth d at the optimized image coordinate (x, y); Represents the probability value of depth d at coordinate (x+i,y+j) in the initial depth probability distribution field; Represents the weight value of the adaptive filter kernel centered at (x, y) at offset (i, j); the weight value is determined by the gradient amplitude at coordinate (x, y) Determine; the summation range is within the neighborhood window defined by the filter kernel size (for example, This filtering process significantly suppresses noise points in the initial depth probability distribution field caused by texture deficiency or mismatching, while preserving depth discontinuities near object boundaries. After multiple iterations of optimization, the module outputs a high-confidence dense depth map, in which each pixel is associated with a specific depth value.

[0034] The spatiotemporal trajectory prediction module takes a dense depth map sequence as input source. The module continuously monitors the connected areas in the depth map that are identified as obstacles. For each tracked obstacle, the module records the three-dimensional coordinates of its center point in continuous time frames. These coordinate points are connected in chronological order to form a sequence of historical movement trajectory fragments of the obstacle. Each trajectory fragment contains the position change data of the obstacle within a fixed time window (for example, the past 1 second, corresponding to 10 frames of depth map). A gated recurrent unit network is used to perform temporal modeling of these trajectory fragments. The network input is a sequence of coordinate increments of consecutive time steps in the trajectory fragment (i.e. , representing the displacement of the obstacle in three-dimensional space from time t-1 to time t). The gated recurrent unit network consists of two hidden layers. The first layer of units processes the input displacement sequence and learns the instantaneous velocity change pattern of the obstacle. The second layer of units receives the output of the first layer and learns the motion trends and pattern characteristics over a longer time span. The gating mechanism (update gate and reset gate) within the unit dynamically controls the retention and forgetting ratio of historical state information. The network ultimately outputs a fixed-dimensional state vector, which serves as the latent variable representation of the obstacle's motion state and encodes the obstacle's current motion characteristics (such as speed, acceleration, directional trend) and its historical motion pattern.

[0035] The conditional random field model receives the aforementioned latent variable representation of the obstacle's motion state as observational evidence. The model constructs a probabilistic graphical model for a sequence of discrete future time points (e.g., predicting the next 2 seconds with a step size of 0.1 seconds, for a total of 20 time nodes). Each time node corresponds to the possible 3D position of the obstacle at a specific future moment. The model defines two core elements: a single-node potential function and a transition potential function between adjacent nodes. The single-node potential function is calculated based on the latent variable representation of the obstacle's motion state and describes the probability of the obstacle being at a certain position at each future time node. Specifically, a fully connected neural network is used to map the motion state latent variables into a prior estimate of the probability distribution of the position at each future time node. The transition potential function defines the probability of state transitions between consecutive time nodes, modeling the continuity and inertia of the obstacle's motion. The transition potential function generally favors state transitions with minimal position change between adjacent moments. The model uses a belief propagation algorithm for probabilistic inference. This algorithm iteratively passes messages through the probabilistic graphical model, from the observation node (the motion state latent variable) to the future time node and between adjacent future nodes. After multiple rounds of message passing, the probability distribution of the obstacle's position at each future time point gradually converges. Ultimately, the model integrates the position probability distributions of all future time points to generate a four-dimensional probability distribution heat map, combining three spatial dimensions with the time dimension. In this heat map, the brightness of each voxel (containing the x, y, and z coordinates) at a specific future time point t represents the probability of the obstacle appearing at that spatial location at that moment. This heat map is the quantified output of the probability distribution of the obstacle's future movement path.

[0036] Example 4: See Figure 4 The heat map of the probability distribution of the obstacle's future movement path is a four-dimensional data structure, containing three-dimensional spatial coordinates (X, Y, Z) and the time dimension T. The conversion module discretizes the heat map into a three-dimensional spatial occupancy grid model: the spatial voxel resolution is set to 0.2 meters (that is, each cube grid has a side length of 0.2 meters) and the time step is 0.2 seconds. For each future time slice (such as t = 0.2s, 0.4s, ... 2.0s), all spatial grid cells are traversed. If the probability value of the corresponding spatial position in the heat map at time t exceeds the threshold of 0.3, the grid is marked as "potentially occupied" and its occupancy probability value is stored. Finally, a set of three-dimensional grid maps arranged in time series is formed.

[0037] The drone's preset flight path planning data is defined as a sequence of waypoints. Each waypoint contains three-dimensional coordinates and an estimated arrival timestamp. The module identifies path turning points and mission critical points, forming a set of key waypoints. Table 1 shows some of the key waypoint attributes in the example scenario.

[0038] Table 1: Some key waypoint attributes in the example scene Waypoint ID Coordinate X (meters) Coordinate Y (meters) Coordinate Z (meters) Estimated time of arrival (seconds) WP1 120.5 85.2 50.0 3.2 WP2 135.7 93.8 52.1 4.1 WP3 148.9 110.5 55.3 5.0 WP4 160.2 125.7 58.0 6.0 WP5 172.8 140.3 60.0 7.2 The waypoint avoidance cost calculation process is as follows: For each point in the key waypoint set (such as WP3), determine its time impact window (for example, estimated time of arrival ± 1 second, i.e. t = 4.0-6.0 seconds). Within this time window, extract all corresponding spatial occupancy grid slices. Calculate the Euclidean distance between the waypoint coordinates and the centers of all grids marked as "potentially occupied". Record the minimum distance value. The cost function is defined as a piecewise function: when When the value is 0, the cost is 0; when rice Meter-hour, cost value ; when Meter-hour, cost value .

[0039] Therefore, in this example, if the distance to the nearest obstacle for WP3 at t=5.0 seconds is 1.5 meters, its avoidance cost is .

[0040] The three-dimensional obstacle avoidance heading correction vector generation process consists of four stages: Candidate heading sampling: Based on the current heading, the heading angle combination (yaw angle) is uniformly sampled at 5-degree intervals within the range of ±30 degrees in the horizontal plane and ±15 degrees in the vertical plane. , pitch angle ), generate an initial candidate heading set.

[0041] Trajectory simulation and cost summation: For each candidate heading (such as , ), based on the UAV’s current velocity dynamics model, simulate its flight trajectory and predict the actual time it takes to pass each key waypoint. Calculate the sum of the avoidance cost function values ​​for all waypoints along the trajectory.

[0042] Altitude constraint verification: Checks whether the predicted flight altitude corresponding to the candidate headings is within the safety boundary (e.g., minimum safe altitude 40 meters, maximum altitude 80 meters). Eliminates headings that cause the altitude to exceed the range of [40, 80] meters.

[0043] Bayesian decision screening: A decision model is constructed, taking as input a feature vector—including the total cost of candidate headings, angular deviation from the preset path, and altitude change rate. A likelihood probability model trained on historical flight data sets is used to assess the feasibility of each candidate heading. The heading with the highest probability is selected as the final 3D obstacle avoidance heading correction vector.

[0044] Specific example: A moving vehicle obstacle is detected during a flight, and its high probability area in the next 2 seconds covers the vicinity of waypoint WP3. , ) The straight path results in a WP3 avoidance cost of 10. Among the candidate courses: Heading A ( , ): Total cost 3.5, height 55 meters, angle deviation medium Course B ( , ): Total cost 1.2, height 38 meters (violating height constraint) Course C ( , ): Total cost 0.8, height 53 meters, small angle deviation The Bayesian model calculates the feasibility probability of course C to be 0.92 (the highest), so the correction vector is output , This vector guides the drone to deflect to the upper right, avoiding areas with dense obstacles while maintaining a safe altitude.

[0045] After receiving the vector, the flight control command converter decomposes it according to the body coordinate system: Yaw correction Convert the tail thruster nozzle to 15° left and increase thrust by 12% Pitch correction The forward rotor group speed is reduced by 5%, and the rear rotor group speed is increased by 5%. Specific execution instructions are generated and transmitted via the control bus. The entire process is completed within 200 milliseconds, achieving dynamic obstacle avoidance path replanning.

[0046] Example 5: See Figure 5 The three-dimensional obstacle avoidance heading correction vector consists of three independent component parameters: the pitch control variable represents the incremental rotation angle around the lateral axis; the roll control variable represents the incremental rotation angle around the longitudinal axis; and the yaw control variable represents the incremental rotation angle around the vertical axis. After receiving this vector, the flight control command converter initiates a multi-channel parallel processing mechanism. For the pitch control component, the converter invokes a pre-set rotor dynamics model. This model establishes a mapping relationship between pitch angle changes and lift differences between the rotor groups. The model calculates the pitch moment required to maintain the target pitch attitude and, based on the current flight speed and air density parameters, derives the lift difference between the left and right symmetrical rotor groups. Based on the rotor aerodynamic characteristics database, this lift difference is converted into a specific rotor speed adjustment command value. This command explicitly specifies the target speed adjustment value for the left main rotor group and the target speed adjustment value for the right main rotor group, with opposite signs, forming differential control. The speed adjustment value is expressed as an absolute speed value or a relative percentage.

[0047] The roll control component is input into the control surface aerodynamic effect model for processing. This model stores the correspondence between the control surface deflection angle and the roll moment coefficient under different flight conditions. The model performs real-time interpolation calculations based on the current airspeed and angle of attack parameters to determine the control surface deflection angle required to generate the target roll moment. For fixed-wing drones using aileron control, the output is the left and right aileron differential deflection command; for multi-rotor configurations, the output is the differential lift adjustment command for the specific rotor group. The command clearly specifies the target deflection angle of the left and right control surfaces, which form an asymmetric deflection. The deflection angle is expressed as a mechanical angle value, accurate to 0.1 degree.

[0048] The yaw control component is processed by the thrust vector control model. This model consists of two parallel computational units: a thrust magnitude calculation unit and a nozzle direction calculation unit. The thrust magnitude calculation unit calculates the required absolute thrust value based on the target yaw moment and current flight altitude, combined with the thrust-fuel consumption characteristic curve of the thruster. The nozzle direction calculation unit solves the nozzle vector angle required to produce the target yaw moment based on the spatial geometry of the thruster installation position. The model output consists of two components: a thrust adjustment command value and a nozzle deflection command value. The thrust adjustment command specifies the thrust target value or throttle valve opening percentage; the nozzle deflection command specifies the target deflection direction and angle. After all control components are converted, the command converter encapsulates the rotor speed control command, the control surface deflection angle command, and the thrust vector command according to the standard protocol to form a complete flight control command set data package.

[0049] This instruction set is transmitted to the execution module via a high-speed flight control bus. The bus utilizes a time-triggered communication protocol to ensure real-time and deterministic command transmission. The rotor speed controller, acting as a bus node, receives rotor speed adjustment commands. The controller incorporates a built-in closed-loop speed control algorithm: It uses Hall sensors to collect the actual rotor speed in real time and compares it with the commanded target value to generate a speed error signal. This error signal is fed into a proportional-integral controller, which outputs an adjustment for the motor drive current. After amplification by the power driver module, the current adjustment signal drives the three-phase windings of the brushless motor, achieving precise rotor speed tracking. The rudder servo controller receives rudder deflection angle commands. The controller uses a high-precision potentiometer or optical encoder to determine the actual rudder position, which is compared to the commanded angle to form a position deviation. This position deviation is processed by the servo control law to generate a pulse-width modulated signal to drive the servo actuator. The servo drives the rudder to the target angle via a gear transmission mechanism, with position feedback forming a closed-loop control loop. The tail thruster controller simultaneously processes both the thrust and directional components of the thrust vectoring command. Based on the thrust adjustment command, the thrust control unit adjusts the propellant flow rate through the fuel metering valve to achieve precise thrust control. The direction control unit drives the stepper motor or hydraulic actuator according to the nozzle deflection command to adjust the thruster nozzle vector direction. The nozzle angle is detected and fed back in real time by a rotary encoder, forming an angle closed-loop control.

[0050] Each execution controller continuously monitors its execution status during command execution. The rotor speed controller provides feedback on actual speed, motor current, and temperature parameters; the rudder servo controller provides feedback on actual rudder angle and servo load torque; and the thruster controller provides feedback on actual thrust, nozzle angle, and fuel flow. These status parameters are transmitted back in real time to the flight control command converter via a feedback channel. The converter compares and verifies the actual execution status with the original command. If it detects that the execution deviation exceeds the tolerance range, it activates the command compensation mechanism or fault handling procedure. The entire command conversion and execution process is completed within strict time constraints, ensuring the real-time responsiveness of obstacle avoidance control. Execution status data is also recorded in the flight data recorder for subsequent analysis and maintenance. All controllers have fault detection capabilities. In the event of abnormal conditions such as sensor failure or actuator jamming, they can automatically switch to a degraded control mode and issue an alarm signal to maintain basic controllability of the drone.

[0051] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0052] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for autonomous obstacle avoidance of UAV based on deep learning and binocular vision, characterized in that: The following steps are involved: The binocular vision acquisition terminal synchronously obtains the left view and right view raw data streams in the flight environment; The deep neural network model synchronously processes the left view and right view raw data streams, and extracts stereo visual feature representations from the binocular images through parallel convolutional coding branches; Inputting the stereoscopic visual feature representation into a three-dimensional space reconstruction module to generate a dense depth map of the flight environment and a set of initial position coordinates of obstacles; The dynamic obstacle analysis engine receives the dense depth map and the obstacle initial position coordinate set, and calculates the obstacle dynamic threat assessment index in combination with the real-time flight attitude parameters of the UAV; Based on the dynamic threat assessment indicators of obstacles, the probability distribution of the obstacle's future movement path is calculated through the spatiotemporal trajectory prediction model; The probability distribution of the future movement path of the obstacle is integrated with the preset navigation path planning data of the UAV to generate a three-dimensional obstacle avoidance heading correction vector; The three-dimensional obstacle avoidance heading correction vector is converted into a flight control instruction set and transmitted to the UAV power system execution module.

2. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 1 is characterized in that: The left view and right view raw data streams contain RGB color channel information and infrared thermal imaging data, and the real-time flight attitude parameters of the UAV include pitch angle, roll angle, yaw angle, altitude and ground speed vector.

3. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 2, characterized in that: The method of extracting stereoscopic visual feature representation from a binocular image by using parallel convolutional coding branches includes: Perform multi-scale feature pyramid convolution operations on the left view original data stream to generate a multi-level feature tensor sequence for the left view; Perform feature enhancement processing of the cross-channel attention mechanism on the right view original data stream to generate a right view optimized feature tensor sequence; The multi-level feature tensor sequence of the left view and the optimized feature tensor sequence of the right view are aligned with each other through cross-view feature matching to form a stereo vision feature representation matrix that integrates binocular disparity.

4. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 3 is characterized in that: The dynamic obstacle analysis engine receives the dense depth map and the obstacle initial position coordinate set, and calculates the obstacle dynamic threat assessment index in combination with the real-time flight attitude parameters of the UAV, including: Analyze the spatial distribution density of the initial position coordinate set of obstacles in the dense depth map; The angle between the ground speed vector direction in the real-time flight attitude parameters of the associated UAV and the normal vector of the obstacle surface; According to the spatial distribution density of obstacles and the angle of surface normal vectors, combined with the preset collision risk level mapping table, the obstacle dynamic threat assessment index is generated.

5. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 4 is characterized in that: The process of generating a dense depth map of the flight environment by the 3D space reconstruction module includes: Decompose the stereo vision feature representation matrix into disparity feature channels and texture feature channels; Perform sub-pixel interpolation calculations on the disparity feature channel through a spatial rasterization algorithm to construct an initial depth probability distribution field; The edge gradient information in the texture feature channel is fused to perform noise filtering optimization on the initial depth probability distribution field and output a dense depth map.

6. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 5, characterized in that: Inputting the stereoscopic visual feature representation into a three-dimensional space reconstruction module to generate a dense depth map of the flight environment includes: Extract the obstacle's historical movement trajectory fragment sequence from the dense depth map; A gated recurrent unit network is used to perform temporal modeling on the sequence of obstacle historical movement trajectory fragments to generate a latent variable representation of the obstacle's motion state. The latent variable representation of the obstacle's motion state is input into the conditional random field model to predict the probability distribution heat map of the obstacle's future movement path.

7. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 6, characterized in that: The integration of the probability distribution of the future movement path of the obstacle and the preset navigation path planning data of the UAV includes: Convert the probability distribution heat map of the obstacle's future movement path into a three-dimensional space occupancy grid model; Mark the key waypoints of the path in the preset navigation path planning data of the UAV; The waypoint avoidance cost function value is calculated based on the spatial overlap area of ​​the three-dimensional space occupancy grid model and the set of key waypoints on the path.

8. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 7, characterized in that: Generating a three-dimensional obstacle avoidance heading correction vector includes: Select the set of alternative headings with the lowest cost according to the result of sorting the waypoint avoidance cost function values; Combined with the altitude constraint in the real-time flight attitude parameters of the UAV, the vertical feasibility of the alternative heading set is verified; The optimal three-dimensional obstacle avoidance heading correction vector is selected from the verified alternative heading set through the Bayesian decision model.

9. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 8, characterized in that: Converting the three-dimensional obstacle avoidance heading correction vector into a flight control instruction set includes: Decompose the three-dimensional obstacle avoidance heading correction vector into pitch control quantity, roll control quantity and yaw control quantity; According to the UAV power system dynamics model, the pitch control variable is converted into the rotor speed adjustment instruction; Convert the roll control quantity into the rudder deflection angle instruction; The yaw control quantity is converted into the tail thruster thrust vector instruction and combined to form a flight control instruction set.

10. The autonomous obstacle avoidance method for UAV based on deep learning and binocular vision according to claim 9, characterized in that: The transmission to the UAV power system execution module includes: Transmitting flight control instruction sets to the rotor speed controller, control surface servo controller, and tail thruster controller via the flight control bus; The rotor speed controller adjusts the motor drive current according to the rotor speed adjustment command; The rudder servo controller drives the servo actuator according to the rudder deflection angle instruction; The tail thruster controller adjusts the thruster nozzle direction and fuel supply according to the thrust vector command.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous obstacle detection system and method based on binocular vision

    CN105222760A

  • Electric power inspection unmanned aerial vehicle obstacle avoidance method based on millimeter wave radar and binocular vision

    CN115933754A

  • Unmanned aerial vehicle obstacle avoidance method based on AI

    CN118897572A

  • Unmanned aerial vehicle perception obstacle avoidance method, system, device and medium

    CN120122709A

  • Unmanned aerial vehicle obstacle avoidance control method and system based on computer vision

    CN120370998A

Cited By

  • Low-slow small target detection and trajectory prediction tracking method based on laser radar

    CN121069407A

  • Unmanned aerial vehicle route intelligent planning and obstacle avoidance method based on deep learning

    CN121070026A

  • Urban low-altitude distribution unmanned aerial vehicle self-adaptive navigation method and system

    CN122261182A