An Unmanned Aerial Vehicle Obstacle Avoidance Method and Device Based on Image Processing

Through the lightweight processing of optical flow calculation and depth estimation model, combined with obstacle avoidance waypoint decisions, the real-time and accuracy of the UAV obstacle avoidance system is solved, and the lightweight and efficient obstacle avoidance effect is achieved.

CN120236215BActive Publication Date: 2025-08-05YIFEI INTELLIGENT TECH (WUHAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510707394.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-05
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional UAV obstacle avoidance systems rely on high-performance sensors such as lidar and stereo cameras, resulting in the lightweight and long battery life of the UAV. The calculation complexity of the monocular camera depth estimation method is high and the real-time performance is insufficient. The reinforcement learning resources and energy are limited, making it difficult to meet the real-time and accuracy requirements of UAV obstacle avoidance.

Method used

A lightweight depth estimation model based on optical flow calculation is adopted, combining the generation adversarial network and the depth separable convolutional layer, through pyramid feature processing and sparse embedding of optical flow information, and combined with the obstacle avoidance waypoint decision model, the autonomous obstacle avoidance decision of the drone is realized.

Benefits of technology

It improves the real-time and accuracy of obstacle avoidance by drones, reduces the computational complexity and resource requirements, and achieves a lightweight obstacle avoidance effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236215B_ABST
    Figure CN120236215B_ABST
Patent Text Reader

Abstract

The present invention discloses an obstacle avoidance method and device for an unmanned aerial vehicle based on image processing, belonging to the technical field of unmanned aerial vehicle obstacle avoidance. The method includes: performing multi-scale processing on the current frame and the next frame in the monocular RGB sequence to obtain corresponding pyramid feature maps, performing optical flow estimation on the pyramid feature maps to obtain an optical flow map, and embedding the brightness information of the optical flow map into the current frame to obtain a fused frame, using a depth estimation model to process the fused frame to obtain a corresponding depth map, and based on the depth map, using an obstacle avoidance waypoint decision model to make an obstacle avoidance decision for the unmanned aerial vehicle. The present invention can improve the real-time performance and accuracy of the obstacle avoidance operation of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of UAV obstacle avoidance, and particularly relates to a UAV obstacle avoidance method and device based on image processing. Background Technique

[0002] Traditional UAV obstacle avoidance systems mainly rely on high-performance sensors such as lidar (LiDAR) and stereo cameras. Although these sensors can provide high-precision depth information and three-dimensional environment perception, due to their large size, heavy weight, high power consumption, and high cost, they are not suitable for being carried on small rotor UAVs, which limits the application potential of UAVs in terms of light weight and long endurance. In contrast, a monocular camera has become an ideal choice for small UAV environmental perception due to its small size, light weight, low power consumption, and relatively low cost. However, a monocular camera cannot directly obtain depth information and must achieve depth estimation through complex algorithm processing. Traditional monocular depth estimation methods, such as those based on geometric calculations, are often affected by problems such as high computational complexity and limited accuracy, and it is difficult to meet the high requirements for real-time performance and accuracy of UAVs during high-speed flight.

[0003] In recent years, with the rise of deep learning technology, monocular depth estimation methods based on generative adversarial networks (GANs) and optical flow technology have gradually become a research hotspot. Through the adversarial training of a generator and a discriminator, GANs can generate realistic depth maps from monocular images, providing a new way for UAVs to obtain depth information. At the same time, optical flow technology can provide important clues about object motion and scene structure by analyzing the pixel motion between consecutive image frames, which helps to improve the accuracy and robustness of depth estimation. However, applying GANs and optical flow technology to the field of UAV obstacle avoidance still faces many challenges. First, the computational complexity of the GANs model is relatively high, posing a high requirement for the computing power of UAVs. Second, optical flow calculation also consumes a large amount of computing resources and has a bottleneck in terms of real-time performance.

[0004] In addition, reinforcement learning has shown great potential in the field of UAV autonomous navigation and obstacle avoidance. By allowing the UAV to perform trial-and-error learning in the environment and aiming to maximize the cumulative reward, it continuously optimizes its behavior strategy to achieve autonomous navigation and obstacle avoidance. However, due to the limited computing resources and energy of UAVs, algorithm complexity, overestimation, slow convergence, etc. have become the main factors restricting the application of reinforcement learning to UAV autonomous obstacle avoidance.

[0005] Therefore, there is an urgent need for a new type of UAV obstacle avoidance method and device based on image processing. This method and device can implement a lightweight monocular depth estimation model based on efficient optical flow calculation and set a reasonable reinforcement learning strategy to give accurate obstacle avoidance decisions for UAVs in real time, thereby improving the real-time performance and accuracy of UAV obstacle avoidance. Summary of the Invention

[0006] In view of the defects existing in the above-mentioned prior art, the present invention provides an obstacle avoidance method for an unmanned aerial vehicle (UAV) based on image processing. The method includes the following steps:

[0007] S1: The UAV acquires a monocular RGB image sequence, which includes a current frame and a next frame. The current frame and the next frame are respectively subjected to multi-scale processing to obtain their corresponding pyramid feature maps, and optical flow estimation is performed on the pyramid feature maps to obtain an optical flow map;

[0008] S2: Extract the luminance information in the optical flow map and embed the luminance information into the current frame to obtain a fused frame;

[0009] S3: Use a depth estimation model to process the fused frame to obtain a depth map corresponding to the fused frame; the depth estimation model includes a generator and a discriminator. The generator includes an encoder, a decoder, skip connections established between the encoder and the decoder, and an output layer. The discriminator includes multiple convolutional layers and a fully connected layer;

[0010] S4: Based on the depth map, the UAV makes an obstacle avoidance decision using an obstacle avoidance waypoint decision model;

[0011] S5: The UAV performs obstacle avoidance flight according to the obstacle avoidance decision result.

[0012] Further, in step S1, the multi-scale processing of the current frame and the next frame respectively to obtain their corresponding pyramid feature maps and the optical flow estimation of the pyramid feature maps to obtain an optical flow map includes:

[0013] Step S11, using a convolutional filter to the current frame and the next frame for multiple image downsamplings to respectively obtain their corresponding L-layer pyramid feature maps; the layer in the L-layer pyramid feature map is obtained by downsampling the layer in the L-layer pyramid feature map through the convolutional filter, ; the bottom layer of the L-layer pyramid feature map is the current frame or the next frame

[0014] Step S12, at the l layer, using the optical flow upsampled by 2 times from the l +1 layer to warp the features of the current frame to the next frame so that the current frame The corresponding pyramid feature map and the corresponding pyramid feature map are aligned;

[0015] Step S13, for any pixel in the pyramid feature map corresponding to the warped current frame Calculate the correlation matching cost between it and the corresponding pixel in the pyramid feature map of the next frame To construct a matching cost volume;

[0016] Step S14, taking the matching cost volume, the pyramid feature map of the current frame, and the upsampled optical flow as inputs, estimate the optical flow of the current layer through a multi-layer CNN;

[0017] Step S15, use a context network to refine the estimated optical flow to obtain a more accurate optical flow estimate;

[0018] Step S16, repeat steps S12 - S15 at different layers of the pyramid feature map to gradually optimize the optical flow estimate;

[0019] Step S17, at the bottom layer of the pyramid, obtain the final optical flow map, which contains the motion vector information of each pixel point in the current frame.

[0020] Furthermore, in step S12, in the layer, use the optical flow upsampled by 2 times from the layer to warp the features of the current frame towards the next frame so that the pyramid feature map corresponding to the current frame is aligned with the pyramid feature map corresponding to the next frame Specifically:

[0021] ;

[0022] where x is the pixel index, represents the optical flow upsampled by 2 times from the layer during the warping process. At the top layer of the pyramid feature map, is set to 0; represents the first feature layer of the L-layer pyramid feature map corresponding to the current frame ; represents the result of warping .

[0023] Furthermore, in step S13, the correlation matching cost is:

[0024] ;

[0025] Among them, represents the correlation matching cost between pixel in the and in the layer pyramid feature map, represents any pixel in the l layer of the pyramid feature map corresponding to the current frame represents any pixel in the layer of the pyramid feature map corresponding to the next frame layer of the pyramid feature map corresponding to the next frame represents the layer of the pyramid feature map corresponding to the next frame layer, N is the vector length of represents the warping result of the layer of the pyramid feature map corresponding to the current frame l layer, and T represents transpose.

[0026] Further, in step S2, the extracting the luminance information in the optical flow map and embedding the luminance information into the current frame to obtain a fused frame includes:

[0027] Step S21: Converting the vector information at each pixel position in the optical flow map into scalar information, where the vector information represents the motion direction and motion speed of the pixel, and the scalar information is the luminance information of the pixel;

[0028] Step S22: Normalizing the luminance information of the optical flow map to convert it into a luminance value range suitable for embedding into the current frame;

[0029] Step S23: Embedding the luminance information of the optical flow map into the current frame at a certain pixel interval in a sparse embedding manner to obtain the fused frame.

[0030] Further, the encoder includes a plurality of depthwise separable convolutional layers, and each depthwise separable convolutional layer is followed by batch normalization and an activation function, and the feature map of the fused frame is gradually extracted through the plurality of depthwise separable convolutional layers;

[0031] A skip connection is established between the encoder and the decoder to transfer the feature maps extracted by each layer of the encoder to the corresponding layers of the decoder;

[0032] The decoder includes a plurality of depthwise separable convolutional layers, and gradually fuses the lower-layer feature maps and the feature maps transferred through the skip connection to obtain a depth map with the same size as the fused frame.

[0033] Furthermore, the loss function of the depth estimation model is:

[0034] ;

[0035] in, Denotes that the generator parameters are optimized to minimize the loss, and the discriminator parameters are optimized to maximize the loss; G denotes the generator, and D denotes the discriminator; represents the generative adversarial network loss; represents the weight factor; represents L1 loss;

[0036] The generative adversarial network loss It is obtained by the following formula:

[0037] ;

[0038] Among them, i represents the input image, gt represents the real image, represents the generated image obtained by the generator, n represents the random noise vector, represents the probability that the discriminator judges the real image as real, represents the probability that the discriminator misjudges the generated image as real, and E represents the expected value;

[0039] The L1 loss It is obtained by the following formula:

[0040] ;

[0041] Wherein, y represents the pixel value of the real image gt, Represents the generated image The corresponding pixel value of represents the L1 norm.

[0042] Furthermore, the obstacle avoidance waypoint decision model is implemented based on the deep Q network, and the reward function of the obstacle avoidance waypoint decision model is , specifically:

[0043] ;

[0044] in, is the speed reward function, is the depth reward function, is the collision reward function.

[0045] Furthermore, the speed reward function :

[0046] ;

[0047] Wherein, represents the linear velocity of the UAV, represents the angular velocity of the UAV, represents the control instruction period of the UAV, is an adjustment constant for adjusting the value range of the speed reward function;

[0048] The depth reward function :

[0049] ;

[0050] ;

[0051] Wherein, represents the absolute distance of the UAV to the nearest obstacle, is a collision threshold for distinguishing the safe area and the potential collision area, represents the relative distance of the UAV to the nearest obstacle, represents the change amount of the relative distance, which is used to encourage the UAV to stay away from the obstacle, represents an offset constant for adjusting the sensitivity value of the depth reward, represents the sign function;

[0052] In the case of a collision of the UAV, the collision reward function .

[0053] The present invention also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods are implemented.

[0054] The present invention performs optical flow calculation on monocular RGB sequences. When performing optical flow calculation, pyramid processing, warping operation, cost calculation, context processing, etc. are integrated, and combined with deep learning. While achieving lightweight optical flow calculation, it ensures the real-time performance and accuracy of optical flow calculation, making it easy to be deployed in UAV devices with limited resources. The present invention embeds the luminance information of the obtained optical flow map into the RGB image, which not only retains the visual features of the RGB image but also introduces the motion information in the optical flow map, thereby improving the accuracy of depth estimation. The present invention replaces the standard convolutional layer of the generator in the existing generative adversarial network with a depthwise separable convolutional layer, significantly reducing the computational amount and the number of parameters of the algorithm. The original encoder-decoder structure and skip connections remain unchanged, and the structure of the discriminator also remains unchanged, thus obtaining a lightweight depth estimation model. According to the above method, the present invention is improved from two aspects of lightweight optical flow calculation and lightweight depth estimation, realizing the reliable application of the generative adversarial network GAN combined with optical flow in UAV obstacle avoidance. The obstacle avoidance waypoint decision model of the present invention improves the accuracy of Q-value estimation in deep Q-learning by separating the estimation of the value and advantage functions, promoting the convergence speed and stability of reinforcement learning. The reward function of the obstacle avoidance waypoint decision model includes a speed reward function, a depth reward function, and a collision reward function, guiding the UAV to balance multiple factors such as speed, safety, and task completion during obstacle avoidance, thereby achieving autonomous obstacle avoidance and efficient flight. Description of the Drawings

[0055] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary but not restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0056] Figure 1 is a flowchart showing a UAV obstacle avoidance method based on image processing according to an embodiment of the present invention.

[0057] Figure 2 is a diagram showing the regional division of the depth reward function according to an embodiment of the present invention. Detailed Embodiments

[0058] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. "Multiple" generally includes at least two.

[0060] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe..., these... should not be limited to these terms. These terms are only used to distinguish.... For example, without departing from the scope of the embodiments of the present invention, the first... may also be referred to as the second..., and similarly, the second... may also be referred to as the first....

[0061] It should be understood that the term "and / or" used herein is only a description of the associative relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0062] Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "when...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined", "in response to determining", "when detected (stated condition or event)", or "in response to detecting (stated condition or event)".

[0063] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or device including the said element.

[0064] As Figure 1 shown, the present invention discloses a method for an unmanned aerial vehicle to avoid obstacles based on image processing. The method includes:

[0065] Step S1: The unmanned aerial vehicle acquires a monocular RGB image sequence, which includes a current frame and a next frame. The current frame and the next frame are respectively subjected to multi-scale processing to obtain their corresponding pyramid feature maps, and optical flow estimation is performed on the pyramid feature maps to obtain an optical flow map.

[0066] In this embodiment, the drone can carry a monocular RGB camera to obtain an RGB image sequence, and the specific selection depends on the task requirements and environmental conditions of the drone application.

[0067] Optical flow represents the motion information of objects in a scene and is estimated by analyzing the pixel changes between two adjacent frames of images. In two consecutive frames of images, the movement of an object causes a change in image intensity. If the time interval between the two frames is short enough, it can be assumed that the luminance of the object surface remains unchanged. Based on this assumption, by analyzing the change in image intensity between two adjacent frames, the motion speed and direction of the object, that is, the optical flow vector, can be estimated.

[0068] In this embodiment, in step S1, multi-scale processing is performed on the current frame and the next frame respectively to obtain their corresponding pyramid feature maps, and optical flow estimation is performed on the pyramid feature maps to obtain an optical flow map, including:

[0069] Step S11, using a convolutional filter for the current frame and the next frame to perform multiple image downsamplings respectively to obtain their corresponding L-layer pyramid feature maps; the th layer in the L-layer pyramid feature map is obtained by downsampling the th layer in the L-layer pyramid feature map by the convolutional filter, ; the bottom layer of the L-layer pyramid feature map is the current frame or the next frame .

[0070] Among them, the number of pyramid layers L can be set as needed. For example, L = 6 can be set, indicating that 6 layers of pyramid feature maps with different resolutions are generated. Starting from the original image size, the image resolution is gradually reduced while the number of feature channels is increased to capture image information at different scales. In this embodiment, a convolutional neural network CNN is used to extract multi-scale features of images. Starting from the bottom layer of the pyramid, the number of feature channels for each layer is 16, 32, 64, 96, 128, and 196 respectively. These feature channels provide rich image information and contribute to subsequent optical flow estimation.

[0071] Step S12, at the th layer, perform 2-fold upsampling on the optical flow from the th layer to unify the optical flow of the th layer to the feature map size of the l th layer. Distort the features of the current frame towards the next frame so that the pyramid feature map corresponding to the current frame matches the next frame Align the corresponding pyramid feature maps.

[0072] In specific implementation, bilinear interpolation, nearest neighbor interpolation or bicubic interpolation can be used to implement the above warping operation. Taking bilinear interpolation as an example, assume that the feature map of the th l layer of the current frame is , and the feature map of the th l layer of the next frame is . For each pixel point in , find the corresponding pixel point in according to its corresponding optical flow value, and use the bilinear interpolation method to obtain the target pixel value based on the values of the four known pixel points around the pixel point in . This target pixel value is the value of the pixel point in . Traverse the pixels in in this way, and finally realize warping the features of the current frame l to the next frame at the th layer.

[0073] This process is specifically expressed as:

[0074] ;

[0075] where x is the pixel index, represents the optical flow upsampled by 2 times from the layer during the warping process. At the top layer of the pyramid feature map, is set to 0; represents the first feature layer of the th layer of the pyramid feature map corresponding to the current frame; represents the result of warping .

[0076] This process is based on the basic assumption of optical flow estimation, that is, pixel points in an image move according to a certain motion pattern between different frames. Through the warping operation, this motion pattern can be explicitly expressed and facilitate subsequent optical flow estimation. This process avoids directly warping the original image and does not contain any learning and training parameters, reducing the computational amount and the size of the algorithm model, which is beneficial to realizing lightweight design.

[0077] Step S13, for the warped current frame For any pixel in the corresponding pyramid feature map, calculate the correlation matching cost between it and the corresponding pixel in the next frame to construct a matching cost volume.

[0078] The correlation matching cost is:

[0079] ;

[0080] where represents the correlation matching cost between pixel in the -th layer of the pyramid feature map and , represents any pixel in the -th layer of the pyramid feature map corresponding to the current frame, l represents any pixel in the -th layer of the pyramid feature map corresponding to the next frame, represents any pixel in the -th layer of the pyramid feature map corresponding to the next frame, represents the -th layer of the pyramid feature map corresponding to the next frame, N is the vector length of , and represents the warping result of the -th layer of the pyramid feature map corresponding to the current frame, and T represents the transpose. l

[0081] When calculating the correlation matching cost, the pixel search range usually needs to be considered. The larger the pixel search range, the richer the correlation matching cost stored in the matching cost volume. However, although increasing the search range can provide more matching information, it will also increase the computational amount. Therefore, in practical applications, it is necessary to balance accuracy and computational efficiency and select a suitable search range. The pixel search range is usually determined by setting a maximum displacement parameter. For example, if the maximum displacement is set to d, the width and height of the search range will both be 2d + 1 (centered on the current pixel and expanding d pixels in all directions). All pixels within this range will be used to calculate the matching cost with the current pixel.

[0082] The matching cost volume is a four-dimensional tensor, where three dimensions respectively correspond to the width, height, and feature channel number of the current frame feature map, and the fourth dimension corresponds to the pixel search range related to the current frame pixel. Thus, the matching cost volume contains rich pixel matching information, providing favorable support for subsequent optical flow estimation.

[0083] It can be seen that the process of constructing the matching cost volume does not involve any learning and training parameters, and a balance between computational efficiency and accuracy is achieved through the pixel search range. Therefore, this process is also conducive to achieving lightweight optical flow estimation.

[0084] Step S14: Take the matching cost volume, the pyramid feature map of the current frame, and the upsampled optical flow as inputs, and estimate the optical flow of the current layer through a multi-layer CNN.

[0085] Among them, the matching cost volume, the pyramid feature map of the current frame, and the upsampled optical flow are all based on the l layer of the current frame feature pyramid. That is, the matching cost volume refers to the matching cost volume corresponding to the l layer, the pyramid feature map of the current frame refers to the feature map corresponding to the l layer of the current frame feature pyramid, and the upsampled optical flow refers to the optical flow result obtained by upsampling the optical flow output from the l +1 layer.

[0086] Among them, the multi-layer CNN can adopt the architecture of DenseNet, and enhance feature transfer and reuse through dense connections, which helps to improve the accuracy of optical flow estimation. Each convolutional layer in the multi-layer CNN may be followed by an activation function to increase the non-linear expression ability of the network. The size and stride of the convolutional kernel will also be adjusted according to the specific task and network design to adapt to feature inputs of different scales and optical flow estimation requirements.

[0087] Step S15: Use the context network to refine the estimated optical flow to obtain a more accurate optical flow estimation.

[0088] Among them, the context network consists of multiple convolutional layers with different dilation coefficients to expand the receptive field of each output unit. The context network takes the estimated optical flow and features of the second-to-last layer in the multi-layer CNN as inputs. These inputs provide preliminary optical flow estimation and related image feature information. The context network processes these inputs through a series of convolutional layers with different dilation coefficients, captures more context information while increasing the receptive field, and finally outputs a refined optical flow field. This refined optical flow field incorporates more context information while retaining the original optical flow information, thereby improving the accuracy and robustness of optical flow estimation.

[0089] Step S16: Repeatedly execute steps S12 - S15 at different layers of the pyramid feature map to gradually optimize the optical flow estimation;

[0090] Step S17: At the bottom layer of the pyramid, obtain the final optical flow map, which contains the motion vector information of each pixel point in the current frame.

[0091] It can be seen that, due to the use of the feature pyramid structure, the optical flow estimation process involved in step S1 can capture motion information at different scales, thereby improving the accuracy of optical flow estimation. Moreover, neither the warping operation nor the construction of the matching cost volume involves any learning and training parameters, reducing the size of the model. And the optical flow estimation result of the previous layer is used to guide the optical flow estimation of the next layer in each layer, achieving efficient operation. At the same time, the above optical flow estimation algorithm can be trained end-to-end, simplifying the training process and further reducing the usage difficulty. The above advantages enable the above optical flow estimation algorithm to achieve lightweight design, which is beneficial for its deployment on devices with limited resources such as drones.

[0092] Step S2: Extract the luminance information from the optical flow map and embed the luminance information into the current frame to obtain a fused frame.

[0093] In step S2, the extracting the luminance information from the optical flow map and embedding the luminance information into the current frame to obtain a fused frame includes:

[0094] Step S21: Convert the vector information at each pixel position in the optical flow map into scalar information, where the vector information represents the motion direction and motion speed of the pixel, and the scalar information is the luminance information of the pixel.

[0095] The optical flow map is mainly used to describe the motion of pixels in an image sequence rather than directly representing luminance information. However, the motion of pixels is usually associated with luminance changes because the optical flow algorithm relies on the assumption of luminance consistency to estimate pixel motion. In this embodiment, the magnitude of the motion vector (i.e., the norm of the vector) in the optical flow map is related to the degree of luminance change. For example, in regions with large luminance changes, the motion vectors are longer. Therefore, in step S21, the vector information at each pixel position in the optical flow map is converted into scalar information to obtain the luminance information of the optical flow map.

[0096] Step S22: Normalize the luminance information of the optical flow map to convert it into a luminance value range suitable for embedding into the current frame.

[0097] Step S23: Embed the luminance information of the optical flow map into the current frame at a certain pixel interval in a sparse embedding manner to obtain the fused frame.

[0098] When directly overlaying the luminance information of the optical flow map onto the current frame, it usually causes image information loss or confusion. Therefore, in this embodiment, a sparse embedding strategy is adopted. That is, instead of embedding each pixel of the optical flow map into the RGB image, a pixel interval is set. For example, a three-pixel interval can be selected, and one pixel is chosen from every three pixels to embed the luminance information of the optical flow map. For each selected pixel, the original RGB value and the optical flow luminance value of this pixel are calculated by a weighted method, and the obtained weighted result is used to replace the RGB value of the selected pixel.

[0099] In this way, a new image that combines the optical flow information and the RGB visual information is generated, that is, the fused frame. While retaining the visual features of the RGB image, the fused frame introduces the motion information in the optical flow map, thereby enhancing the accuracy of depth estimation.

[0100] Step S3: Process the fused frame using a depth estimation model to obtain the depth map corresponding to the fused frame; the depth estimation model includes a generator and a discriminator. The generator includes an encoder, a decoder, skip connections established between the encoder and the decoder, and an output layer. The discriminator includes multiple convolutional layers and a fully connected layer.

[0101] Among them, the encoder includes multiple depthwise separable convolutional layers. After each depthwise separable convolutional layer, batch normalization and an activation function are followed. The feature maps of the fused frame are gradually extracted through the multiple depthwise separable convolutional layers.

[0102] Skip connections are established between the encoder and the decoder to transfer the feature maps extracted by each layer of the encoder to the corresponding layers of the decoder.

[0103] The decoder includes multiple depthwise separable convolutional layers, and gradually fuses the lower-layer feature maps and the feature maps transferred through the skip connections to obtain a depth map with the same size as the fused frame.

[0104] Among them, the generator can be implemented based on the known U-Net network structure. To achieve lightweight processing, the standard convolutional layers in the original U-Net network structure are replaced with depthwise separable convolutional layers, reducing the computational complexity and the number of parameters of the model. At the same time, the encoder-decoder structure and skip connections in the U-net network structure are retained, thereby obtaining a lightweight depth estimation model. The lightweight depth estimation model not only maintains high depth estimation accuracy but also reduces the model complexity, making it more suitable for running on resource-constrained UAV platforms.

[0105] The loss function of the depth estimation model is:

[0106] ;

[0107] Among them, represents optimizing the generator parameters to minimize the loss and optimizing the discriminator parameters to maximize the loss; G represents the generator, and D represents the discriminator; represents the loss of the generative adversarial network; represents the weight factor; represents the L1 loss;

[0108] The loss of the generative adversarial network is obtained by the following formula:

[0109] ;

[0110] Among them, i represents the input image, gt represents the real image, represents the generated image obtained by the generator, n represents the random noise vector, represents the probability that the discriminator judges the real image as real, represents the probability that the discriminator misjudges the generated image as real, and E represents the expected value;

[0111] The L1 loss is obtained by the following formula:

[0112] ;

[0113] Among them, y represents the pixel value of the real image gt, represents the generated image of the corresponding pixel value; represents the L1 norm.

[0114] Step S4: Based on the depth map, the drone makes an obstacle avoidance decision using the obstacle avoidance waypoint decision model.

[0115] Among them, the obstacle avoidance waypoint decision model is implemented based on the deep Q network. The obstacle avoidance waypoint decision model includes two neural network modules: a value evaluation module and an advantage estimation module. The value evaluation module is used to estimate the value of taking each action in a given state, and the target network module is used to estimate the advantage of taking each action relative to the average action in a given state. The above two modules work together to achieve effective learning and optimization of the drone obstacle avoidance decision-making task.

[0116] In this step, it is first necessary to extract the depth map features and encode them as the state inputs of the obstacle avoidance waypoint decision model. These states include the three-dimensional information of the current position of the UAV and its surrounding environment. The above states are input into the obstacle avoidance waypoint decision model, and the Q-values of each possible action in the action space of the UAV are calculated using a neural network. The ε-greedy strategy or the softmax strategy is used to select the optimal action, that is, the action with the highest Q-value. In the exploration phase, the algorithm has a certain probability of randomly selecting actions to explore unknown states.

[0117] The UAV executes the selected action, thereby changing its position in the environment. At this time, according to the flight state of the UAV and the obstacle avoidance effect, the corresponding reward value is calculated based on a pre-designed reward function. The design of the reward function aims to guide the UAV to avoid obstacles and reach the destination as soon as possible.

[0118] Subsequently, the state transition process of the UAV, that is, the current state, action, reward, and next state, is stored in the experience replay buffer. During the training process of the obstacle avoidance waypoint decision model, a batch of experience data is randomly sampled from the experience replay buffer for the training of the obstacle avoidance waypoint decision model. During the training process, using the adopted experience data, the parameters of the obstacle avoidance waypoint decision model are optimized through the backpropagation algorithm and the gradient descent method to reduce the error between the predicted Q-value and the target Q-value. In addition, the parameters of the value evaluation module need to be copied to the advantage estimation module regularly to achieve a slow update of the advantage estimation module, which helps to maintain the stability of the training process.

[0119] The above process is continuously executed in a loop. As the training progresses, the obstacle avoidance waypoint decision model gradually learns how to make optimal obstacle avoidance decisions based on the depth map information.

[0120] The reward function plays a crucial role in the obstacle avoidance waypoint decision model. It determines the behavioral tendency of the UAV during the obstacle avoidance process. By quantifying the behavioral performance of the UAV during the obstacle avoidance process, it provides a learning signal for the UAV to guide it to optimize the obstacle avoidance strategy. In order to be able to guide the UAV to achieve autonomous obstacle avoidance in a complex environment, the following reward function is designed in this embodiment as , specifically:

[0121] ;

[0122] where is the speed reward function, is the depth reward function, is the collision reward function.

[0123] The purpose of the speed reward function is to encourage the UAV to fly at a relatively fast speed to reach the target as soon as possible while avoiding unnecessary speed fluctuations.

[0124] ;

[0125] wherein, represents the linear velocity of the UAV, represents the angular velocity of the UAV, represents the control instruction period of the UAV, is an adjustment constant for adjusting the value range of the speed reward function;

[0126] The depth reward function :

[0127] ;

[0128] ;

[0129] wherein, represents the absolute distance of the UAV to the nearest obstacle, is a collision threshold for distinguishing the safe area and the potential collision area, represents the relative distance of the UAV to the nearest obstacle, represents the change amount of the relative distance, which is used to encourage the UAV to stay away from the obstacle, represents an offset constant for adjusting the sensitivity value of the depth reward, represents the sign function;

[0130] In the case of a collision of the UAV, the collision reward function . Its purpose is to immediately give a large negative reward (-1) to punish such bad behavior when the UAV collides with an obstacle, which helps the UAV learn collision avoidance strategies during the training process.

[0131] Another embodiment of the present invention discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods are implemented.

[0132] The above introduces the preferred embodiments of the present invention, aiming to make the spirit of the present invention clearer and easier to understand, rather than to limit the present invention. Any modifications, substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope defined by the appended claims of the present invention.

Claims

1. A drone obstacle avoidance method based on image processing, characterized in that: The method comprises the following steps: S1: The drone acquires a monocular RGB image sequence, the monocular RGB image sequence including a current frame and a next frame, performs multi-scale processing on the current frame and the next frame respectively to obtain corresponding pyramid feature maps, and performs optical flow estimation on the pyramid feature maps to obtain an optical flow map; S2: extracting brightness information from the optical flow map and embedding the brightness information into the current frame to obtain a fused frame; S3: Processing the fused frame using a depth estimation model to obtain a depth map corresponding to the fused frame; the depth estimation model includes a generator and a discriminator, the generator includes an encoder, a decoder, a skip connection established between the encoder and the decoder, and an output layer, and the discriminator includes multiple convolutional layers and a fully connected layer; S4: Based on the depth map, the UAV uses an obstacle avoidance waypoint decision model to make an obstacle avoidance decision; S5: The UAV performs obstacle avoidance flight according to the obstacle avoidance decision result; In step S2, extracting brightness information from the optical flow map and embedding the brightness information into the current frame to obtain a fused frame includes: Step S21: converting the vector information of each pixel position in the optical flow map into scalar information, wherein the vector information represents the movement direction and movement speed of the pixel, and the scalar information is the brightness information of the pixel; Step S22: normalizing the brightness information of the optical flow map to convert it into a brightness value range suitable for embedding into the current frame; Step S23: using a sparse embedding method, embedding the brightness information of the optical flow map into the current frame at a certain pixel interval, thereby obtaining the fused frame.

2. The image processing-based obstacle avoidance method for a UAV according to claim 1, wherein: In step S1, multi-scale processing is performed on the current frame and the next frame respectively to obtain corresponding pyramid feature maps, and optical flow estimation is performed on the pyramid feature maps to obtain optical flow maps, including: Step S11, using a convolution filter to process the current frame I A and the next frame I B Perform multiple image downsampling to obtain the corresponding L-layer pyramid feature maps; the first l The layer passes the convolution filter to the L-layer pyramid feature map l -1 layer downsampling is performed, 1≤ l ≤L; the bottom layer of the L-layer pyramid feature map is the current frame I A or the next frame I B ; Step S12, in the l layer, using l +1 layer upsamples the optical flow by a factor of 2 to convert the current frame I A The feature of the next frame I B Distort so that the current frame I A The corresponding pyramid feature map and the next frame I B The corresponding pyramid feature maps are aligned; Step S13, for the distorted current frame I A For any pixel in the corresponding pyramid feature map, calculate its sum with the next frame I B The correlation matching cost between corresponding pixels in the corresponding pyramid feature map is used to construct the matching cost volume. Step S14, taking the matching cost volume, the pyramid feature map of the current frame and the upsampled optical flow as input, and estimating the optical flow of the current layer through a multi-layer CNN; Step S15, using the context network to refine the estimated optical flow to obtain a more accurate optical flow estimation; Step S16, repeatedly performing steps S12 to S15 at different layers of the pyramid feature map to gradually optimize the optical flow estimation; Step S17: obtaining a final optical flow map at the bottom layer of the pyramid, wherein the optical flow map includes motion vector information of each pixel in the current frame.

3. The image processing-based obstacle avoidance method for unmanned aerial vehicles according to claim 2, wherein: In step S12, the l layer, using l +1 layer upsamples the optical flow by a factor of 2 to convert the current frame I A The feature of the next frame I B Distort so that the current frame I A The corresponding pyramid feature map and the next frame I B The corresponding pyramid feature map alignment is as follows: ; Where x is the pixel index, up2(w l+1 ) indicates the distortion process from l +1 layer upsamples the optical flow by 2 times, at the top layer of the pyramid feature map, up2(w l+1 ) is set to 0; c l A Indicates the current frame I A The corresponding l-th feature layer of the L-layer pyramid feature map; c l w Indicates c l A The result of distortion.

4. The image processing-based obstacle avoidance method for unmanned aerial vehicles according to claim 3, wherein: In step S13, the correlation matching cost is: ; in, Indicates the l Pixels in the layer pyramid feature map and The correlation matching cost between Indicates the current frame I A The corresponding pyramid feature map l Any pixel of the layer, Indicates the next frame I B The corresponding pyramid feature map l Any pixel of the layer, Indicates the next frame I B The corresponding pyramid feature map l Layer, N is The length of the vector, Indicates the current frame I A The corresponding pyramid feature map l The warped result of the layer, T stands for transpose.

5. The method for avoiding obstacles in a UAV based on image processing according to claim 1, wherein: The encoder comprises a plurality of depth-wise separable convolutional layers, each of the depth-wise separable convolutional layers is followed by batch normalization and an activation function, and feature maps of the fused frame are gradually extracted through the plurality of depth-wise separable convolutional layers; Establishing a skip connection between the encoder and the decoder to transfer the feature maps extracted by each layer of the encoder to the corresponding layer of the decoder; The decoder includes multiple depth-wise separable convolutional layers, which gradually fuse the lower-layer feature maps with the feature maps transmitted through the jump connection to obtain a depth map with the same size as the fused frame.

6. The method for avoiding obstacles in a UAV based on image processing according to claim 5, characterized in that: The loss function of the depth estimation model is: ; in, Denotes that the generator parameters are optimized to minimize the loss, and the discriminator parameters are optimized to maximize the loss; G denotes the generator, and D denotes the discriminator; represents the generative adversarial network loss; represents the weight factor; represents L1 loss; The generative adversarial network loss It is obtained by the following formula: ; Among them, i represents the input image, gt represents the real image, represents the generated image obtained by the generator, n represents the random noise vector, represents the probability that the discriminator judges the real image as real, represents the probability that the discriminator misjudges the generated image as real, and E represents the expected value; The L1 loss It is obtained by the following formula: ; Wherein, y represents the pixel value of the real image gt, Represents the generated image The corresponding pixel value of represents the L1 norm.

7. The method for avoiding obstacles in a UAV based on image processing according to claim 1, wherein: The obstacle avoidance waypoint decision model is implemented based on the deep Q network, and the reward function of the obstacle avoidance waypoint decision model is , specifically: ; in, is the speed reward function, is the depth reward function, is the collision reward function.

8. The image processing-based obstacle avoidance method for unmanned aerial vehicles according to claim 7, characterized in that: The speed reward function : ; in, represents the linear velocity of the drone, represents the angular velocity of the drone, represents the control instruction cycle of the UAV, is a regulating constant used to adjust the value range of the speed reward function; The deep reward function : ; ; in, Indicates the absolute distance from the drone to the nearest obstacle, is the collision threshold, used to distinguish between the safe area and the potential collision area. Indicates the relative distance from the drone to the nearest obstacle, Indicates the change in the relative distance, used to encourage the drone to stay away from obstacles. Represents the bias constant, which is used to adjust the sensitivity of the depth reward. represents a symbolic function; In the event of a collision between the drones, the collision reward function .

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Unsupervised monocular depth estimation method based on optical flow mask

    CN115187638A

  • Driver behavior strategy generation method, device and equipment and readable storage medium

    CN118627276A