Obstacle avoidance method and apparatus for unmanned aerial vehicle, electronic device, and storage medium

By integrating millimeter-wave radar and cameras onto drones, and combining attention mechanisms and deep learning networks for feature-level data fusion, the problem of inaccurate obstacle recognition by drones in low-visibility environments has been solved, improving the flight safety and adaptability of drones.

WO2025260493A1PCT designated stage Publication Date: 2025-12-26TIANMUSHAN LABORATORY

Patent Information

Application Number
PCT/CN2024/113610
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2024-08-21
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing drone obstacle avoidance technology struggles to accurately identify obstacles in complex environments with low visibility, resulting in insufficient flight safety.

Method used

By combining millimeter-wave radar and high-speed cameras, environmental images and point cloud data are acquired. Then, attention mechanisms and deep learning networks are used to perform feature-level data fusion to achieve precise obstacle localization and avoidance.

Benefits of technology

It improves the accuracy of obstacle detection and flight safety of UAVs in complex environments, and enhances the adaptability and robustness of UAVs in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024113610_26122025_PF_FP_ABST
    Figure CN2024113610_26122025_PF_FP_ABST
Patent Text Reader

Abstract

An obstacle avoidance method and apparatus for an unmanned aerial vehicle, an electronic device, and a storage medium. During flight of a target unmanned aerial vehicle, an environment image and point cloud data that correspond to the surrounding environment of the target unmanned aerial vehicle are acquired (S101); an obstacle in the environment image is obtained, and when an obstacle involved in the point cloud data matches the obtained obstacle, a fused image is determined on the basis of the environment image and the point cloud data (S102); and position information corresponding to the obstacle is determined on the basis of the fused image, and the position information is transmitted to the target unmanned aerial vehicle, enabling the target unmanned aerial vehicle to avoid the obstacle (S103). In the obstacle avoidance method for an unmanned aerial vehicle, the surroundings of the unmanned aerial vehicle can be monitored in real time by means of the environment image and the point cloud data, so as to achieve comprehensive sensing and accurate positioning of the obstacle in the presence of the obstacle, thereby accurately avoiding the obstacle, and improving the travelling safety of the unmanned aerial vehicle during travelling.
Need to check novelty before this filing date? Find Prior Art

Description

Unmanned aerial vehicle (UAV) obstacle avoidance methods, devices, electronic equipment and storage media

[0001] This disclosure is based on Chinese Patent Application No. 202410810516.6, filed on June 20, 2024, entitled "Unmanned Aerial Vehicle Obstacle Avoidance Method, Apparatus, Electronic Device and Storage Medium", and claims priority to that Chinese Patent Application, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the technical field of drone driving, and more particularly to a drone obstacle avoidance method, device, electronic device, and storage medium. Background Technology

[0003] Unmanned aerial vehicles (UAVs), also known as drones, are unmanned aircraft controlled by wireless remote control equipment and their own program control devices.

[0004] In related technologies, cameras can be used to collect images of the drone's surroundings during flight, allowing the drone to avoid obstacles if they are present in the images. However, when using cameras to collect images of the drone's surroundings, in complex traffic environments with low visibility, objective factors such as lighting, occlusion, or shadows may result in poor image clarity, leading to poor accuracy in obstacle recognition.

[0005] Summary of the Invention

[0006] In view of this, the present disclosure provides an obstacle avoidance method, apparatus, electronic device, and storage medium for unmanned aerial vehicles (UAVs) to solve the problems existing in the related technologies.

[0007] A first aspect of this disclosure provides an obstacle avoidance method for an unmanned aerial vehicle (UAV). The method includes: acquiring environmental images and point cloud data corresponding to the surrounding environment of the UAV during flight; acquiring obstacles in the environmental images; and determining a fused image based on the environmental images and point cloud data when the obstacles involved in the point cloud data match the obstacles; determining the location information corresponding to the obstacles based on the fused image; and sending the location information to the UAV so that the UAV avoids the obstacles.

[0008] A second aspect of this disclosure provides an obstacle avoidance device for unmanned aerial vehicles (UAVs), applied to the obstacle avoidance method of the first aspect. The device includes: an acquisition module, configured to acquire environmental images and point cloud data corresponding to the surrounding environment of the target UAV during flight; a determination module, configured to acquire obstacles in the environmental images, and determine a fused image based on the environmental images and point cloud data when the obstacles involved in the point cloud data match; and a control module, configured to determine the position information corresponding to the obstacles based on the fused image, and send the position information to the target UAV, so that the target UAV avoids the obstacles.

[0009] A third aspect of this disclosure provides an electronic device including at least one processor; a memory for storing at least one processor-executable instruction; wherein the at least one processor is used to execute the instruction to implement the steps of the above-described drone obstacle avoidance method.

[0010] A fourth aspect of this disclosure provides a computer-readable storage medium that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the steps of the above-described drone obstacle avoidance method.

[0011] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the above-described drone obstacle avoidance method.

[0012] The at least one technical solution adopted in this disclosure can achieve the following beneficial effects: During the flight of the target drone, environmental images and point cloud data corresponding to the surrounding environment are acquired; obstacles in the environmental images are acquired; when obstacles in the point cloud data match, a fused image is determined based on the environmental images and point cloud data; the location information corresponding to the obstacles is determined based on the fused image, and the location information is sent to the target drone, enabling the target drone to avoid the obstacles. Based on this, the surroundings of the drone can be monitored in real time using environmental images and point cloud data, thereby achieving comprehensive perception and accurate positioning of obstacles in the presence of obstacles, thus accurately avoiding obstacles and improving the driving safety of the drone during flight.

[0013] To make the above-described objects, features, and advantages of this disclosure more apparent and understandable, specific embodiments are described below in conjunction with the accompanying drawings. It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0015] Figure 1 is a flowchart illustrating an exemplary embodiment of the present disclosure of an obstacle avoidance method for unmanned aerial vehicles (UAVs).

[0016] Figure 2 is a flowchart illustrating an image data matching method provided in an exemplary embodiment of this disclosure;

[0017] Figure 3 is a schematic diagram of a truncated cone association method provided in an exemplary embodiment of this disclosure;

[0018] Figure 4 is a schematic diagram of a point cloud data pillaring structure provided by an exemplary embodiment of the present disclosure;

[0019] Figure 5 is a flowchart illustrating a center point detection network provided in an exemplary embodiment of this disclosure;

[0020] Figure 6 is a flowchart illustrating an obstacle localization method provided in an exemplary embodiment of this disclosure;

[0021] Figure 7 is a structural schematic diagram of an obstacle avoidance device for a drone provided in an exemplary embodiment of this disclosure;

[0022] Figure 8 is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of the present disclosure;

[0023] Figure 9 is a schematic diagram of the structure of a computer system provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0026] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0027] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0028] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0029] In today's era of rapid technological development, drone technology has become an important tool in many fields such as aviation, military, logistics, and agriculture. However, with the increasingly widespread application of drones, their flight safety issues are becoming more and more prominent, especially in complex flight environments where it is difficult to accurately avoid obstacles, meaning that flight safety cannot be guaranteed during drone flight.

[0030] Initially used primarily in the military field, drones have rapidly penetrated the civilian sector due to advancements in related technologies and reductions in cost. For example, in agriculture, drones are used for crop monitoring, pest and disease control, and yield assessment; in the construction industry, drones are used for construction progress monitoring and building inspection; and in rescue operations, drones can quickly reach hard-to-access areas for search and rescue.

[0031] In the three scenarios described above, drones may fly in complex three-dimensional spaces during missions. These complex spaces may include various obstacles such as buildings, trees, and utility poles. Therefore, drones may collide with obstacles during flight due to untimely obstacle avoidance, resulting in damage. To address this issue, existing technology provides a drone obstacle detection technology for detecting obstacles during drone flight. This allows the drone to autonomously identify these obstacles without human control, thereby reducing collision risks and ensuring flight safety.

[0032] Existing drone obstacle avoidance technologies primarily rely on basic sensing devices and algorithms, which played a crucial role in the early stages of drone development. Specifically, drones can be equipped with ultrasonic sensors. These sensors measure the distance between the drone and obstacles by emitting sound waves and receiving their echoes. Ultrasonic sensors are relatively inexpensive to produce and suitable for short-range obstacle avoidance, but they can be affected by wind speed and ambient noise in outdoor environments. Infrared sensors can also be mounted on drones. These sensors emit infrared light and detect the reflected light, commonly used for obstacle detection during low-altitude flight, especially in indoor environments. However, their detection range and angle are limited, potentially requiring multiple sensors to cover a larger area. Monocular cameras can also be used. While providing two-dimensional images with rich information, they struggle to accurately identify obstacles in complex traffic environments with low visibility due to variations in lighting, occlusion, and shadows. Additionally, model-based obstacle avoidance methods exist. These methods rely on predefined environmental models, allowing the drone to plan its path and avoid known obstacles. This approach works well in environments with relatively fixed structures.

[0033] Single-sensor devices such as cameras, lidar, or millimeter-wave radar often have limited capabilities in obstacle detection. For example, while cameras capture images rich in information, in complex traffic environments with low visibility, variations in lighting, obstructions, and shadows can make it difficult to obtain effective images, thus hindering accurate obstacle identification. Lidar, on the other hand, provides highly accurate distance and velocity measurements, but its cost is relatively high, and traditional lidar systems can be bulky and heavy, making installation difficult. Therefore, because different sensor devices have different applicable conditions, installing only a single sensor is insufficient for various environments.

[0034] To address the aforementioned issues, existing technologies employ multiple sensors working together to achieve more comprehensive and accurate obstacle detection. Extensive research has been conducted on multi-source data fusion algorithms corresponding to multiple sensors, yielding significant results. Various advanced algorithms have been successfully applied in various fields. Currently, these algorithms are primarily used in vehicle driving, and their application in the drone field urgently needs verification and promotion. For example, Tesla's Autopilot system integrates multiple sensors, including cameras, radar, and ultrasonic sensors. Data from these sensors is fused to provide a comprehensive view of the vehicle's surroundings, enabling functions such as adaptive cruise control, lane keeping, automatic lane changing, and collision prevention. In the drone field, drones like DJI's use multi-source data fusion technology, combining visual sensors, ultrasonic sensors, and GPS data to achieve automatic obstacle avoidance and precise hovering. Therefore, if a multi-source information fusion drone obstacle detection system can be designed, combining the advantages of millimeter-wave radar, cameras, and other possible sensing devices, and using data fusion technology, the accuracy of obstacle detection and the reliability of the system can be improved, enabling drones to perform tasks safely and effectively in more complex and dynamic environments, which has great application prospects.

[0035] Compared to radar or cameras alone, radar-camera fusion sensing systems leverage the advantage of radar's rapid measurement of object dynamics while utilizing the visual recognition of obstacles' physical attributes, such as size and location. In practical applications, current fusion systems typically operate at the data or decision level. Data-level fusion, also known as early fusion, involves coordinate transformation and synchronization calibration of raw data from various sensors before direct combination. This type of fusion requires high levels of system synchronization and calibration and has relatively low flexibility. Decision-level fusion, also known as late-stage fusion, occurs in the later stages of data processing, integrating the local decision results from each sensor to form a global decision. At the decision level, each sensor processes data independently, potentially leading to the loss of useful information during independent processing, thus hindering the utilization of complementary information between different sensors.

[0036] Based on this, in the current UAV obstacle detection technology, since different sensing devices have their own unique detection capabilities and coverage, multiple different types of sensing devices can be fused. For the processing of the fused data, existing technologies usually perform fusion at the data level or decision level. Such fusion methods are usually very sensitive to data spatial and temporal deviations and cannot fully fuse information at the feature level.

[0037] To address the aforementioned issues, this disclosure provides an obstacle avoidance method, apparatus, electronic device, and storage medium for unmanned aerial vehicles (UAVs). These methods can integrate data from different sensing devices and fuse the data at the feature level, thereby enabling a more comprehensive perception of the surrounding environment and improving the accuracy and reliability of obstacle detection.

[0038] The drone obstacle avoidance method provided in this disclosure can be executed by a terminal or by a chip applied to the terminal.

[0039] For example, the aforementioned terminals may include one or more of the following: mobile phones, tablets, wearable devices, in-vehicle devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, handheld computers (PDAs), and wearable devices based on augmented reality (AR) and / or virtual reality (VR) technologies. They may also include, but are not limited to, remote control devices, wearable devices, streetlights, home appliances, and other smart terminals. This disclosure does not impose specific limitations on these aspects.

[0040] Figure 1 is a flowchart illustrating an exemplary embodiment of this disclosure of an obstacle avoidance method for unmanned aerial vehicles (UAVs). As shown in Figure 1, the method specifically includes:

[0041] S101 acquires environmental images and point cloud data corresponding to the surrounding environment of the target drone during its flight.

[0042] In some embodiments, because radar detection devices have strong environmental adaptability and image acquisition devices can provide rich texture information, embodiments of this disclosure can mount at least one radar detection device and at least one image acquisition device on a drone. Compared with other sensing devices, millimeter-wave radar has unique advantages: it can remotely detect the presence of obstacles and can operate under any lighting and weather conditions, exhibiting strong environmental adaptability. Therefore, the radar detection device in embodiments of this disclosure can be a millimeter-wave radar, etc., while the image acquisition device can be a device with image acquisition capabilities, such as a high-speed camera.

[0043] In practical applications, radar detection equipment is an active sensing device that senses the surrounding environment by emitting radio waves and measures reflected waves to determine the position and velocity of obstacles. Millimeter-wave radar, operating at a frequency of 77 GHz, can measure multi-dimensional information such as the coordinates, distance, and velocity of obstacles. Image acquisition equipment, on the other hand, is a visual sensing device that can capture high-resolution images, providing rich color, texture, and shape information. By analyzing these images, the shape and approximate size of obstacles in the image can be determined. Combined with stereo vision technology, the position and distance of obstacles in three-dimensional space can be obtained.

[0044] Based on this, embodiments of this disclosure can utilize image acquisition devices to collect environmental data corresponding to the environment surrounding the target drone, generate environmental images, and determine whether obstacles exist around the target drone using these images. When obstacles are identified, their location in the real environment and their distance from the target drone's current location can be determined. Simultaneously, radar detection devices can be used to acquire point cloud data corresponding to the environment surrounding the target drone. This point cloud data can also be used to obtain information such as the location of obstacles in the real environment and their distance from the target drone's current location. Therefore, through the combined action of multiple sensing devices, accurate detection of obstacles in different scenarios can be achieved.

[0045] S102, Obtain obstacles in the environmental image, and determine the fused image based on the environmental image and the point cloud data if the obstacles involved in the point cloud data match.

[0046] In some embodiments, since different sensing devices have different applicable conditions, this disclosure uses two sensing devices to detect the surrounding environment of the target UAV and acquire image data and point cloud data corresponding to the surrounding environment. At the same time, in order to make full use of the data acquired by the sensing devices, the environmental image and point cloud data can be fused to obtain a fused image. The fused image can include all data from the environmental image and point cloud data, thereby obtaining rich data and realizing a more comprehensive analysis of obstacles and more accurate positioning.

[0047] Specifically, an attention mechanism can be used to fuse information from environmental images and point cloud data. Here, the use of the attention mechanism fusion module can further enhance the fusion network's ability to extract key information. Furthermore, by integrating data detected by different sensors through image fusion, the surrounding environment can be perceived more comprehensively, and obstacles around the target drone can be detected more comprehensively. This allows for precise obstacle localization, helping the target drone avoid obstacles during flight and ensuring its flight safety.

[0048] S103: Determine the location information of the obstacle based on the fused image, and send the location information to the target drone so that the target drone can avoid the obstacle.

[0049] In some embodiments, the present disclosure can determine obstacles in the current surrounding environment by fusing images and determine the location information corresponding to the obstacles, and send the location information of the obstacles to the target drone so that the target drone can avoid the obstacles in time, thereby ensuring the flight safety of the target drone.

[0050] Based on this, during the flight of the target drone, environmental images and point cloud data corresponding to the surrounding environment can be acquired; obstacles in the environmental images can be identified; when obstacles in the point cloud data match, a fused image can be determined based on the environmental images and point cloud data; the location information corresponding to the obstacles can be determined based on the fused image, and the location information can be sent to the target drone, enabling the target drone to avoid the obstacles. Therefore, the embodiments of this disclosure can monitor the drone's surroundings in real time using environmental images and point cloud data, thereby achieving comprehensive perception and accurate positioning of obstacles in the presence of obstacles, thus accurately avoiding obstacles and improving the drone's flight safety.

[0051] In some embodiments of this disclosure, when obstacles in the point cloud data are matched with each other, a fused image can be determined based on the environmental image and the point cloud data, including: associating obstacles in the point cloud data with each other to obtain an association result; and determining a fused image based on the association result, the environmental image, and the point cloud data.

[0052] Specifically, during the fusion process, data corresponding to the same obstacles need to be fused. Therefore, it is necessary to first match one or more obstacles in the environmental image with one or more obstacles in the point cloud data to determine the data corresponding to the same obstacle under different acquisition methods. Then, the data corresponding to the same obstacle under different acquisition methods are associated. Finally, the data is fused based on the association results, so that the data corresponding to the same obstacle can be accurately fused, ensuring the accuracy of the data.

[0053] In some embodiments of this disclosure, the location information corresponding to the obstacle can be determined based on the environmental image, and the bounding box corresponding to the obstacle can be determined based on the location information corresponding to the obstacle; the point cloud data in the bounding box can be determined as the target point cloud data; and a fused image can be determined based on the environmental image and the target point cloud data.

[0054] Specifically, this disclosure describes the processing of environmental images using a deep learning network for object detection. The detection network is a deep learning network for real-time object detection that simplifies the object detection problem by modeling the object as a single point in the image; specifically, it can be a CenterNet network. In the CenterNet network, the bounding box corresponding to the obstacle can be obtained by predicting the offset of the object's center point and the obstacle's size, thereby enabling accurate obstacle analysis.

[0055] In some embodiments, when determining the bounding box, the present disclosure embodiments can determine a center point heatmap based on an environmental image, wherein the center point heatmap is used to characterize the center point position of an obstacle; and the bounding box corresponding to the obstacle is determined based on the center point position.

[0056] Specifically, the CenterNet network can be used to process the environmental image to obtain a center point heatmap of the environmental image. Then, the center point position of the obstacle can be determined based on the center point heatmap. Finally, the bounding box of the obstacle can be determined by the center point offset corresponding to the center point position and the size of the obstacle.

[0057] Figure 2 is a flowchart illustrating an image data matching method provided in an exemplary embodiment of this disclosure. Figure 3 is a structural diagram illustrating a frustum association method provided in an exemplary embodiment of this disclosure. As shown in Figures 2 and 3, the frustum association method can be used to associate point cloud data 202 corresponding to the same obstacle with data in the environmental image. Specifically, for the first obstacle in the environmental image, a frustum 301 of the obstacle can be created based on the bounding box 201 corresponding to the obstacle, the size information of the obstacle, and the estimated depth of the bounding box 201. The environmental image can be a depth image, also known as a distance image, which refers to an image where the distance from the image acquisition device to each point in the scene is used as pixel values.

[0058] As shown in Figure 3, in order to further improve the robustness of image depth estimation, the size of the truncated cone 301 can be controlled by the parameter δ. For example, by using the parameter δ to enlarge the truncated cone, the number of point cloud data 202 within the truncated cone 301 can be increased.

[0059] Figure 4 is a schematic diagram of a point cloud data pillaring structure provided by an exemplary embodiment of this disclosure. As shown in Figure 4, in order to further improve the problem of inaccurate height information of radar detection equipment, this embodiment of the disclosure also introduces a pillar expansion method to preprocess the point cloud data 401. Specifically, each point cloud data 401 can be expanded into a point cloud pillar 402 of fixed size. The point cloud pillar expansion of the point cloud data can better represent the obstacles detected by the radar detection equipment, and can also better correlate with obstacles in the environmental image. Based on this, for an obstacle A detected in the environmental image, if all or part of the point cloud pillar corresponding to the point cloud data detected by the radar detection equipment is within the truncated cone 301 corresponding to obstacle A, then it is determined that the obstacle matches obstacle A, that is, the obstacle and obstacle A are the same obstacle.

[0060] In some embodiments, the present disclosure may extract features from an environmental image to obtain an environmental feature image; use the data corresponding to the environmental feature image to perform augmentation processing on point cloud data, and determine a point cloud feature image based on the augmented point cloud data; and fuse the point cloud feature image with the environmental feature image to obtain a fused image.

[0061] Specifically, in this embodiment of the present disclosure, when extracting features from environmental images, a deep learning network for object detection can be used. This detection network is a deep learning network for real-time object detection that simplifies the object detection problem by modeling obstacles as individual points in the environmental image; specifically, the CenterNet algorithm can be used. In the CenterNet algorithm, each obstacle can be represented as the center point of its bounding box, rather than the traditional bounding box representation, which greatly simplifies the detection process. Then, based on the center point, other target attributes of the obstacle can be determined. Here, other target attributes can specifically include the obstacle's size, 3D position, orientation, and even pose. CenterNet is a single-stage object detection algorithm, an end-to-end algorithm that is faster and more accurate than anchor-box-based object detection algorithms.

[0062] Figure 5 is a flowchart illustrating a center point detection network provided in an exemplary embodiment of this disclosure. As shown in Figure 5, an environmental image in standard RGB image format is used as input. The dimensions of an RGB image are typically W×H×C, where W represents the width of the RGB image, H represents the height of the RGB image, and C represents the number of image channels. The backbone network adopts an encoder-decoder network architecture based on the ResNet network structure. The main goal of the backbone network is to generate a center point heatmap. Here, R represents the downsampling step size, C represents the number of image channels, and the generated obstacle size and center point offset are also included. The center point heatmap represents the center point position of each obstacle in the RGB image. The obstacle size output is used to regress the specific dimensions of the obstacle (such as height and width). The center point offset represents the systematic error, used to further improve the accuracy of obstacle center point localization. Furthermore, when determining the obstacle position, the predicted center point coordinates can be added to the corresponding center point offset to obtain a more accurate obstacle center point location. The ResNet network structure has wide applications in various fields such as image classification, object detection, and semantic segmentation. For example, in image classification tasks, ResNet can improve classification accuracy by continuously increasing network depth; in object detection tasks, ResNet can be used as a feature extraction network to extract image features and input them into the object detector for object detection; in semantic segmentation tasks, ResNet can be used as an encoder to extract image features and pass the features to the decoder for pixel-level semantic segmentation.

[0063] Specifically, in this embodiment of the present disclosure, when extracting features from point cloud data, an encoding network, such as a ResNet network, can be used. After extracting features from the point cloud data using a ResNet network, the extracted point cloud data can be augmented using data corresponding to the environmental feature image to obtain a point cloud feature image; that is, data augmentation of the point cloud data is performed using data corresponding to the environmental feature image. Next, a multi-head attention mechanism and a cross-attention mechanism can be used to fuse the point cloud feature image and the environmental feature image to obtain a fused image.

[0064] In practical applications, data corresponding to environmental feature images and point cloud feature images can be fused within a set of local image windows. It should be understood that the local image window can adaptively adjust based on the distance between the point cloud or pixels and the obstacle, thus enabling accurate obstacle identification when the drone approaches. The adaptive size of the local image window τ(d) is similar to the inverse function of spatially increasing discretization, as shown below:

[0065] Where d represents the distance between the radar detection device and the obstacle, α represents the lower limit of the radar detection device's detection range, β represents the upper limit of the radar detection device's detection range, and W represents the scaling factor.

[0066] Specifically, for example, W = 3.5, α = 2, β = 55 can be set, and the size of the local image window τ(d) can be adaptively adjusted through this function. More specifically, the characteristics of a radar detection device can be given. Features extracted adaptively from point cloud data w represents the size of the local image window, C represents the number of image channels, i represents the image features, k represents the k-th point cloud or pixel, and r represents the identifier of the radar detection device.

[0067] In some embodiments, the present disclosure can also acquire the detection depth and detection speed corresponding to the radar detection device, and determine complementary features based on the detection depth and detection speed, wherein the complementary features are used to characterize the supplementary features of the detection depth and detection speed to the fused image, and the radar detection device is used to acquire point cloud data; and determine the location information of the obstacle based on the complementary features and the fused image.

[0068] Specifically, in order to improve the accuracy of image recognition, after fusing point cloud data and environmental images, this embodiment of the present disclosure can also create complementary features by using the detection depth and detection speed corresponding to the radar detection device. Then, by using the complementary features to supplement the data in the fused image, the content and semantic information of the image can be described more comprehensively, thereby further improving the positioning accuracy of obstacles.

[0069] Figure 6 is a flowchart illustrating an obstacle localization method provided in an exemplary embodiment of this disclosure. As shown in Figure 6, the method specifically includes:

[0070] The S601 uses image acquisition equipment and radar detection equipment to acquire environmental images and point cloud data around the target UAV.

[0071] In some embodiments, the radar detection device is a millimeter-wave radar, and the image acquisition device can be a camera.

[0072] S602, associate the obstacles corresponding to the point cloud data with the obstacles in the environmental image, and fuse the environmental image and point cloud data according to the association result to obtain a fused image.

[0073] In some embodiments, deep learning networks for object detection can be used to process environmental images. These detection networks are deep learning networks for real-time object detection that simplify the object detection problem by modeling obstacles as individual points in the environmental image. Taking CenterNet as an example, each obstacle can be represented as a center point instead of a traditional bounding box, thus simplifying the obstacle detection process. Then, based on the center point, other target attributes of the obstacle, such as size, 3D position, orientation, and even pose, can be determined.

[0074] As shown in Figure 5, a standard RGB image format environment image is used as input. The dimensions of an RGB image are typically W×H×C, where W represents the width of the RGB image, H represents the height of the RGB image, and C represents the number of images. The backbone network adopts a ResNet-based encoder-decoder network architecture. The main goal of the backbone grid is to generate a center point heatmap. Where R represents the downsampling step size, C represents the number of images, and the generated obstacle size and center point offset are also specified. The center point heatmap represents the center point position of each obstacle in the RGB image. The obstacle size output is used to regress the specific dimensions of the obstacle (such as height and width). The center point offset is a systematic error used to further improve the accuracy of obstacle center point localization. Furthermore, when determining the obstacle position, the predicted center point coordinates can be added to the corresponding center point offset to obtain a more accurate obstacle center point position.

[0075] The key to the fusion mechanism is the precise association between point cloud data and obstacles in the environmental image. The CenterNet center point target detection network generates a center point heatmap for each obstacle in the image. Peaks in the center point heatmap represent possible center points of the object, and the image features corresponding to these center points can be used to estimate other attributes of the obstacle. To fully utilize point cloud data in this context, it is necessary to map the point cloud data corresponding to obstacles detected by radar detection equipment onto the corresponding obstacles in the environmental image for association. Then, based on the association results, the point cloud data corresponding to the same obstacle can be precisely fused with the environmental image.

[0076] As shown in Figures 2 and 3, the frustum association method can be used to associate the point cloud data 202 corresponding to the same obstacle with the data in the environmental image. Specifically, for the first obstacle in the environmental image, a frustum 301 of the obstacle can be created based on the bounding box 201 corresponding to the obstacle, the size information of the obstacle, and the estimated depth of the bounding box 201. The environmental image can be a depth image, also known as a distance image, which refers to an image where the distance from the image acquisition device to each point in the scene is used as the pixel value.

[0077] As shown in Figure 3, in order to further improve the robustness of image depth estimation, the size of the truncated cone 301 can be controlled by the parameter δ. For example, by using the parameter δ to enlarge the truncated cone, the number of point cloud data 202 within the truncated cone 301 can be increased.

[0078] As shown in Figure 4, to further improve the accuracy of height information in radar detection equipment, this embodiment introduces a pillar expansion method to preprocess the point cloud data 401. Specifically, each point cloud data 401 can be expanded into a fixed-size point cloud pillar 402. The point cloud pillar expansion can better represent obstacles detected by the radar detection equipment and can also better correlate with obstacles in the environmental image. Based on this, for obstacle A detected in the environmental image, if all or part of the point cloud pillar corresponding to the point cloud data detected by the radar detection equipment is within the truncated cone 301 corresponding to obstacle A, then it is determined that the obstacle matches obstacle A, that is, the obstacle and obstacle A are the same obstacle.

[0079] S603 determines complementary features based on the detection depth and detection speed of the radar detection equipment.

[0080] In some embodiments, after associating point cloud data with data from the environmental image, it is also necessary to use the detection depth and detection speed of the radar detection device to create complementary features for the fused image, further improving the obstacle localization accuracy. Here, an attention mechanism is used to fuse the environmental image and point cloud data. The use of the attention mechanism fusion module can further enhance the fusion network's ability to extract key information.

[0081] Specifically, in this embodiment of the present disclosure, when extracting features from point cloud data, an encoding network, such as a ResNet network, can be used. After extracting features from the point cloud data using a ResNet network, the extracted point cloud data can be augmented using data corresponding to the environmental feature image to obtain a point cloud feature image; that is, data augmentation of the point cloud data is performed using data corresponding to the environmental feature image. Next, a multi-head attention mechanism and a cross-attention mechanism can be used to fuse the point cloud feature image and the environmental feature image to obtain a fused image.

[0082] In practical applications, data corresponding to environmental feature images and point cloud feature images can be fused within a set of local image windows. It should be understood that the local image window can adaptively adjust based on the distance between the point cloud or pixels and the obstacle, thus enabling accurate obstacle identification when the drone approaches. The adaptive size of the local image window τ(d) is similar to the inverse function of spatially increasing discretization, as shown below:

[0083] Where d represents the distance between the radar detection device and the obstacle, α represents the lower limit of the radar detection device's detection range, β represents the upper limit of the radar detection device's detection range, and W represents the scaling factor.

[0084] Specifically, for example, W = 3.5, α = 2, β = 55 can be set, and the size of the local image window τ(d) can be adaptively adjusted through this function. More specifically, the characteristics of a radar detection device can be given. Features extracted adaptively from point cloud data w represents the size of the local image window, C represents the number of image channels, i represents the image features, k represents the k-th point cloud or pixel, and r represents the identifier of the radar detection device.

[0085] This embodiment also uses Layer Normalization (LN) during the fusion process. LN reduces internal covariate bias by normalizing the input of each layer in the neural network, ensuring that the distribution of input data for each layer remains relatively stable. This helps to accelerate convergence and improve the model's generalization ability. The calculation process is as follows:

[0086] Where x is the input feature, μ is the feature mean, σ is the feature standard deviation, and γ and β are learned parameters used to scale and translate the normalized data to maintain the model's expressive power.

[0087] S604, determine the location information of the obstacle based on complementary features and the fused image.

[0088] In some embodiments, this disclosure may also use a regression prediction network to perform secondary regression prediction, finely adjusting the positional offset of the obstacle, thereby more accurately locating the obstacle.

[0089] Based on this, this disclosure proposes a drone obstacle avoidance method that integrates radar and visual sensing devices, achieving comprehensive obstacle perception. This not only improves the accuracy of obstacle detection by the drone but also enhances its adaptability and robustness in complex environments. Furthermore, compared to traditional data-level or decision-level fusion methods, this disclosure employs a feature-level fusion strategy, deeply integrating data from different sensing devices at the feature level, and more effectively utilizing the complementary information from different sensors. This strategy can more effectively utilize the key information from each sensing device, reduce data redundancy, and improve the flexibility and accuracy of drone control. Simultaneously, the use of a cross-attention mechanism enhances the ability to extract key information, enabling point cloud data to be more accurately fused with image data, thus optimizing the accuracy of obstacle localization.

[0090] The foregoing mainly describes the solutions provided by the embodiments of this disclosure. It is understood that, in order to achieve the above functions, the electronic device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0091] By dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a drone obstacle avoidance device, which can be a server or a chip applied to a server. Figure 7 is a schematic structural diagram of a drone obstacle avoidance device provided in an exemplary embodiment of this disclosure. As shown in Figure 7, the drone obstacle avoidance device 700 includes:

[0092] The acquisition module 701 is used to acquire environmental images and point cloud data corresponding to the environment around the target drone during the flight of the target drone;

[0093] The determination module 702 is used to acquire obstacles in the environmental image, and when the obstacles involved in the point cloud data match the obstacles, determine a fused image based on the environmental image and the point cloud data;

[0094] The control module 703 is used to determine the location information corresponding to the obstacle based on the fused image, and send the location information to the target drone so that the target drone avoids the obstacle.

[0095] In an alternative embodiment, the determining module 702 is further configured to associate the obstacles involved in the point cloud data with the obstacles to obtain an association result; and to determine a fused image based on the association result, the environmental image, and the point cloud data.

[0096] In an optional manner, the determining module 702 is further configured to extract features from the environmental image to obtain an environmental feature image; augment the point cloud data using the data corresponding to the environmental feature image, and determine a point cloud feature image based on the augmented point cloud data; and fuse the point cloud feature image with the environmental feature image to obtain a fused image.

[0097] In an optional embodiment, the determining module 702 is further configured to acquire the detection depth and detection speed corresponding to the radar detection device, and determine complementary features based on the detection depth and the detection speed, wherein the complementary features are used to characterize the supplementary features of the detection depth and the detection speed to the fused image, the radar detection device is used to acquire point cloud data; and determine the location information of the obstacle based on the complementary features and the fused image.

[0098] In an optional embodiment, the determining module 702 is further configured to determine the location information corresponding to the obstacle based on the environmental image, and determine the bounding box corresponding to the obstacle based on the location information corresponding to the obstacle; determine the point cloud data in the bounding box as target point cloud data; and determine a fused image based on the environmental image and the target point cloud data.

[0099] In an alternative embodiment, the determining module 702 is further configured to determine a center point heatmap based on the environmental image, wherein the center point heatmap is used to characterize the center point position of the obstacle; and to determine the bounding box corresponding to the obstacle based on the center point position.

[0100] This disclosure also provides an electronic device, including: at least one processor; a memory for storing at least one processor-executable instruction; wherein the at least one processor is used to execute the instruction to implement the steps of the method disclosed in this disclosure.

[0101] Figure 8 is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of the present disclosure. As shown in Figure 8, the electronic device 800 includes at least one processor 801 and a memory 802 coupled to the processor 801. The processor 801 can execute the corresponding steps in the methods disclosed in the embodiments of the present disclosure.

[0102] The processor 801 described above can also be referred to as a Central Processing Unit (CPU), which can be an integrated circuit chip with signal processing capabilities. Each step in the method disclosed in this embodiment can be implemented by the integrated logic circuitry in the processor 801 or by software instructions. The processor 801 can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this embodiment can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in the memory 802, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor 801 reads information from the memory 802 and, in conjunction with its hardware, completes the steps of the method described above.

[0103] Furthermore, various operations / processes according to this disclosure, when implemented through software and / or firmware, can be installed from a storage medium or network onto a computer system with a dedicated hardware architecture, such as the computer system 900 shown in FIG. 9. When various programs are installed, this computer system is capable of performing various functions, including those described above. FIG. 9 is a schematic diagram of the structure of a computer system provided in an exemplary embodiment of this disclosure.

[0104] Computer system 900 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0105] As shown in Figure 9, the computer system 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. The RAM 903 can also store various programs and data required for the operation of the computer system 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0106] Multiple components in the computer system 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device capable of inputting information into the computer system 900. The input unit 906 can receive input numerical or character information and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 908 may include, but is not limited to, a hard disk and an optical disk. The communication unit 909 allows the computer system 900 to exchange information / data with other devices via a network such as the Internet, and may include, but is not limited to, a modem, network card, infrared communication device, wireless communication transceiver, and / or chipset, such as Bluetooth™ device, WiFi device, WiMax device, cellular communication device, and / or the like.

[0107] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the methods disclosed in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 902 and / or communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the methods disclosed in this disclosure by any other suitable means (e.g., by means of firmware).

[0108] This disclosure also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the methods disclosed in this disclosure.

[0109] The computer-readable storage medium in this disclosure can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The aforementioned computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specifically, the aforementioned computer-readable storage medium may include electrical connections based on one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0110] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0111] This disclosure also provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements the methods disclosed in the embodiments of this disclosure.

[0112] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer.

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.

[0115] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0116] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0117] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A method for obstacle avoidance by unmanned aerial vehicles (UAVs), characterized in that, The method includes: During the flight of the target drone, environmental images and point cloud data corresponding to the environment around the target drone are acquired; Obstacles in the environmental image are acquired, and if the obstacles involved in the point cloud data match the obstacles, a fused image is determined based on the environmental image and the point cloud data; The location information corresponding to the obstacle is determined based on the fused image, and the location information is sent to the target drone so that the target drone can avoid the obstacle.

2. The method according to claim 1, characterized in that, When the obstacle involved in the point cloud data matches the obstacle, determining the fused image based on the environmental image and the point cloud data includes: The obstacles involved in the point cloud data are associated with the obstacles to obtain the association results; The fused image is determined based on the association results, the environmental image, and the point cloud data.

3. The method according to claim 1, characterized in that, The method further includes: Feature extraction is performed on the environmental image to obtain an environmental feature image; The point cloud data is augmented using the data corresponding to the environmental feature image, and the point cloud feature image is determined based on the augmented point cloud data. The point cloud feature image is fused with the environmental feature image to obtain a fused image.

4. The method according to claim 1, characterized in that, The method further includes: The detection depth and detection speed of the radar detection device are obtained, and complementary features are determined based on the detection depth and the detection speed. The complementary features are used to characterize the supplementary features of the detection depth and the detection speed to the fused image. The radar detection device is used to acquire point cloud data. The location information of the obstacle is determined based on the complementary features and the fused image.

5. The method according to claim 1, characterized in that, The method further includes: The location information of the obstacle is determined based on the environmental image, and the bounding box corresponding to the obstacle is determined based on the location information of the obstacle. The point cloud data within the bounding box is identified as the target point cloud data. A fused image is determined based on the environmental image and the target point cloud data.

6. The method according to claim 5, characterized in that, The method further includes: A center point heatmap is determined based on the environmental image, wherein the center point heatmap is used to characterize the obstacle. Center point location; The bounding box corresponding to the obstacle is determined based on the position of the center point.

7. An obstacle avoidance device for unmanned aerial vehicles (UAVs), characterized in that, include: The acquisition module is used to acquire environmental images and point cloud data corresponding to the environment around the target drone during the flight of the target drone; A determination module is used to acquire obstacles in an environmental image, and, if the obstacles involved in the point cloud data match the obstacles, to determine a fused image based on the environmental image and the point cloud data; The control module is used to determine the location information corresponding to the obstacle based on the fused image, and send the location information to the target drone so that the target drone can avoid the obstacle.

8. An electronic device, characterized in that, include: At least one processor; Memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Man-machine moving obstacle monitoring method, readable storage medium and unmanned aerial vehicle

    CN110568861A

  • Unmanned aerial vehicle control method and system, unmanned aerial vehicle equipment and remote control equipment

    CN111984021A

  • Automatic driving environment sensing method and system

    CN112101092A

  • Obstacle detection and marking method and device for automatic driving and storage medium

    CN112419494A

  • Obstacle information generation method, device and equipment and computer readable storage medium

    CN115236672A

Cited By

  • Aerial obstacle avoidance method and system based on multi-machine vibration mutual inspection, medium and equipment

    CN122111086A