Air-ground adaptive fusion perception method

By acquiring visual data from unmanned logistics vehicles and drones, passable grid maps and pedestrian detection data are generated. Multi-level feature fusion is performed by determining fusion weights based on pitch angles, which solves the problem of perspective differences in air-ground collaborative perception and improves perception accuracy and safety.

CN120747703BActive Publication Date: 2025-12-23HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202511240114.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-23
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

In unmanned logistics vehicles, when using air-ground collaborative perception, the huge difference in perspective between the drone and the vehicle causes image projection distortion. Directly fusing air-ground data will introduce errors and affect perception accuracy.

Method used

An air-ground adaptive fusion perception method is adopted. By acquiring visual data from unmanned logistics vehicles and drones, passable grid maps and pedestrian detection data are generated. The fusion weights are determined based on the target pitch angle, and the ORP-Pyramid engine is used to perform multi-level feature fusion and information compensation to improve accuracy.

Benefits of technology

Effectively integrating air-to-ground data with large viewing angle differences improves the accuracy of air-to-ground adaptive fusion perception, ensuring the safe operation of unmanned logistics vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747703B_ABST
    Figure CN120747703B_ABST
Patent Text Reader

Abstract

The application provides an air-ground adaptive fusion perception method, and relates to the technical field of multi-sensor data fusion. The method determines the fusion weight of the unmanned aerial vehicle data collection according to the target pitch angle of the unmanned aerial vehicle, and the fusion weight and the size of the target pitch angle are negatively correlated. The fusion weight of the unmanned flow vehicle data collection is the value obtained by subtracting the fusion weight of the unmanned aerial vehicle data collection from 1. On this basis, multi-level feature fusion is performed with the help of an ORP-Pyramid engine, so that the air-ground data with a large visual angle difference can be effectively fused, and the precision of the air-ground adaptive fusion perception is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-sensor data fusion, and in particular to an air-ground adaptive fusion perception method. BACKGROUND

[0002] At present, in the field of unmanned flow vehicles, a single vehicle-mounted sensor, such as a camera or a laser radar, has inherent limitations in complex traffic scenes. Specifically, the line of sight of a ground vehicle is easily blocked by other traffic participants or roadside obstacles (such as buildings, fences), thereby forming a blind area of perception, which brings serious safety hazards to driving safety.

[0003] In order to overcome the limitations of ground view, the industry has proposed an air-ground collaborative perception scheme, that is, using the wide and unobstructed overhead view of an unmanned aerial vehicle or other aerial platform to assist the ground unmanned flow vehicle in environmental perception. However, in air-ground collaborative perception, there is usually a large view angle difference between the view angle of the unmanned aerial vehicle and the view angle of the vehicle-mounted sensor (the pitch angle of the unmanned aerial vehicle often changes in a large range between 30° and 75°), which will cause serious image projection distortion, especially when the pitch angle of the unmanned aerial vehicle is greater than 60°, the projection distortion rate is greater than 25%. In this case, if the feature fusion of air-ground data is directly performed, a large error will be introduced, which will seriously affect the perception accuracy.

[0004] Based on this, how to effectively fuse the air-ground data with a large view angle difference and improve the accuracy of air-ground adaptive fusion perception has become a technical problem to be solved. SUMMARY

[0005] Therefore, in order to solve the above technical problems, the present application provides an air-ground adaptive fusion perception method.

[0006] The present application adopts the following technical scheme:

[0007] An air-ground adaptive fusion perception method comprises:

[0008] acquiring first visual data collected by a first visual data collection device and second visual data collected by a second visual data collection device; the first visual data collection device is arranged on an unmanned flow vehicle, and the second visual data collection device is arranged on an unmanned aerial vehicle; the unmanned aerial vehicle flies at a target pitch angle at a target distance in front of the unmanned flow vehicle;

[0009] generating a passable grid map according to the first visual data, in which passable regions, unknown regions and obstacles are marked by grid units in the passable grid map;

[0010] generating pedestrian detection data according to the second visual data;

[0011] determine a first fusion weight of the first visual data and a second fusion weight of the second visual data according to the target pitch angle; a sum of the first fusion weight and the second fusion weight is 1; the second fusion weight is negatively correlated with a size of the target pitch angle;

[0012] perform multi-level feature fusion based on an ORP-Pyramid engine according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map, and the pedestrian detection data, to generate a fused environment representation;

[0013] identify an occluded area in the fused environment representation that originates from the first visual data, and compensate information of the occluded area by using the second visual data to obtain a target grid map;

[0014] output the target grid map.

[0015] Optionally, the passable grid map is generated according to the first visual data, and specifically includes:

[0016] input the first visual data into a BEVFormer encoder to obtain a passable grid map in a bird's-eye view perspective output by the BEVFormer encoder.

[0017] Optionally, the pedestrian detection data is generated according to the second visual data, and specifically includes:

[0018] input the second visual data into a YOLO-LE detector to obtain pedestrian detection data output by the YOLO-LE detector;

[0019] When there is a pedestrian in the occluded area, the pedestrian detection data includes a pedestrian 3D bounding box.

[0020] Optionally, the YOLO-LE detector is obtained by using a multi-modal knowledge distillation technology to distill knowledge learned by the BEVFormer encoder to a YOLO detector.

[0021] Optionally, the multi-level includes an object level, a region level, and a pixel level.

[0022] perform multi-level feature fusion based on an ORP-Pyramid engine according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map, and the pedestrian detection data, to generate a fused environment representation, and specifically includes:

[0023] In the object layer, the pedestrian detection data is fused with auxiliary positioning pedestrian information in the first visual data according to the second fusion weight and the first fusion weight, to obtain an object layer feature vector;

[0024] In the region layer, occluded region information in the second visual data is fused with the passable grid map according to the second fusion weight and the first fusion weight, to obtain a region layer feature vector;

[0025] In the pixel layer, the first visual data is fused with the second visual data according to the first fusion weight and the second fusion weight, to obtain a pixel layer feature vector;

[0026] The object layer feature vector, the region layer feature vector and the pixel layer feature vector are spliced in sequence, to obtain the fusion environment representation.

[0027] Optionally, the occluded region is compensated for information by using the second visual data, and the compensation specifically includes:

[0028] An occlusion compensation convolution operation is adopted to use target information located in the occluded region and detected in the second visual data to perform risk labeling or information filling on the fusion environment representation.

[0029] The above technical solutions are adopted, the fusion weight of the unmanned aerial vehicle target pitch angle is determined, the fusion weight and the size of the target pitch angle are negatively correlated, the fusion weight of the unmanned aerial vehicle target pitch angle is determined, the fusion weight of the unmanned aerial vehicle target pitch angle is determined, and the fusion weight of the unmanned aerial vehicle target pitch angle is determined. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0031] Figure 1 is a flowchart of an air-ground adaptive fusion perception method provided by an embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0033] Figure 1 is a flowchart of an air-ground adaptive fusion perception method provided by an embodiment of the present application. As shown in the figure, the flowchart includes: Figure 1

[0034] Step 101: acquiring first visual data collected by a first visual data collection device and second visual data collected by a second visual data collection device; the first visual data collection device is arranged on an unmanned vehicle, and the second visual data collection device is arranged on an unmanned aerial vehicle; the unmanned aerial vehicle flies at a target distance in front of the unmanned vehicle and at a target pitch angle.

[0035] Specifically, the first visual data collection device can be a vehicle-mounted surround-view camera system, which is composed of 8 cameras distributed at different positions around the vehicle, and can realize 360° dead-angle-free collection of visual data of the environment around the vehicle, providing omnidirectional information for the vehicle to perceive the surrounding road conditions.

[0036] The second visual data collection device can be a 4K high-definition gimbal camera carried on the unmanned aerial vehicle, which has a stable gimbal damping function and can keep the shooting picture stable during the flight of the unmanned aerial vehicle. The 4K resolution can provide high-definition visual data to meet the shooting requirements of long-distance and large-range scenes.

[0037] The target distance can be 50 meters or other distances, which are not specifically limited in the present application. The target pitch angle can be 75° or other angles, which are not specifically limited in the present application.

[0038] Step 102: generating a passable grid map according to the first visual data, in which the passable grid map marks passable areas, unknown areas and obstacles with grid units.

[0039] Specifically, the unmanned vehicle is equipped with a vehicle-mounted computing platform, which is responsible for processing the first visual data collected by the first visual data collection device and running a BEVFormer (Bird's-Eye-View Former, bird's-eye view generation model) encoder.

[0040] Based on this, in the embodiments of the present application, the passable grid map is generated according to the first visual data, which can specifically include:

[0041] ​The first visual data is input into the BEVFormer encoder to obtain a passable grid map in the bird's eye view perspective output by the BEVFormer encoder. In the passable grid map, the passable area, unknown area and obstacle are marked by grid units, wherein the accuracy of the grid unit can be 0.1 m / pixel, the passable area can be marked by green, the unknown area can be marked by gray, and the obstacle can be marked by red. The obstacle includes a construction fence.

[0042] Step 103: generating pedestrian detection data according to the second visual data.

[0043] Specifically, an airborne computing platform is carried on the unmanned aerial vehicle, responsible for processing the second visual data collected by the second visual data collection device, and running a YOLO (You Only Look Once) -LE (Lightweight and Efficient) detector.

[0044] Based on this, in the embodiment of the application, the pedestrian detection data is generated according to the second visual data, and specifically can include:

[0045] The second visual data is input into the YOLO-LE detector to obtain pedestrian detection data output by the YOLO-LE detector.

[0046] Wherein, when the total number of pixel points of the occluded area of the detected pedestrian image is greater than a preset number threshold, it is determined that the occluded area exists a pedestrian, and a pedestrian 3D bounding box is generated. The preset number threshold can be equal to 15. The pedestrian 3D bounding box contains the three-dimensional coordinates, size and confidence information of the pedestrian.

[0047] In the embodiment of the application, the YOLO-LE detector obtains the knowledge learned by the BEVFormer encoder by using a multi-modal knowledge distillation technology. In this way, the YOLO-LE detector can greatly reduce the computing load of the unmanned aerial vehicle end while ensuring the perception accuracy.

[0048] Step 104: determining a first fusion weight of the first visual data and a second fusion weight of the second visual data according to the target pitch angle; the sum of the first fusion weight and the second fusion weight is 1; the second fusion weight is negatively correlated with the size of the target pitch angle.

[0049] Specifically, the calculation formula of the second fusion weight can be:

[0050] ......(1)

[0051] Wherein, ; is the second fusion weight, Target pitch angle.

[0052] Step 105: Based on the ORP-Pyramid (Object-Relationship Pyramid) engine, multi-level feature fusion is performed according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map and the pedestrian detection data, to generate a fused environment representation.

[0053] In the embodiment of the present application, the multi-level includes an object layer, a region layer and a pixel layer.

[0054] Based on the ORP-Pyramid engine, multi-level feature fusion is performed according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map and the pedestrian detection data, to generate a fused environment representation, which can specifically include:

[0055] (1051) In the object layer, the pedestrian detection data is fused with the auxiliary positioning pedestrian information in the first visual data according to the second fusion weight and the first fusion weight, to obtain an object layer feature vector. The auxiliary positioning pedestrian information can include positioning related data of the vehicle itself and the surrounding environment, such as relative position, distance, etc.

[0056] (1052) In the region layer, the occluded region information in the second visual data is fused with the passable grid map according to the second fusion weight and the first fusion weight, to obtain a region layer feature vector.

[0057] (1053) In the pixel layer, the first visual data is fused with the second visual data according to the first fusion weight and the second fusion weight, to obtain a pixel layer feature vector, through the E2E-MFD (End-to-End Multi-Modal Fusion and Detection) technology.

[0058] (1054) The object layer feature vector, the region layer feature vector and the pixel layer feature vector are sequentially spliced to obtain a fused environment representation.

[0059] Step 106: An occluded region in the fused environment representation originated from the first visual data is identified, and information compensation is performed on the occluded region by using the second visual data, to obtain a target grid map.

[0060] In the embodiment of the present application, the information compensation on the occluded region by using the second visual data can specifically include:

[0061] The occlusion compensation convolution operation is used to label the risk or fill the information of the fused environment representation by using the target information detected in the second visual data located in the occluded area.

[0062] In a specific example, the perspective of the first visual data acquisition device of the unmanned flow vehicle is blocked by a large construction fence, and the situation behind the fence cannot be seen at all. The overhead perspective of the unmanned aerial vehicle is not blocked at all, and it can clearly detect that there is a pedestrian behind the fence. At this time, a 3x3 Gaussian filter can be used as a convolution kernel to perform convolution operation in the occluded area of the fused environment representation to generate a probability grid map in the occluded area. According to the distribution of the pedestrian pixel points (such as 18 pixels), a high probability value (such as ≥0.8) is assigned to the corresponding grid, and is marked as a "high-risk area" (red grid), thereby achieving information filling of the occluded area.

[0063] The target grid map is a grid map with a resolution of 0.1 meters / pixel, which fuses high-precision local information of ground perspective and unobstructed global information of overhead perspective, and clearly labels potential risks in the occluded area. In addition, the target grid map also plans a passable path with a lateral offset of 1.2 m and a curvature radius > 6 m, which is used to guide the unmanned flow vehicle to travel along the passable path and ensure the unmanned flow vehicle to avoid the occluded target safely.

[0064] The lateral offset of 1.2 m means that the passable path deviates from the originally planned base path (such as the center line of the road or the initial reference path) by a distance of 1.2 m. The curvature radius > 6 m means that the degree of bending when the path turns is small, and the curvature radius of the curve is greater than 6 meters. Such path design can ensure that the unmanned flow vehicle maintains stable driving while avoiding obstacles, reduces the risk of rollover, and ensures driving safety.

[0065] Step 107: output the target grid map.

[0066] It can be understood that the same or similar parts in the above embodiments can be mutually referred to, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0067] It should be noted that in the description of the present application, the terms "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is at least two.

[0068] Any process or method described in a flowchart illustration or described elsewhere in this specification can be understood and can be practiced as a set of steps that include one or more steps for performing the described logic functions and processes. The scope of preferred embodiments of the present application encompasses also other implementations involving additional or fewer steps, and alternative arrangements of the steps used or described, as will occur to those skilled in the art. The various steps of the processes described herein can be performed in the order shown or in a different order. All such modifications and variations are considered to be within the scope of the present application.

[0069] In the description of the specification, the description using the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the particular feature, structure, material or characteristic being described is included in at least one embodiment or example of the present application. The illustrative appearance of the above terms in various places in the specification are not necessarily intended to refer to the same embodiment or example. Moreover, the particular features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0070] Although the embodiments of the present application have been shown and described above, it is to be understood that the above-described embodiments are merely exemplary, and are not to be taken in a limiting sense, but the scope of the present application is not to be understood as being limited to the above-described embodiments but can be varied, modified, replaced and changed by those skilled in the art within the scope of the present application.

Claims

1. An air-ground adaptive fusion perception method, characterized in that, The method comprises the following steps: acquiring first visual data collected by a first visual data collection device and second visual data collected by a second visual data collection device; the first visual data collection device is arranged on an unmanned logistics vehicle, and the second visual data collection device is arranged on an aerial unmanned aerial vehicle; the aerial unmanned aerial vehicle flies at a target distance in front of the unmanned logistics vehicle at a target pitch angle; generating a passable grid map according to the first visual data, wherein in the passable grid map, passable areas, unknown areas and obstacles are marked by grid units; generating pedestrian detection data according to the second visual data; determining a first fusion weight of the first visual data and a second fusion weight of the second visual data according to the target pitch angle; the sum of the first fusion weight and the second fusion weight is 1; the second fusion weight is negatively correlated with the size of the target pitch angle; based on an ORP-Pyramid engine, performing multi-level feature fusion according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map and the pedestrian detection data to generate a fused environment representation; the multi-level includes an object layer, a region layer and a pixel layer; based on the ORP-Pyramid engine, performing multi-level feature fusion according to the first fusion weight, the second fusion weight, the first visual data, the second visual data, the passable grid map and the pedestrian detection data to generate a fused environment representation, specifically including: in the object layer, weighting and fusing the pedestrian detection data according to the second fusion weight and auxiliary positioning pedestrian information in the first visual data according to the first fusion weight to obtain an object layer feature vector; in the region layer, weighting and fusing occluded region information in the second visual data according to the second fusion weight and the passable grid map according to the first fusion weight to obtain a region layer feature vector; in the pixel layer, weighting and fusing the first visual data according to the first fusion weight and the second visual data according to the second fusion weight to obtain a pixel layer feature vector; sequentially splicing the object layer feature vector, the region layer feature vector and the pixel layer feature vector to obtain the fused environment representation; identifying an occluded region in the fused environment representation that originates from the first visual data, and compensating information of the occluded region by using the second visual data to obtain a target grid map; outputting the target grid map.

2. The air-ground adaptive fusion perception method according to claim 1, characterized in that, Generating a passable grid map according to the first visual data specifically comprises: inputting the first visual data into a BEVFormer encoder to obtain a passable grid map in a bird's eye view perspective output by the BEVFormer encoder.

3. The air-ground adaptive fusion perception method according to claim 2, characterized in that, Generating pedestrian detection data according to the second visual data specifically comprises: input the second visual data into a YOLO-LE detector to obtain pedestrian detection data output by the YOLO-LE detector; The pedestrian detection data comprises a 3D bounding box of a pedestrian when the occluded area exists the pedestrian.

4. The air-ground adaptive fusion perception method according to claim 3, characterized in that, The YOLO-LE detector distills the knowledge learned by the BEVFormer encoder to a YOLO detector by using a multi-modal knowledge distillation technology.

5. The air-ground adaptive fusion perception method according to claim 1, characterized in that, The second visual data is used to compensate information of the occluded area, specifically comprising: An occlusion compensation convolution operation is used to label risks or fill information of the fusion environment representation by using target information located in the occluded area and detected in the second visual data.