Large-range multi-resolution map construction method based on air-ground view cross matching

By using geometric transformation and deep learning-based methods to fuse heterogeneous information from UAVs and unmanned vehicles, the problem of unmanned vehicles acquiring complete maps in complex environments was solved, and large-scale multi-resolution maps were generated.

CN121783112APending Publication Date: 2026-04-03ZHONGBING INTELLIGENT INNOVATION RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Unmanned vehicles cannot acquire complete environmental map information in complex environments. How to effectively integrate the local mapping information of the unmanned vehicle itself, the aerial perception information of the drone, and the perception information of the operation payload, overcome the problem of heterogeneous information matching from different perspectives, dimensions, and resolutions, and build a large-scale, multi-resolution map.

Method used

A large-scale, multi-resolution map is generated by employing a cross-view image transformation based on geometric transformation and a multi-domain feature extraction method based on deep learning, combined with cross-domain feature alignment optimization, fusing aerial images from UAVs and ground images from unmanned vehicles, and combining 3D point cloud modeling and operational payload perception information.

Benefits of technology

It achieved heterogeneous information matching from different perspectives, dimensions, and resolutions, and completed cross-domain cross-matching and fusion of perception information from ground unmanned vehicles and aerial drones, generating a large-scale multi-resolution map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121783112A_ABST
    Figure CN121783112A_ABST
Patent Text Reader

Abstract

A large-range multi-resolution map construction method based on air-ground view cross matching mainly comprises the following steps: respectively transforming an air unmanned aerial vehicle image and an unmanned vehicle ground image to obtain an air polar coordinate image and a ground aerial view image; performing feature extraction on the aerial unmanned aerial vehicle image, the unmanned vehicle ground image, the aerial polar coordinate image and the ground aerial view image by adopting a deep learning-based multi-domain feature extraction method; geometric alignment is carried out on the extracted features, and a cross matching fusion result of the unmanned aerial vehicle aerial image information and the unmanned vehicle ground image information is obtained through optimization; on the basis of matching and fusion of aerial image information of the unmanned aerial vehicle and ground image information of the unmanned aerial vehicle, three-dimensional point cloud modeling information of a local area of the ground and operation load sensing information of the unmanned aerial vehicle are fused, and a large-range multi-resolution map is generated. According to the method, the sensing information of the ground unmanned vehicle and the air unmanned aerial vehicle can be matched and fused in a cross-domain manner, and a large-range multi-resolution map is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned vehicles or air-ground cooperative unmanned systems, and in particular to a method for constructing large-scale multi-resolution maps based on cross-matching of air and ground views. Background Technology

[0002] In complex environments with interference, obstruction, and non-line-of-sight conditions, unmanned vehicles cannot acquire complete environmental map information by relying solely on their own perception information. Therefore, how to effectively integrate the local mapping information of the unmanned vehicle itself, the aerial perception information of the UAV, and the perception information of the unmanned vehicle's operational payload based on air-ground collaboration, overcome the problem of heterogeneous information matching from different perspectives, dimensions, and resolutions, and collaboratively construct a large-scale, multi-resolution map is a key technical problem that unmanned vehicles need to solve for large-scale perception sharing in complex environments. Summary of the Invention

[0003] This disclosure provides a method for constructing large-scale multi-resolution maps based on cross-matching of air and ground views, which can cross-domain match and fuse perception information from ground unmanned vehicles and aerial drones to construct large-scale multi-resolution maps.

[0004] This method first uses geometric transformation-based cross-view image transformation to transform the aerial UAV image and the ground image of the unmanned vehicle, respectively, to obtain the aerial polar coordinate image and the ground bird's-eye view image; Based on this, feature extraction is performed on aerial drone images, unmanned vehicle ground images, aerial polar coordinate images, and ground bird's-eye view images using a deep learning-based multi-domain feature extraction method to obtain corresponding feature maps. Then, a feature alignment optimization method based on cross-domain features is used to geometrically align the extracted features, and the cross-matching fusion result of UAV aerial image information and unmanned vehicle ground image information is obtained through optimization. Based on this, the 3D point cloud modeling information of the local ground area of ​​the unmanned vehicle and the perception information of the operation load are integrated to finally generate a large-scale multi-resolution map.

[0005] Compared with existing technologies, the beneficial effects of this disclosure are: ① it solves the problem of heterogeneous information matching between air-to-ground unmanned platforms from different perspectives, dimensions, and resolutions; ② it realizes cross-domain cross-matching and fusion of perception information from ground unmanned vehicles and airborne unmanned aerial vehicles; ③ it can complete the construction and generation of large-scale multi-resolution maps. Attached Figure Description

[0006] The above and other objects, features and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments of this disclosure taken in conjunction with the accompanying drawings, in which the same reference numerals generally represent the same components.

[0007] Figure 1 Transform drone images into aerial polar coordinate images; Figure 2 This is a schematic diagram of the BEV feature extraction network; Figure 3 This is a schematic diagram of a two-stream feature extraction network; Figure 4 This is a schematic diagram of angle pre-alignment; Figure 5 This is the overall flowchart. Detailed Implementation

[0008] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0009] This disclosure provides a method for constructing large-scale, multi-resolution maps based on cross-matching of air and ground views. The main steps include: S1 uses a cross-view image transformation based on geometric transformation to transform the aerial UAV image and the ground image of the unmanned vehicle respectively, to obtain the aerial polar coordinate image and the ground bird's-eye view image; S2 uses a deep learning-based multi-domain feature extraction method to extract features from aerial drone images, unmanned vehicle ground images, aerial polar coordinate images, and ground bird's-eye view images to obtain corresponding feature maps. S3, adopts a feature alignment optimization method based on cross-domain features to perform geometric alignment of the extracted features, and obtains the cross-matching fusion result of UAV aerial image information and unmanned vehicle ground image information through optimization; S4, based on the matching and fusion of UAV aerial image information and unmanned vehicle ground image information, integrates 3D point cloud modeling information of local ground areas and unmanned vehicle operation load perception information to generate a large-scale multi-resolution map.

[0010] In one exemplary embodiment, a method for constructing a large-scale multi-resolution map based on air-ground view cross-matching according to the present disclosure is provided, as shown in the appendix. Figure 5 As shown, the specific steps include: Step 1: Transformation of the cross-view image of the open space based on geometric transformation To effectively address the issue of significant differences in the projected positions of scene objects during matching due to substantial visual differences between aerial drone and unmanned vehicle views, a cross-view image transformation based on geometric transformation is employed to transform both the aerial drone image and the ground image of the unmanned vehicle, resulting in an aerial polar coordinate image and a ground bird's-eye view image.

[0011] 1) Polar coordinate transformation of aerial UAV images By employing a rectangular projection method to adjust the UAV image coordinate system to be perpendicular to the ground, the UAV image can be considered orthogonal to the aerial UAV image captured by the aerial platform. Pixels along the horizontal straight line of the UAV ground image have similar depth. The aerial UAV image is then transformed into an aerial polar coordinate image through polar coordinate transformation, as shown below. Figure 1 As shown, it can be seen that the spatial layout of the ground image and the aerial polar coordinate image of the unmanned vehicle is roughly aligned.

[0012] 2) Inverse perspective transformation of unmanned vehicle ground images Due to the limitations of UAV flight altitude, the orthogonality between aerial UAV images and unmanned vehicle ground images cannot be strictly maintained. Therefore, the matching accuracy between the aerial polar coordinate image obtained by transforming the aerial UAV image and the unmanned vehicle ground image is not high, and in this project, it is mainly used for angle pre-alignment. To achieve accurate matching, inverse perspective transformation (IPM) will be used to transform the unmanned vehicle ground image into a ground bird's-eye view image. Inverse perspective transformation eliminates the perspective effect of the unmanned vehicle ground image, making the ground bird's-eye view image and the aerial UAV image have similar perspectives, thus enabling accurate matching between the ground bird's-eye view image and the aerial UAV image.

[0013] Step 2: Multi-domain feature extraction based on deep learning Because of the significant difference in perspective between aerial drone images and ground images of unmanned vehicles, traditional feature extraction methods would severely limit the accuracy and robustness of matching and localization. Therefore, a multi-domain feature extraction method based on a deep learning BEV feature extraction network is adopted to extract multi-domain invariant features.

[0014] like Figure 2 As shown, aerial drone images and ground-based bird's-eye view images are directly processed by the same U-Net network to extract features. Figure 3 As shown, for the ground images and aerial polar coordinate images of the unmanned vehicle, the features are extracted using dual-stream VGG (Visual Geometry Group). Two VGG networks are used to extract the features of the ground images of the unmanned vehicle and the aerial polar coordinate images of the unmanned drone, respectively.

[0015] Step 3: Feature alignment optimization of unmanned vehicle ground images and aerial UAV images based on cross-domain features To achieve accurate matching and fusion between the ground images of the unmanned vehicle (V2V) and the aerial images of the drone, the features of the two domains are directly aligned. To ensure the stability of the optimization, an angle pre-alignment method is first used to obtain the approximate relative angles between the V2V ground images and the aerial drone images. Therefore, the V2V ground image is used as a sliding window, and the correlation between the features of the V2V ground image and the features of the aerial polar coordinate image is calculated along the azimuth axis of the aerial polar coordinate image, generating a similarity score at each angle, such as... Figure 4 The red curve in the aerial polar coordinate feature image features shows that the red arrow indicates the middle position of the unmanned vehicle ground image. It can be seen that the position of the maximum similarity score corresponds to the potential relative azimuth angle of the unmanned vehicle ground image with respect to the aerial drone image.

[0016] After obtaining the pre-alignment results, the features of the unmanned vehicle ground image are aligned with the features of the aerial polar coordinate image, and the features of the aerial drone image are aligned with the features of the ground bird's-eye view image through geometric relationships. Finally, the results of the angle pre-alignment are used as the initial relative angles for joint optimization to obtain the accurate relative angles and relative positions of the aerial drone image and the unmanned vehicle ground image. Therefore, it is possible to perform accurate fusion and matching of air and ground cross views.

[0017] For aerial drone images and unmanned vehicle ground images at different resolutions, since image transformation and feature extraction operate on each pixel of the image, changes in resolution only affect the number of pixels to be operated on in the two steps of geometric transformation-based cross-view image transformation and deep learning-based multi-domain feature extraction. However, because aerial drone images and unmanned vehicle ground images at different resolutions correspond to different camera intrinsic parameters, different camera intrinsic parameters need to be used for alignment and optimization during feature alignment.

[0018] Step 4: Generate a large-scale, multi-resolution map Based on aerial imagery from UAVs and ground imagery from unmanned vehicles, a large-scale, multi-resolution map is generated by integrating 3D point cloud modeling information of local ground areas and the operational payload information of unmanned vehicles.

[0019] 1) Information fusion of 3D point cloud modeling in local ground areas Based on the matching and fusion of aerial UAV images and unmanned vehicle ground images, local 3D point cloud model information fusion is performed. First, multiple frames of unmanned vehicle ground point clouds in a local area need to be associated with the unmanned vehicle ground images and aerial UAV images. After obtaining the pose of each frame of point cloud using a combined inertial navigation system, a method of projecting each frame of point cloud onto each frame of image is used to obtain the precise matching relationship between multiple frames of unmanned vehicle ground point clouds and unmanned vehicle ground images. Similarly, the relative positional relationship between the UAV and the unmanned vehicle obtained by matching and fusing aerial UAV images and unmanned vehicle ground images can be used to obtain the precise matching relationship between multiple frames of unmanned vehicle ground point clouds and aerial UAV image patches. To fuse the 3D point clouds, a local binding optimization method is adopted. That is, in the existing fusion of aerial UAV images and unmanned vehicle ground images, information association between multiple frames of unmanned vehicle point clouds and unmanned vehicle ground images and aerial UAV images is added, and an optimization equation is constructed to obtain the fused map of the unmanned vehicle ground point clouds and images.

[0020] 2) Generation of large-scale, multi-resolution maps Building upon this foundation, the system integrates aerial drone imagery, local environmental information from ground-based unmanned vehicles, and shared perception information from operational payloads to create a 2.5D battlefield map with different perspectives, resolutions, and dimensions. The fusion of aerial drone and ground-based unmanned vehicle images from different viewpoints allows the map to encompass information from various perspectives, enabling multi-view observation. Because it allows for the matching and fusion of images at different resolutions, the map can have varying resolutions to meet the needs of different application scenarios, ultimately resulting in the construction and generation of a large-scale, multi-resolution map.

[0021] The above technical solutions are merely exemplary embodiments of the present invention. For those skilled in the art, based on the application methods and principles disclosed in the present invention, it is easy to make various types of improvements or modifications, and not limited to the methods described in the specific embodiments of the present invention. Therefore, the methods described above are merely preferred and not restrictive.

Claims

1. A method for constructing large-scale multi-resolution maps based on cross-matching of air and ground views, characterized in that, Includes the following steps: S1 uses a cross-view image transformation based on geometric transformation to transform the aerial UAV image and the ground image of the unmanned vehicle respectively, to obtain the aerial polar coordinate image and the ground bird's-eye view image; S2 uses a deep learning-based multi-domain feature extraction method to extract features from aerial drone images, unmanned vehicle ground images, aerial polar coordinate images, and ground bird's-eye view images to obtain corresponding feature maps. S3, adopts a feature alignment optimization method based on cross-domain features to perform geometric alignment of the extracted features, and obtains the cross-matching fusion result of UAV aerial image information and unmanned vehicle ground image information through optimization; S4, based on the matching and fusion of UAV aerial image information and unmanned vehicle ground image information, integrates 3D point cloud modeling information of local ground areas and unmanned vehicle operation load perception information to generate a large-scale multi-resolution map.

2. The method according to claim 1, characterized in that, Step S1 specifically includes: 1) Polar coordinate transformation of aerial UAV images The UAV image coordinate system is adjusted to be perpendicular to the ground using a rectangular projection method; Pixels along the horizontal straight line of the unmanned vehicle ground image have similar depths. By transforming the aerial drone image into an aerial polar coordinate image through polar coordinate transformation, the spatial layout of the unmanned vehicle ground image and the aerial polar coordinate image are roughly aligned. 2) Inverse perspective transformation of ground images of unmanned vehicles By employing inverse perspective transformation, the ground image of the unmanned vehicle is transformed into a ground bird's-eye view image, so that the ground bird's-eye view image and the aerial drone image have a similar perspective.

3. The method according to claim 1 or 2, characterized in that, In step S2: Aerial drone images and ground-based bird's-eye view images are processed by the same U-Net network to extract their respective features; Features were extracted from ground and aerial polar coordinate images of the unmanned vehicle using dual-stream VGG.

4. The method according to claim 1, characterized in that, Step S3 specifically includes: First, an angle pre-alignment method is used to obtain the approximate relative angles between the unmanned vehicle ground image and the aerial drone image: using the unmanned vehicle ground image as a sliding window, the correlation between the features of the unmanned vehicle ground image and the features of the aerial polar coordinate image is calculated along the azimuth axis of the aerial polar coordinate image, and a similarity score is generated at each angle; After obtaining the pre-alignment results, the ground image features of the unmanned vehicle are aligned with the aerial polar coordinate image features, and the aerial drone image features are aligned with the ground bird's-eye view image features using geometric relationships. Finally, the results of angle pre-alignment are used as the initial relative angles for joint optimization to obtain the accurate relative angles and relative positions of aerial UAV images and unmanned vehicle ground images; Specifically, different camera intrinsic parameters are used for alignment and optimization of aerial drone images and unmanned vehicle ground images with different resolutions during feature alignment optimization.

5. The method according to claim 1, characterized in that, Step S4 specifically includes: 1) Information fusion of 3D point cloud modeling in local ground areas Based on the matching and fusion of aerial drone images and unmanned vehicle ground images, local area 3D point cloud model information fusion is performed: First, the multi-frame unmanned vehicle ground point cloud of a local area is associated with the unmanned vehicle ground image and the aerial drone image. Based on the pose of each frame of point cloud obtained by using a combined inertial navigation system, the method of projecting each frame of point cloud onto each frame of image is used to obtain the accurate matching relationship between the multi-frame unmanned vehicle ground point cloud and the unmanned vehicle ground image. Similarly, based on the relative positional relationship between the drone and the vehicle obtained by matching and fusing aerial drone images and unmanned vehicle ground images, the precise matching relationship between multiple frames of unmanned vehicle ground point cloud and aerial drone image patches is obtained. A local bundling optimization method is used to fuse 3D point clouds. That is, in the fusion of existing aerial UAV images and unmanned vehicle ground images, multiple frames of unmanned vehicle point cloud information are added to associate with the information of unmanned vehicle ground images and aerial UAV images, and an optimization equation is constructed to obtain a fused map of unmanned vehicle ground point cloud and images. 2) Generation of large-scale, multi-resolution maps Based on this, by integrating aerial drone image information, ground unmanned vehicle local environmental information, and shared perception information of operational payloads, 2.5D battlefield maps with different perspectives, resolutions, and dimensions are obtained.