Pose estimation method and device based on air-ground cross-view image matching

By generating multi-view simulation forward-view images and combining magnetic field azimuth angle for feature points matching, the matching difficulties caused by the difference in imaging angles and distances in the image matching across view angles of space is solved, and high-precision pose estimation and position pose estimation are achieved, which is suitable for detection tasks in unknown areas.

CN120495416AActive Publication Date: 2025-08-15AEROSPACE INFORMATION RES INST CAS

Patent Information

Application Number
CN202510980779.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-15
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

The existing open-ground cross-view image matching method has difficulty in matching and low pose estimation accuracy due to the difference in imaging angle and imaging distance in unknown area detection tasks, especially in the absence of satellite positioning and navigation systems, which are difficult to achieve high-precision position and posture estimation.

Method used

By generating a global digital elevation model and dividing the grid, combining the magnetic field azimuth angle and the visual sensor parameters of the ground moving platform, a multi-view simulation forward image is generated, and a magnetic field azimuth angle is used to filter similar simulated forward images for feature point matching, solving the pose matrix, and finally optimizing the motion trajectory.

Benefits of technology

The accuracy and reliability of cross-view angle matching are improved, the position and attitude estimation problem under the satellite-free positioning system in unknown areas is solved, and high-precision pose estimation is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495416A_ABST
    Figure CN120495416A_ABST
Patent Text Reader

Abstract

The invention discloses a pose estimation method and device based on air-ground cross-view image matching, and belongs to the technical field of image matching and positioning. According to the method, a global digital elevation model (DEM) and an orthographic image are generated by obtaining an aerial image sequence of a target area and are divided into a plurality of grids, and a multi-view-angle simulation forward view image is generated for each grid. After a visual sensor carried by a ground mobile platform obtains a ground front-view image, a simulation front-view image similar to the ground front-view image is screened by using a magnetic field azimuth angle, and an optimal grid is determined through feature point extraction and matching. And resolving a pose matrix of the ground mobile platform based on the feature point coordinates of the optimal grid, and further improving pose estimation precision through track reconstruction and optimization. According to the method, the problems of matching difficulty and low pose estimation precision caused by relatively large difference between imaging angles and imaging content scales of the aerial image and the ground front view image are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image matching and positioning, and in particular relates to a method and device for posture estimation based on air-ground cross-view image matching. Background Art

[0002] Cross-view image matching involves matching two images with significantly different imaging angles and distances. Aerial images are typically calculated from a sequence of overhead images of the target area, acquired by cameras aboard aerial vehicles such as drones and satellites. Their imaging plane is nearly parallel to the ground, and due to their high altitude, orthophotos capture a large target area. Ground-based front-view images, on the other hand, are typically acquired from a front-view perspective by cameras carried or mounted on ground-based mobile platforms, such as personnel and vehicles, entering the target area. Their imaging plane is nearly perpendicular to the ground, and due to their proximity to the target area, the captured area is smaller. Due to the differences in imaging angles and distances between aerial orthophotos and ground-based front-view images, their image content and scale differ significantly.

[0003] In many exploration missions into unknown areas, where global navigation satellite systems are often unavailable or unavailable, personnel often utilize aerial observations of unknown target areas using aircraft such as drones, satellites, rovers, and landers to ensure mission safety. By processing the resulting aerial bird's-eye view image sequences, they can obtain data on the target area's topography and landforms, which can be used for subsequent exploration missions. Cross-view image matching and positioning technology based on aerial orthophotos and ground-based front-view images plays a crucial role in exploration missions into unknown areas without satellite positioning and navigation systems.

[0004] Currently, commonly used pose estimation methods based on cross-view matching between aerial and ground-view images mainly fall into three categories: those based on traditional image feature matching, those based on deep learning feature matching, and those based on image transformation. The drawbacks of traditional image feature matching methods lie in the differences in imaging angle and distance between aerial orthophotos and ground-view images, resulting in significant differences in image content and scale. Therefore, traditional feature point extraction and matching methods struggle to meet the requirements for high-precision pose estimation. Pose estimation methods based on deep learning feature matching offer superior performance compared to traditional image features. However, like traditional features, the significant differences in viewpoint and scale between aerial and ground-view images lead to poor feature matching results, making them difficult to meet pose estimation requirements. While pose estimation methods based on image transformation narrow the viewpoint gap between ground-view and satellite images, they fail to consider the position, attitude, and intrinsic parameters of the forward-view sensor. Consequently, significant image distortion occurs after image transformation, resulting in less than ideal matching and localization results. In addition, some methods require roughly determining the position range of the ground front view image and the aerial image based on prior information such as GNSS navigation positioning, and then performing more accurate cross-perspective matching. This will be difficult to apply to unknown application scenarios that lack satellite positioning information. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a pose estimation method and device based on air-ground cross-view image matching, which can solve the existing problem of matching and positioning difficulties caused by the differences in imaging angles and imaging distances between orthophotos and ground front-view images, and better complete the position and pose estimation of ground mobile platforms in unknown area detection tasks without satellite positioning and navigation systems.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A pose estimation method based on air-ground cross-view image matching, the method comprising:

[0008] Step 1: Obtain an aerial image sequence of the target area and generate a global digital elevation model (DEM) and orthophoto image;

[0009] Step 2: Grid the global DEM and orthoimage, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the visual sensor internal parameters of the ground mobile platform;

[0010] Step 3: The ground mobile platform acquires ground front-view images at different orientations, and uses the magnetic field azimuth to select simulated front-view images similar to the ground front-view images;

[0011] Step 4: extracting and matching feature points of the similar ground front view image and the simulated front view image, and determining the grid where the simulated front view image with the largest number of matching feature points is located as the optimal grid;

[0012] Step 5: Calculate the pose matrix of the ground mobile platform according to the coordinates of the feature points corresponding to the optimal grid;

[0013] Step 6: Optimize the pose matrix and output the motion trajectory of the ground mobile platform.

[0014] On the other hand, the present invention provides a pose estimation device based on air-ground cross-view image matching, comprising:

[0015] The acquisition unit is used to obtain aerial image sequences of the target area and generate global digital elevation model (DEM) and orthophoto images;

[0016] The simulation unit is used to grid the global DEM and orthophoto image, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the internal parameters of the visual sensor of the ground mobile platform;

[0017] A selection unit is used for the ground mobile platform to obtain ground front-view images of different orientations, and to select a simulated front-view image similar to the ground front-view image by using the magnetic field azimuth angle;

[0018] A matching unit is used to extract and match feature points of the similar ground front view image and the simulated front view image, and determine that the grid where the simulated front view image with the largest number of matched feature points is located is the optimal grid;

[0019] A solving unit, used to solve the pose matrix of the ground mobile platform according to the coordinates of the feature points corresponding to the optimal grid;

[0020] The output unit is used to optimize the posture matrix and output the motion trajectory of the ground mobile platform.

[0021] In a third aspect, the present invention provides an electronic device comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned pose estimation method based on air-ground cross-view image matching.

[0022] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the aforementioned pose estimation method based on air-ground cross-view image matching.

[0023] The beneficial effects of the present invention are:

[0024] The method of the present invention divides the DEM and orthophoto of the target area generated by aerial images into grid areas, and generates multiple simulated front-view images at different perspectives within each grid. The angles and scales of these grid-scale simulated images are close to those of the front-view images of the ground mobile platform. The image matching range is further narrowed by the magnetic field azimuth of the ground mobile platform, resulting in a cross-perspective matching result with higher accuracy and better reliability.

[0025] The method of the present invention does not directly convert the aerial image into the perspective image of the ground forward-looking image, but combines the magnetic field azimuth and intrinsic parameters of the forward-looking sensor to generate a grid-scale multi-perspective simulated forward-looking image, effectively solving the scale and image distortion problems after conversion.

[0026] The method of the present invention does not require a satellite positioning navigation system to obtain the initial position of a ground mobile platform. Image matching and positioning are performed through grid-scale simulated forward-looking images, solving the problem of position and attitude estimation in detection tasks in unknown areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flow chart of the pose estimation method based on air-ground cross-view image matching of the present invention;

[0028] Figure 2 Generate schematic diagrams for grid-scale simulation front-view images;

[0029] Figure 3 Determine the schematic diagram for simulating the front view image for the ground vision sensor;

[0030] Figure 4 Updated schematic for ground vision sensor trajectories. DETAILED DESCRIPTION

[0031] The present invention will be further described below with reference to the accompanying drawings and examples.

[0032] like Figure 1 As shown, the present invention provides a pose estimation method based on air-ground cross-view image matching, and the specific implementation steps are as follows:

[0033] Step 1: Obtain an aerial image sequence of the target area and generate a global DEM and orthophoto images. Use visual sensors such as ground observation cameras carried by drones and other aircraft to obtain an aerial image sequence of the target area in advance, complete image alignment by extracting and matching feature points between images, and then use common three-dimensional reconstruction methods such as multi-view stereo geometry to generate dense point clouds and textures. Then, use coordinate system projection transformation to further generate a global DEM and orthophoto images of the target area.

[0034] Step 2: Divide the global DEM and orthophoto image into grids, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the visual sensor internal parameters of the ground mobile platform; Figure 2 As shown, the global DEM and orthophoto images are divided into grids, and obtain the DEM and local orthophoto of each grid. At the same time, each grid is numbered, that is, {1, 2, 3, …, The ground mobile platform is equipped with indivual( ≥1) The same visual sensors have the same internal parameters such as focal length and principal point, but have different installation azimuths and positions on the platform. The center point of each grid is then used as the simulation position. , with the magnetic field azimuth angle at intervals of 5 degrees as the simulation attitude , that is, {0°, 5°, 10°,…, 350°, 355°}, the intrinsic parameters of the visual sensor are used as the intrinsic parameters of the simulated visual sensor, and the local DEM and orthophoto image are used for rendering in each grid to obtain N (here based on the setting of 5°, N=72) simulated front view images of different simulation postures, and the entire target area is obtained. A simulated front view image.

[0035] Step 3: The ground mobile platform passes A visual sensor obtains ground front-view images in different directions, and uses the magnetic field azimuth to screen simulated front-view images similar to the ground front-view images; Figure 3 As shown, equipped with indivual( ≥1) The ground mobile platform with a visual sensor enters the target area, where the visual sensor in the front direction of the ground mobile platform serves as the standard visual sensor ( = 1), its position and attitude are used as the position and attitude of the ground mobile platform. There is a certain azimuth difference between other visual sensors and the standard visual sensor, and the extrinsic parameters between other visual sensors and the standard visual sensor are fixed. These extrinsic parameter matrices are recorded as , which means the The transformation matrix of the position and posture of the visual sensor and the standard visual sensor is hour, is the unit matrix. Multiple visual sensors are used to obtain ground front-view images of different orientations in real time, and the magnetic field azimuth sensor carried by the ground mobile platform is used to obtain the magnetic field azimuth of the standard visual sensor. The magnetic field azimuth of each visual sensor is then calculated by the extrinsic parameters between the other visual sensors and the standard visual sensor. Then, based on the magnetic field azimuth of each visual sensor, the closest set of simulated postures in the set of N=72 simulated front-view images of the target area is found. A simulated front view image is obtained, and the entire ground mobile platform will Groups of different azimuths A simulated front view image.

[0036] Step 4: Extract and match the feature points of the similar ground front view image and the simulated front view image, determine the grid where the simulated front view image with the largest number of feature point matches is located as the optimal grid, and obtain the 3D coordinates corresponding to the 2D feature points in the optimal grid DEM; use the ground front view images of different azimuths obtained by multiple visual sensors in step 3 and the corresponding 3D coordinates of each visual sensor. Extract and match the image feature points of the simulated front view images, and select the grid number of the simulated image with the largest number of successful feature point matches as the potential grid number of the current visual sensor. A visual sensor will get grid numbers, find the grid number with the largest number The grid where the ground mobile platform is currently located. The target area DEM is determined as the best matching grid DEM. The ground front view image obtained by the visual sensor is consistent with the ground front view image in the grid. The corresponding 2D feature point coordinates of the simulated front view image that are successfully matched and ,in, For the A set of 2D feature point pixel coordinates on the ground front view image obtained by a visual sensor , For its grid A set of 2D feature point pixel coordinates on the simulated front view image corresponding to the location. For each visual sensor, according to its The position used when rendering the corresponding simulation image and posture ,turn up A set of 3D point coordinates corresponding to the DEM image , which is the coordinates of the 2D feature points on the ground front view image The corresponding set of 3D point coordinates.

[0037] Step 5: Construct a set of equations based on the 2D feature point coordinates of the ground front view image corresponding to the optimal grid and their corresponding 3D coordinates to solve the pose matrix of the ground mobile platform; use step 4 to obtain The 2D feature point coordinates of each visual sensor on the ground front view image and its corresponding set of 3D point coordinates , where each pair of 2D-3D points can construct an equation as shown in formula (1):

[0038] (1)

[0039] In formula (1), express The pixel coordinates of a 2D point in , express The coordinates of the corresponding 3D point in, Represents the intrinsic parameter matrix of each visual sensor, with a size of , Represents the position and posture matrix of each visual sensor , the size is , subscript It means taking the first 3 rows of the matrix. and The 2D-3D point pairs in the can construct a system of equations to calculate The position and attitude matrices of each visual sensor , at least 6 pairs of 2D-3D point pairs are required for calculation.

[0040] Then, the external parameters of the pose between the visual sensor and the standard visual sensor are , calculated according to formula (2) to obtain the standard visual sensor latent position and pose matrices ,have A vision sensor can calculate the standard vision sensor latent position and pose matrices , then indivual Determine the position and attitude matrix of a unique standard vision sensor statistically , that is, the position and attitude matrix of the ground mobile platform Specific statistical methods that can be used include mean, median, random sampling consistency, etc.

[0041] (2)

[0042] Step 6: Optimize the pose matrix, output the motion trajectory of the ground mobile platform, reconstruct the DEM based on the trajectory, and construct the constrained optimized motion trajectory by combining the DEM generated by the aerial image; during the detection process of the target area, the ground mobile platform Obtain the position and posture of the standard vision sensor at each fixed moment in real time through steps 3 to 5 , that is, the position and attitude matrix of the ground mobile platform , and finally generate the motion trajectory of the ground mobile platform .

[0043] like Figure 4 As shown in FIG, after obtaining the trajectory points of at least three moments of the ground mobile platform through step 6, the ground front view image sequence of these moments is obtained by using the standard visual sensor The position and attitude matrix corresponding to each front view image , the dense point cloud and texture of this section of the trajectory are generated by common multi-view stereo geometry and other 3D reconstruction methods, and then the DEM of this section of the trajectory is further generated by coordinate system projection transformation, which is recorded as The local DEM of the target area corresponding to all the grids passed during the journey is spliced together and recorded as ,use and Construct constraint equations to further optimize the position and attitude of the standard vision sensor. Indicates that a 2D feature point on the current ground front view image is The corresponding 3D point coordinates, In construction It has been associated with the 2D feature points of the image. Indicates that these 2D feature points are The corresponding 3D point coordinates, Then it is associated with the 2D feature points on the image in step 5. A total of right and ,make For the optimization matrix that needs to be optimized, the optimization equation shown in formula (3) can be constructed for optimization and solution:

[0044] (3)

[0045] The optimized matrix is calculated Then, the position and attitude matrix of the ground mobile platform in the target area at the current moment is calculated by formula (4): Optimize and output the latest optimized position and posture .

[0046] (4)

[0047] Repeat steps 3 to 6 above to obtain and optimize the trajectory of the ground mobile platform in real time at certain time intervals.

[0048] On the other hand, the present invention provides a pose estimation device based on air-ground cross-view image matching, which includes various modules capable of implementing the various steps of the aforementioned method, specifically including:

[0049] The acquisition unit is used to obtain aerial image sequences of the target area and generate global digital elevation model (DEM) and orthophoto images;

[0050] The simulation unit is used to grid the global DEM and orthophoto image, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the internal parameters of the visual sensor of the ground mobile platform;

[0051] A selection unit is used for the ground mobile platform to obtain ground front-view images of different orientations, and to select simulated front-view images similar to the ground front-view images by using the magnetic field azimuth angle;

[0052] A matching unit is used to extract and match feature points of the similar ground front view image and the simulated front view image, and determine that the grid where the simulated front view image with the largest number of matched feature points is located is the optimal grid;

[0053] A solving unit, used to solve the pose matrix of the ground mobile platform according to the coordinates of the feature points corresponding to the optimal grid;

[0054] The output unit is used to optimize the posture matrix and output the motion trajectory of the ground mobile platform.

[0055] In a third aspect, the present invention provides an electronic device comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned pose estimation method based on air-ground cross-view image matching.

[0056] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the aforementioned pose estimation method based on air-ground cross-view image matching.

[0057] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A pose estimation method based on air-ground cross-view image matching, characterized in that: The method comprises: Step 1: Obtain an aerial image sequence of the target area and generate a global digital elevation model (DEM) and orthophoto image; Step 2: Grid the global DEM and orthoimage, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the visual sensor internal parameters of the ground mobile platform; Step 3: The ground mobile platform acquires ground front-view images at different orientations, and uses the magnetic field azimuth to select simulated front-view images similar to the ground front-view images; Step 4: extracting and matching feature points of the similar ground front view image and the simulated front view image, and determining the grid where the simulated front view image with the largest number of matching feature points is located as the optimal grid; Step 5: Calculate the pose matrix of the ground mobile platform according to the coordinates of the feature points corresponding to the optimal grid; Step 6: Optimize the pose matrix and output the motion trajectory of the ground mobile platform.

2. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 1 comprises: performing image alignment on the aerial image sequence, generating a dense point cloud by a three-dimensional reconstruction method, and generating a global DEM and an orthophoto image of the target area based on a coordinate system projection.

3. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 2 includes: The obtained global DEM and orthophoto images are divided into Grids, each grid has its corresponding local DEM and orthophoto image; The center point of each grid is used as the simulation position, the preset multi-directional magnetic field azimuth is used as the simulation posture, and the visual sensor internal parameters of the ground mobile platform are used as the simulated visual sensor internal parameters. Local DEM and orthophoto images are used for rendering in each grid to obtain N simulated front view images with different postures.

4. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 3 comprises: Using ground mobile platforms The visual sensors acquire the ground front view images in different directions in real time. The visual sensor in the front direction of the ground mobile platform is regarded as the standard visual sensor, and the extrinsic parameters of the other visual sensors and the standard visual sensor are fixed. According to the magnetic field azimuth of each visual sensor, find the closest set of N simulated poses to the target area simulated front view image set A simulated front view image is obtained, and the entire ground mobile platform will Groups of different azimuths A simulated front view image.

5. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 4 comprises: The ground front view images at different azimuths obtained by multiple visual sensors are consistent with the corresponding images of each visual sensor. Extract and match the image feature points of the simulated front view images, select the grid number of the simulated image with the largest number of successful feature point matches as the potential grid number of the current visual sensor, and determine the grid where the ground mobile platform is located based on the matching results of multiple visual sensors; According to the local DEM of the grid, find the 3D feature point coordinates corresponding to the 2D feature point coordinates of the matching simulation image.

6. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 5 comprises: Using step 4, we can get Calculate the coordinates of the 2D feature points on the ground front view image of each visual sensor and the corresponding set of 3D feature point coordinates. The position and posture matrix of each visual sensor; The standard vision sensor’s pose parameters are calculated by comparing the visual sensor with the standard vision sensor. potential position and pose matrices, and then The potential position and attitude matrices are statistically used to determine the unique position and attitude matrix of the standard vision sensor, that is, the position and attitude matrix of the ground mobile platform.

7. The pose estimation method based on air-ground cross-view image matching according to claim 1, characterized in that: The step 6 comprises: During the detection process of the target area, the ground mobile platform obtains the position and posture of the standard visual sensor at each fixed time in real time at a fixed time interval, that is, the position and posture matrix of the ground mobile platform, and finally generates the motion trajectory of the ground mobile platform; Using the ground front view image sequence obtained by the standard visual sensor and the position and posture matrix corresponding to each ground front view image, the dense point cloud and texture of the motion trajectory of the ground mobile platform are generated through the 3D reconstruction method. Then, the digital elevation model of this trajectory is generated through the coordinate system projection transformation, which is recorded as , the local DEM of the target area corresponding to all the grids passed during the journey is spliced and recorded as ,use and Construct constraint equations to optimize the position and posture of the standard vision sensor and output the optimized motion trajectory of the ground mobile platform.

8. A pose estimation device based on air-ground cross-view image matching, characterized in that: include: The acquisition unit is used to obtain aerial image sequences of the target area and generate global digital elevation model (DEM) and orthophoto images; The simulation unit is used to grid the global DEM and orthophoto image, and generate a multi-view simulated front view image for each grid by combining the preset magnetic field azimuth and the internal parameters of the visual sensor of the ground mobile platform; A selection unit is used for the ground mobile platform to obtain ground front-view images of different orientations, and to select a simulated front-view image similar to the ground front-view image by using the magnetic field azimuth angle; A matching unit is used to extract and match feature points of the similar ground front view image and the simulated front view image, and determine that the grid where the simulated front view image with the largest number of matched feature points is located is the optimal grid; A solving unit, used to solve the pose matrix of the ground mobile platform according to the coordinates of the feature points corresponding to the optimal grid; The output unit is used to optimize the posture matrix and output the motion trajectory of the ground mobile platform.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; Wherein, when one or more programs are executed by the one or more processors, the one or more processors implement the pose estimation method based on air-ground cross-view image matching as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that Executable instructions are stored thereon, which, when executed by a processor, enable the processor to implement the pose estimation method based on air-ground cross-view image matching as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Ground-to-unmanned aerial vehicle laser point cloud cross-view relocation method based on semantic map

    CN115797422A

  • Air-ground heterogeneous collaborative mapping method and device, equipment and storage medium

    CN117191005A

  • Vehicle-mounted system positioning method based on air-ground view angle image collaboration, terminal and storage medium

    CN117422764A

  • Method of using image warping for geo-registration feature matching in vision-aided positioning

    US20150199556A1

  • Calibration of angle measuring sensors (IMU) in portable devices using elevation model map (DEM) and landform signature

    WO2024196332A1

Cited By

  • Ground camera trusted visual positioning method and device based on satellite reference base map

    CN121921360A

  • Satellite reference map-based ground camera reliable visual positioning method and device

    CN121921360B