Pose estimation method and apparatus based on air-to-ground cross-view image matching
By generating multi-view simulated front view images and combining them with magnetic field azimuth angles for feature point matching, the problem of high-precision pose estimation in cross-view image matching between air and ground was solved, and high-precision position and attitude estimation in unknown areas was achieved.
Patent Information
- Application Number
- CN202510980779.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing air-to-ground cross-view image matching methods struggle to achieve high-precision pose estimation in unknown area exploration missions, especially in the absence of satellite positioning and navigation systems. Traditional and deep learning feature matching methods are ineffective, resulting in image distortion after transformation and requiring prior information that is difficult to obtain.
By generating a global digital elevation model and dividing it into grids, and combining the magnetic field azimuth angle and the intrinsic parameters of the visual sensors of the ground mobile platform, a multi-view simulated front view image is generated. The magnetic field azimuth angle is used to filter similar images for feature point matching, solve the pose matrix, and optimize the motion trajectory.
It improves the accuracy and reliability of cross-view matching, solves the problem of position and attitude estimation in unknown areas, and can achieve high-precision pose estimation without the need for a satellite positioning system.
Smart Images

Figure CN120495416B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image matching and localization technology, specifically relating to a pose estimation method and apparatus based on cross-view image matching between air and ground. Background Technology
[0002] Cross-view image matching refers to matching two images with significantly different imaging angles and distances. Aerial images are typically calculated from a sequence of top-down views of the target area acquired by cameras mounted on aerial vehicles such as drones and satellites. Their imaging plane is almost parallel to the ground, and due to their higher altitude, orthophotos capture a large target area. In contrast, ground-based forward-looking images are usually acquired from forward-looking cameras carried or mounted on ground-based mobile platforms such as personnel or vehicles entering the target area. Their imaging plane is nearly perpendicular to the ground, and due to their closer proximity to the target area, they capture a smaller area. Because of the differences in imaging angles and distances between aerial and ground-based orthophoto images, their content and scale differ considerably.
[0003] In many exploration missions targeting unknown areas, which often lack or cannot utilize global navigation satellite systems, researchers typically employ drones, satellites, rovers, and landers to observe the target area from the air to ensure mission safety. By processing the aerial image sequences obtained from these observations, topographical and other data can be extracted for subsequent exploration tasks. Cross-view image matching and localization technology based on aerial orthophotos and ground-based frontal images plays a crucial role in exploring unknown areas without satellite positioning and navigation systems.
[0004] Currently, commonly used pose estimation methods based on cross-viewpoint matching of aerial and ground-based front-view images mainly fall into three categories: pose estimation methods based on traditional image feature matching, pose estimation methods based on deep learning feature matching, and pose estimation methods based on image transformation. The drawback of traditional image feature matching-based pose estimation methods lies in the differences in imaging angles and distances between aerial orthophotos and ground-based front-view images. These images exhibit significant differences in content and scale, making it difficult for traditional feature point extraction and matching methods to meet the requirements of high-precision pose estimation. While deep learning feature matching-based pose estimation methods outperform traditional image features, like traditional features, the significant viewpoint and scale differences between aerial and ground images result in poor feature matching results, failing to meet the needs of pose estimation. Although image transformation-based pose estimation methods reduce the viewpoint difference between ground-based front-view images and satellite images, they do not consider the position, attitude, and intrinsic parameters of the forward-looking sensor. This leads to significant image distortion after image transformation, resulting in less than ideal matching and localization performance. In addition, some methods require the approximate location range of the ground front view image and the aerial image based on prior information such as GNSS navigation and positioning, and then perform more accurate cross-view matching, which will be difficult to apply to unknown application scenarios that lack satellite positioning information. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a pose estimation method and apparatus based on cross-view image matching between air and ground. This method can solve the existing problem of difficulty in matching and positioning caused by the differences in imaging angles and imaging distances between orthophotos and ground frontal images. It can better complete the position and attitude estimation of ground mobile platforms in unknown area exploration missions without satellite positioning and navigation systems.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A pose estimation method based on cross-view image matching between air and ground, the method comprising:
[0008] Step 1: Acquire aerial image sequences of the target area and generate a global digital elevation model (DEM) and orthophotos;
[0009] Step 2: Divide the global DEM and orthophoto image into grids, and generate multi-view simulated front view images for each grid by combining the preset magnetic field azimuth angle and the intrinsic parameters of the visual sensor of the ground mobile platform.
[0010] Step 3: The ground mobile platform acquires front view images of the ground from different directions, and uses the magnetic field azimuth angle to filter out simulated front view images that are similar to the ground front view images;
[0011] Step 4: Extract and match feature points between the similar ground front view images and the simulated front view images, and determine the grid where the simulated front view image with the most matched feature points is located as the optimal grid.
[0012] Step 5: Calculate the pose matrix of the ground mobile platform based on the coordinates of the feature points corresponding to the optimal mesh;
[0013] Step 6: Optimize the pose matrix and output the motion trajectory of the ground mobile platform.
[0014] On the other hand, the present invention provides a pose estimation device based on air-to-ground cross-view image matching, comprising:
[0015] The acquisition unit is used to acquire aerial image sequences of the target area and generate a global digital elevation model (DEM) and orthophotos.
[0016] The simulation unit is used to divide the global DEM and orthophoto into grids, and generate multi-view simulated front view images for each grid by combining the preset magnetic field azimuth angle and the visual sensor intrinsic parameters of the ground mobile platform.
[0017] The selection unit is used by the ground mobile platform to acquire ground front view images from different orientations and to filter simulated front view images that are similar to the ground front view images using the magnetic field azimuth angle.
[0018] The matching unit is used to extract and match feature points between the similar ground front view image and the simulation front view image, and determine the grid where the simulation front view image with the most matched feature points is located as the optimal grid.
[0019] The calculation unit is used to calculate the pose matrix of the ground mobile platform based on the coordinates of the feature points corresponding to the optimal grid.
[0020] The output unit is used to optimize the pose matrix and output the motion trajectory of the ground mobile platform.
[0021] Thirdly, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned pose estimation method based on air-to-ground cross-view image matching.
[0022] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned pose estimation method based on cross-view image matching between air and ground.
[0023] The beneficial effects of this invention are as follows:
[0024] The method described in this invention divides the target area DEM generated from aerial images and orthophotos into grid regions. Multiple simulated front view images with different perspectives are generated within each grid. The simulated images at these grid scales are close to the front view images of the ground mobile platform in terms of angle and scale. Furthermore, the magnetic field azimuth angle of the ground mobile platform is used to narrow the image matching range, resulting in cross-view matching results with higher accuracy and better reliability.
[0025] The method described in this invention does not directly convert aerial images into ground-based forward-looking images. Instead, it combines information such as the magnetic field azimuth and intrinsic parameters of the forward-looking sensor to generate multi-view simulated forward-looking images at a grid scale, effectively solving the problems of scale and image distortion after conversion.
[0026] The method described in this invention does not require a satellite positioning and navigation system to obtain the initial position of the ground mobile platform. Instead, it performs image matching and positioning through grid-scale simulated front view images, thus solving the problem of position and attitude estimation in exploration missions in unknown areas. Attached Figure Description
[0027] Figure 1 This is a flowchart of the pose estimation method based on cross-view image matching between air and ground according to the present invention.
[0028] Figure 2 A schematic diagram for generating a front view image for grid-scale simulation;
[0029] Figure 3 A schematic diagram for determining the simulated front view image of a ground vision sensor;
[0030] Figure 4 This is a schematic diagram of the trajectory update for ground-based visual sensors. Detailed Implementation
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0032] like Figure 1 As shown, this invention provides a pose estimation method based on cross-view image matching between air and ground, and the specific implementation steps are as follows:
[0033] Step 1: Acquire aerial image sequences of the target area and generate a global DEM and orthophoto. Use visual sensors such as ground observation cameras mounted on UAVs or other aircraft to acquire aerial image sequences of the target area in advance. Complete image alignment by extracting and matching feature points between images. Then, use common 3D reconstruction methods such as multi-view solid geometry to generate dense point clouds and textures. Finally, generate a global DEM and orthophoto of the target area through coordinate system projection transformation.
[0034] Step 2: Divide the global DEM and orthophoto image into a grid. Combine the preset magnetic field azimuth angle and the intrinsic parameters of the ground mobile platform's visual sensors to generate a multi-view simulated frontal image for each grid. For example... Figure 2 As shown, the obtained global DEM and orthophoto image are divided into [various parts] at fixed intervals. The algorithm generates a DEM and a local orthophoto image for each grid cell, and assigns a number to each grid cell: {1, 2, 3, …,}. The ground mobile platform is equipped with... indivual( ≥1) Identical vision sensors with the same focal length, principal point, and other intrinsic parameters, but with different mounting azimuths and positions on the platform. The center point of each grid is then used as the simulation position. The magnetic field azimuth angle at 5-degree intervals is used as the simulation attitude. The simulation uses the intrinsic parameters of the visual sensor, i.e., {0°, 5°, 10°, ..., 350°, 355°}, as the intrinsic parameters of the visual sensor. Within each grid, the local DEM and orthophoto are used for rendering, resulting in N (N=72 here based on a 5° setting) simulated front view images of different simulated poses. The entire target area is then obtained. A simulated front view image.
[0035] Step 3: The ground mobile platform passes through A visual sensor acquires frontal view images of the ground from different orientations, and uses magnetic field azimuth angles to filter simulated frontal view images similar to the ground frontal view images; such as Figure 3 As shown, equipped with indivual( ≥1) The ground-based mobile platform with a vision sensor enters the target area, wherein the vision sensor directly in front of the ground-based mobile platform is used as the standard vision sensor. When =1), its position and attitude are taken as the position and attitude of the ground mobile platform. Other visual sensors have certain azimuth angle differences from the standard visual sensor, and the pose extrinsic parameters between other visual sensors and the standard visual sensor are fixed. These pose extrinsic parameter matrices are denoted as... , which indicates the first The position and orientation transformation matrix of a visual sensor compared to a standard visual sensor, when hour, The system uses an identity matrix. Multiple visual sensors acquire real-time frontal images of the ground from different orientations. A magnetic field azimuth sensor mounted on a ground-based mobile platform is used to obtain the magnetic field azimuth of the standard visual sensor. The magnetic field azimuth of each visual sensor is then calculated using the pose extrinsic parameters between the other visual sensors and the standard visual sensor. Finally, based on the magnetic field azimuth of each visual sensor, the set of N=72 simulated frontal images of the target area that most closely matches the simulated poses is found. A simulated front view image will be obtained, and the entire ground mobile platform will receive... Groups of different azimuth angles A simulated front view image.
[0036] Step 4: Extract and match feature points between the similar ground front view images and the simulated front view images. Determine the grid containing the simulated front view image with the most matched feature points as the optimal grid, and obtain the 3D coordinates of the 2D feature points in the optimal grid DEM. Then, use the ground front view images at different azimuth angles obtained from multiple vision sensors in Step 3 and their corresponding coordinates from each vision sensor. Image feature points are extracted and matched from each simulated front view image. The grid number of the simulated image with the highest number of successfully matched feature points is selected as the potential grid number of the current vision sensor. At this time... A visual sensor will get Given a grid number, find the grid number with the largest number of occurrences. The grid in which the ground mobile platform is currently located. The target area DEM was determined to be the optimal matching grid DEM. (The last part is incomplete and likely refers to a separate process.) The frontal view of the ground acquired by each visual sensor and the grid The coordinates of the 2D feature points that were successfully matched in the simulation front view image are shown below. and ,in, For the first A set of 2D feature point pixel coordinates on a ground frontal image acquired by a visual sensor. , For it in the grid The corresponding 2D feature point pixel coordinates on the simulated front view image. For each vision sensor, based on its position on the grid... The position used when rendering the corresponding simulation image. and posture ,turn up A set of 3D point coordinates corresponding to the DEM image That is, the coordinates of 2D feature points on the front view image of the ground. The corresponding set of 3D point coordinates.
[0037] Step 5: Construct a system of equations based on the 2D feature point coordinates and corresponding 3D coordinates of the ground front view image corresponding to the optimal mesh, and solve for the pose matrix of the ground mobile platform; using the results obtained in Step 4... Coordinates of 2D feature points on the ground frontal image of each visual sensor and its corresponding set of 3D point coordinates Each pair of 2D-3D points can construct an equation as shown in formula (1):
[0038] (1)
[0039] In formula (1), express The pixel coordinates of a 2D point in the image. express The coordinates of the corresponding 3D point in the image. This represents the intrinsic parameter matrix for each vision sensor, with a size of [value missing]. , Represents the position and orientation matrix of each vision sensor Size is subscript This indicates retrieving the first 3 rows of the matrix. and A system of equations can be constructed from 2D-3D point pairs to calculate... The position and orientation matrices of each visual sensor At least 6 pairs of 2D-3D points are required for the calculation.
[0040] Then, the pose extrinsic parameters between the visual sensor and the standard visual sensor are used. The standard visual sensor is calculated according to formula (2). A matrix of potential locations and attitudes ,have A single vision sensor can calculate the performance of a standard vision sensor. A potential position and orientation matrix Then to indivual The position and orientation matrix of a unique standard visual sensor was determined using statistical methods. That is, the position and attitude matrix of the ground mobile platform. Statistical methods can be used, including mean, median, and random sampling consistency.
[0041] (2)
[0042] Step 6: Optimize the pose matrix, output the motion trajectory of the ground mobile platform, reconstruct the DEM based on the trajectory, and construct a constrained optimized motion trajectory by combining the DEM generated from the aerial image; during the detection process of the ground mobile platform in the target area, at fixed time intervals... The position and orientation of the standard visual sensor are acquired in real time at each fixed moment through steps 3-5. That is, the position and attitude matrix of the ground mobile platform. Ultimately, the motion trajectory of the ground mobile platform is generated. .
[0043] like Figure 4 As shown, after obtaining the trajectory points of the ground mobile platform at least three times in step 6, the ground front view image sequence at these times is acquired using a standard vision sensor. The position and pose matrix corresponding to each front view image The dense point cloud and texture of this trajectory are generated using common 3D reconstruction methods such as multi-view solid geometry. Then, a DEM of this trajectory is further generated through coordinate system projection transformation, denoted as... The local DEMs of the target area corresponding to all the grids traversed during the journey are stitched together and denoted as... ,use and We construct constraint equations to further optimize the position and pose of the standard vision sensor. Let... This indicates that a certain 2D feature point on the front view image of the ground at the current moment is in The corresponding 3D point coordinates, In construction It has already been associated with the 2D feature points of the image. Let This indicates that these 2D feature points are in The corresponding 3D point coordinates, This involves associating the data with 2D feature points on the image in step 5. A total of [number] points can be found. right and ,make For the optimization matrix that needs to be optimized, the optimization equation shown in formula (3) can be constructed for optimization:
[0044] (3)
[0045] The optimization matrix is obtained through calculation. Then, the position and attitude matrix of the ground mobile platform in the target area at the current moment is obtained by formula (4). Perform optimization and output the optimized latest position and pose. .
[0046] (4)
[0047] Repeat steps 3 to 6 above to acquire and update the trajectory of the ground mobile platform in real time at certain time intervals.
[0048] On the other hand, the present invention provides a pose estimation device based on air-to-ground cross-view image matching, which includes modules capable of implementing the steps of the aforementioned method, specifically including:
[0049] The acquisition unit is used to acquire aerial image sequences of the target area and generate a global digital elevation model (DEM) and orthophotos.
[0050] The simulation unit is used to divide the global DEM and orthophoto into grids, and generate multi-view simulated front view images for each grid by combining the preset magnetic field azimuth angle and the visual sensor intrinsic parameters of the ground mobile platform.
[0051] The selection unit is used by the ground mobile platform to acquire ground front view images from different orientations and to filter simulated front view images that are similar to the ground front view images using the magnetic field azimuth angle.
[0052] The matching unit is used to extract and match feature points between the similar ground front view image and the simulation front view image, and determine the grid where the simulation front view image with the most matched feature points is located as the optimal grid.
[0053] The calculation unit is used to calculate the pose matrix of the ground mobile platform based on the coordinates of the feature points corresponding to the optimal grid.
[0054] The output unit is used to optimize the pose matrix and output the motion trajectory of the ground mobile platform.
[0055] Thirdly, the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned pose estimation method based on air-to-ground cross-view image matching.
[0056] Fourthly, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the aforementioned pose estimation method based on cross-view image matching between air and ground.
[0057] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A pose estimation method based on cross-view image matching between air and ground, characterized in that, The method includes: Step 1: Acquire aerial image sequences of the target area and generate a global digital elevation model (DEM) and orthophotos; Step 2: Divide the global DEM and orthophoto image into grids, and generate multi-view simulated front view images for each grid by combining the preset magnetic field azimuth angle and the intrinsic parameters of the visual sensor of the ground mobile platform. Step 3: The ground mobile platform acquires front view images of the ground from different directions, and uses the magnetic field azimuth angle to filter out simulated front view images that are similar to the ground front view images; Step 4: Extract and match feature points between the similar ground front view images and the simulated front view images, and determine the grid where the simulated front view image with the most matched feature points is located as the optimal grid. Step 5: Calculate the pose matrix of the ground mobile platform based on the coordinates of the feature points corresponding to the optimal mesh; Step 6: Optimize the pose matrix and output the motion trajectory of the ground mobile platform.
2. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 1 includes: after aligning the aerial image sequence, generating a dense point cloud using a 3D reconstruction method, and generating a global DEM and orthophoto image of the target area based on coordinate system projection.
3. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 2 includes: The obtained global DEM and orthophoto image are divided into sections at fixed intervals. Each grid contains its corresponding local DEM and orthophoto image; The center point of each grid is used as the simulation position, the preset multi-directional magnetic field azimuth angle is used as the simulation attitude, and the intrinsic parameters of the ground mobile platform's visual sensor are used as the simulation's intrinsic parameters of the visual sensor. Within each grid, local DEM and orthophotos are used for rendering to obtain N simulated front view images with different attitudes.
4. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 3 includes: Utilizing ground mobile platforms Each visual sensor acquires frontal images of the ground from different directions in real time. Among them, the visual sensor in the front direction of the ground mobile platform is regarded as the standard visual sensor, and the pose extrinsic parameters of the other visual sensors are fixed with respect to the standard visual sensor. Based on the magnetic field azimuth angle of each visual sensor, find the set of N simulated poses that are closest to the simulated front view images of the target area. A simulated front view image will be obtained, and the entire ground mobile platform will receive... Groups of different azimuth angles A simulated front view image.
5. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 4 includes: Frontal images of the ground at different azimuth angles acquired by multiple visual sensors and their respective corresponding images for each visual sensor. The simulation front view image is used to extract and match image feature points. The grid number of the simulation image with the most successful feature point matching is selected as the potential grid number of the current vision sensor. The grid where the ground mobile platform is located is determined based on the matching results of multiple vision sensors. Based on the local DEM of the grid, find the 3D feature point coordinates corresponding to the 2D feature point coordinates of the matching simulation image.
6. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 5 includes: Obtained using step 4 The coordinates of 2D feature points on the ground frontal image of each visual sensor and their corresponding set of 3D feature point coordinates are calculated. The position and attitude matrices of each visual sensor; The standard visual sensor is obtained by calculating the pose extrinsic parameters between the visual sensor and the standard visual sensor. A potential position and pose matrix, then... The potential position and attitude matrix is determined statistically to be the unique position and attitude matrix of the standard vision sensor, i.e., the position and attitude matrix of the ground mobile platform.
7. The pose estimation method based on cross-view image matching between air and ground as described in claim 1, characterized in that, Step 6 includes: During the detection process in the target area, the ground mobile platform acquires the position and attitude of the standard visual sensor at each fixed time interval in real time, that is, the position and attitude matrix of the ground mobile platform, and finally generates the motion trajectory of the ground mobile platform. Using a sequence of ground front view images acquired by a standard vision sensor and the position and attitude matrix corresponding to each ground front view image, a dense point cloud and texture of the motion trajectory of the ground mobile platform are generated through 3D reconstruction methods. Then, a digital elevation model of this trajectory is generated through coordinate system projection transformation, denoted as […]. The local DEMs of the target area corresponding to all the grids traversed during the journey are stitched together, denoted as ,use and Constraint equations are constructed to optimize the position and orientation of the standard vision sensor, and the optimized motion trajectory of the ground mobile platform is output.
8. A pose estimation device based on cross-view image matching between air and ground, characterized in that, include: The acquisition unit is used to acquire aerial image sequences of the target area and generate a global digital elevation model (DEM) and orthophotos. The simulation unit is used to divide the global DEM and orthophoto into grids, and generate multi-view simulated front view images for each grid by combining the preset magnetic field azimuth angle and the visual sensor intrinsic parameters of the ground mobile platform. The selection unit is used by the ground mobile platform to acquire ground front view images from different orientations and to filter simulated front view images that are similar to the ground front view images using the magnetic field azimuth angle. The matching unit is used to extract and match feature points between the similar ground front view image and the simulation front view image, and determine the grid where the simulation front view image with the most matched feature points is located as the optimal grid. The calculation unit is used to calculate the pose matrix of the ground mobile platform based on the coordinates of the feature points corresponding to the optimal grid. The output unit is used to optimize the pose matrix and output the motion trajectory of the ground mobile platform.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the pose estimation method based on cross-view image matching between air and ground as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, enable the processor to implement the pose estimation method based on cross-view image matching between air and ground as described in any one of claims 1-7.
Citation Information
Patent Citations
Ground-to-unmanned aerial vehicle laser point cloud cross-view relocation method based on semantic map
CN115797422A
Air-ground heterogeneous collaborative mapping method and device, equipment and storage medium
CN117191005A