Three-dimensional reconstruction method and terminal device

By using planning and global bundling adjustment and optimization technology for UAVs and ground equipment, the problem of insufficient coordinate system alignment accuracy in the fusion of aerial and terrestrial photogrammetry was solved, and high-precision 3D model reconstruction was achieved.

CN120747388BActive Publication Date: 2026-02-06SHENZHEN XGRIDS-INNOVATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511271430.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-02-06
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

In existing technologies, the coordinate system alignment accuracy of aerial and terrestrial photogrammetry fusion is insufficient, which makes the fused area prone to geometric distortion and misalignment, making it difficult to achieve high-precision and large-scale 3D model reconstruction.

Method used

By planning the flight path of the UAV and the ground scanning plan, we ensure that there are overlapping areas between the images captured by the UAV and the images captured by the ground scanning equipment. We also use global binding adjustment and optimization technology to combine aerial sparse point clouds and ground sparse point clouds to perform image matching and pose information optimization, and use control point constraints to improve the coordinate system alignment accuracy.

Benefits of technology

It improves the coordinate system alignment accuracy during air-to-ground fusion, reduces geometric distortion and misalignment in the fusion area, and achieves high-precision 3D model reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747388B_ABST
    Figure CN120747388B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of three-dimensional reconstruction, and discloses a three-dimensional reconstruction method and a terminal device, which comprise the following steps: unmanned aerial vehicle route planning and ground terminal scanning planning are performed on a scanning area; aerial images and ground images are acquired; an aerial model is reconstructed according to the aerial images to obtain aerial sparse point clouds and first pose information, and a ground model is reconstructed according to the ground images to obtain ground sparse point clouds and second pose information; target aerial images and target ground images in the aerial images and the ground images that have an overlapping area are matched to obtain relative pose information between the target aerial images and the target ground images; and the pose information and the sparse point positions of each image in the target aerial images and the target ground images that have the overlapping area are adjusted and optimized by using global bundling with control point constraints according to the first pose information, the second pose information and the relative pose information, so that a fusion model is obtained. The application improves the air-ground coordinate alignment accuracy of three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of three-dimensional reconstruction, in particular to a three-dimensional reconstruction method and a terminal device. BACKGROUND

[0002] Three-dimensional reconstruction technology refers to a process of recovering a three-dimensional model and spatial structure of an object or a scene by extracting geometric information from a single or multiple two-dimensional images or point cloud data. The technology is one of the core research directions in the fields of computer vision, digital photogrammetry and graphics, and has been widely applied in many fields such as robot navigation, unmanned driving, virtual reality, cultural relic digitization, industrial detection and medical imaging.

[0003] The three-dimensional reconstruction technology used in the fields of surveying and mapping, planning and the like mainly includes aerial photogrammetry, ground photogrammetry and fusion of aerial photogrammetry and ground photogrammetry.

[0004] Aerial photogrammetry usually adopts a drone (single lens or multiple lenses) plus a gimbal, and carries a real-time dynamic positioning (RTK) device to ensure stability to obtain a large range of overhead images. The advantage of aerial photogrammetry is that the coverage range is large and the image acquisition efficiency is high, and the reconstruction is relatively stable, but the detail accuracy of building facades and the like is not enough due to the long shooting distance. RTK can avoid the problem of large traditional GPS positioning error, and provides higher absolute accuracy for aerial models.

[0005] Ground photogrammetry usually uses a laser device to carry a panoramic camera, a handheld camera or a vehicle-mounted camera to obtain street view or close-up images, which can capture high-precision facade details, but the coverage range is limited and the operation is relatively cumbersome. Among them, the laser can provide high-precision three-dimensional geometric information (point cloud), and the camera provides real texture color information, and the two complement each other.

[0006] In the fusion mode of aerial photogrammetry and ground photogrammetry, since the ground data and the aerial data use different coordinates, the fusion of the coordinate systems of the two is difficult, and the alignment accuracy of the coordinate systems is insufficient, which leads to misalignment of the model splicing. Moreover, the ground data and the aerial data are directly fused based on all images as a whole, and geometric distortion and misalignment are prone to occur in the fusion area.

[0007] Therefore, a three-dimensional reconstruction method capable of improving the coordinate alignment accuracy of the fusion of aerial photogrammetry and ground photogrammetry, realizing accurate fusion of aerial and ground images, and obtaining a three-dimensional model with high precision and a large range is needed. SUMMARY

[0008] In view of the above problems, the embodiments of the present application provide a three-dimensional reconstruction method and a terminal device, which are used to solve the problems of insufficient coordinate system alignment accuracy and easy geometric distortion and misplacement in the fusion area in the prior art.

[0009] According to an aspect of the embodiments of the present application, a three-dimensional reconstruction method is provided, which comprises:

[0010] The scanning area is subjected to unmanned aerial vehicle flight path planning and ground end scanning planning to obtain an unmanned aerial vehicle shooting flight path and a ground end scanning path, wherein the unmanned aerial vehicle shooting flight path comprises an unmanned aerial vehicle take-off and landing point and an unmanned aerial vehicle shooting path, and the ground end scanning path comprises a scanning path of a ground scanning device, and the scanning path of the ground scanning device passes through the unmanned aerial vehicle take-off and landing point;

[0011] A plurality of aerial images with different perspectives obtained by continuous aerial shooting of the unmanned aerial vehicle according to the unmanned aerial vehicle shooting flight path, and a plurality of ground images obtained by ground shooting of the ground scanning device according to the ground end scanning path are acquired, wherein the coordinate system of the ground images is the same as that of the aerial images;

[0012] An aerial model is reconstructed according to the plurality of aerial images to obtain an aerial sparse point cloud and first pose information of each aerial image, and a ground model is reconstructed according to the plurality of ground images to obtain a ground sparse point cloud and second pose information of each ground image;

[0013] Based on the aerial sparse point cloud and the ground sparse point cloud, target aerial images and target ground images in the plurality of aerial images and the plurality of ground images that have an overlapping area are matched to obtain relative pose information between the target aerial images and the target ground images;

[0014] According to the first pose information, the second pose information and the relative pose information, the pose information and the sparse point position of each image in the target aerial images and the target ground images that have the overlapping area are adjusted and optimized by using global bundle adjustment with control point constraints to obtain a fusion model.

[0015] Optionally, the aerial model reconstruction according to the plurality of aerial images to obtain the aerial sparse point cloud and the first pose information of each aerial image comprises:

[0016] Position and attitude POS data of each aerial image are acquired;

[0017] Feature points in each aerial image and feature descriptors of the feature points are determined, wherein the feature points comprise corner points, edge intersection points and centers of high-contrast regions;

[0018] determining a plurality of pairs of aerial image pairs from the plurality of aerial images according to the POS data, wherein each pair of the aerial image pairs comprises two aerial images having an overlapping area;

[0019] determining, for each pair of the aerial image pairs, corresponding feature points in the two aerial images to obtain a plurality of pairs of matching feature points composed of the corresponding feature points;

[0020] determining relative pose information between the two aerial images to which the matching feature points belong according to the matching feature points, determining initial first pose information of each of the aerial images and an initial aerial sparse point cloud according to the relative pose information between the two aerial images, and optimizing the initial first pose information and the initial aerial sparse point cloud through global bundle adjustment to obtain the first pose information and the aerial sparse point cloud.

[0021] Optionally, the ground scanning device is provided with a laser radar and a camera, and the ground images are images captured by the camera.

[0022] reconstructing a ground model according to the plurality of ground images to obtain a ground sparse point cloud and second pose information of each ground image, comprising:

[0023] determining a laser track according to point cloud data collected by the laser radar, wherein the laser track comprises a pose sequence of the laser radar;

[0024] transferring the laser track to an image to obtain an image track by using pre-labeled extrinsic parameters between the laser radar and the camera, wherein the image track comprises initial estimated pose information of each of the ground images.

[0025] performing global bundle adjustment optimization on the pose information of the ground images through feature point extraction and matching based on the image track to obtain the second pose information and the ground sparse point cloud.

[0026] Optionally, based on the aerial sparse point cloud and the ground sparse point cloud, matching target aerial images and target ground images having an overlapping area in the plurality of aerial images and the plurality of ground images to obtain relative pose information between the target aerial images and the target ground images, comprising:

[0027] determining target aerial images and target ground images having an overlapping area according to the first pose information of the plurality of target aerial images and the second pose information of the plurality of target ground images.

[0028] determining same-named sparse points in the target aerial image and the target ground image located in the existing overlapping area based on the aerial sparse point cloud and the ground sparse point cloud, to obtain a plurality of pairs of sparse point pairs composed of same-named sparse points, wherein each pair of sparse point pairs includes an aerial sparse point and a ground sparse point;

[0029] For each pair of sparse point pairs, determining a target aerial image corresponding to the aerial sparse point in the pair of sparse points from a plurality of the aerial images, to obtain a plurality of target aerial images, and determining a target ground image corresponding to the ground sparse point in the pair of sparse points from a plurality of the ground images, to obtain a plurality of target ground images;

[0030] Combining the plurality of target aerial images and the plurality of target ground images two by two to obtain a plurality of pairs of aerial-ground image pairs, wherein each pair of aerial-ground image pairs includes a target aerial image and a target ground image;

[0031] For each pair of aerial-ground image pairs, performing a plurality of scale zooms on the target aerial image therein to obtain a plurality of aerial zoom images of different scales, and performing a plurality of scale zooms on the target ground image therein to obtain a plurality of ground zoom images of different scales;

[0032] Determining aerial zoom images and ground zoom images of the same scale from the plurality of aerial zoom images and the plurality of ground zoom images;

[0033] Performing feature point extraction and matching based on view angle weight on the aerial zoom images and the ground zoom images of the same scale to obtain relative pose information between the target aerial image and the target ground image, wherein the view angle weight of a feature point extracted from an aerial zoom image with a lower shooting height is greater than the view angle weight of a feature point extracted from an aerial zoom image with a higher shooting height.

[0034] Optionally, the determining same-named sparse points in the target aerial image and the target ground image located in the existing overlapping area based on the aerial sparse point cloud and the ground sparse point cloud comprises:

[0035] Based on the aerial sparse point cloud, obtaining three-dimensional coordinates of the aerial sparse points in the target aerial image located in the existing overlapping area, and based on the ground sparse point cloud, obtaining three-dimensional coordinates of the ground sparse points in the target ground image located in the existing overlapping area;

[0036] Determining the aerial sparse points and the ground sparse points with similar three-dimensional coordinates as the same-named sparse points.

[0037] Optionally, the adjusting and optimizing the pose information and the sparse point position of each of the target aerial image and the target ground image in the overlapping region comprises:

[0038] determining the control points from all the sparse points included in the aerial sparse point cloud and the ground sparse point cloud;

[0039] adjusting and optimizing the pose information and the sparse point position of each of the target aerial image and the target ground image in the overlapping region by global bundle adjustment with the position of the control points as constraints according to the first pose information, the second pose information and the relative pose information, to obtain a fusion model, wherein the position of the control points in the fusion model remains unchanged before optimization.

[0040] Optionally, the determining the control points from all the sparse points included in the aerial sparse point cloud and the ground sparse point cloud comprises:

[0041] obtaining the number of images corresponding to all the sparse points included in the aerial sparse point cloud and the ground sparse point cloud;

[0042] determining the sparse points as the control points, which have a number of corresponding images greater than a first preset value and a number of images in which the overlapping region does not exist less than or equal to a second preset value.

[0043] Optionally, the fusion model comprises the optimized plurality of sparse points and positions thereof, and the optimized pose information of each image.

[0044] Optionally, the unmanned aerial vehicle route planning and the ground end scanning planning for the scanning region to obtain the unmanned aerial vehicle shooting route and the ground end scanning path comprises:

[0045] dividing the scanning region into a plurality of grids;

[0046] determining the vertex of each grid as the unmanned aerial vehicle take-off and landing point;

[0047] determining the starting shooting height and the ending shooting height of the unmanned aerial vehicle to obtain the unmanned aerial vehicle shooting path, wherein the unmanned aerial vehicle shooting path is a path in the vertical direction, the starting point of which is a position in the vertical direction at a distance of the starting shooting height from the unmanned aerial vehicle take-off and landing point, and the ending point of which is a position in the vertical direction at a distance of the ending shooting height from the unmanned aerial vehicle take-off and landing point, and the difference between the starting shooting height and the shooting height of the ground scanning device is within a preset range, so that the target aerial image and the target ground image in which the overlapping region exists are included in the plurality of aerial images and the plurality of ground images.

[0048] determine a scanning path of the ground scanning device based on the UAV landing point, so that the scanning path of the ground scanning device passes through the UAV landing point.

[0049] According to another aspect of the embodiments of the present application, a terminal device is provided, which comprises a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the three-dimensional reconstruction method as described above.

[0050] The embodiments of the present application make the scanning path of the ground scanning device pass through the UAV landing point through UAV flight path planning and ground end scanning planning for the scanning area, so that there is an overlapping area between the image of the landing part in the aerial image taken by the UAV and the ground image taken by the ground scanning device, feature point extraction and matching in the overlapping area between the aerial and ground images can be performed, the image in the overlapping area is added to the global bundle adjustment optimization process, the alignment accuracy of the coordinate system when fusing the aerial and ground images is improved, the consistency of the aerial and ground poses is improved, and compared with simple alignment for only two scenes of aerial and ground images, higher accuracy is achieved. Moreover, the control point constraint is used when performing global bundle adjustment optimization, the reconstruction stability is improved, the image without overlapping area on the ground will not be optimized incorrectly, and similarly, the image without overlapping area in the air will not be incorrect, the geometric distortion and misplacement problems in the fusion area are reduced, and the accuracy is further improved.

[0051] The above description is only a summary of the technical solutions of the embodiments of the present application, in order to more clearly understand the technical means of the embodiments of the present application, the embodiments of the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS

[0052] The accompanying drawings are included to provide a further understanding of the embodiments of the present application, and are incorporated herein and constitute a part of the detailed description. In the drawings:

[0053] Figure 1 An application scenario of the embodiments of the present application is shown;

[0054] Figure 2 A flowchart of the three-dimensional reconstruction method provided by the embodiments of the present application is shown;

[0055] Figure 3 A grid map of the divided scanning area is shown;

[0056] Figure 4 A schematic diagram of the UAV shooting path is shown;

[0057] Figure 5A structure schematic diagram of a terminal device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0058] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.

[0059] The present application provides a three-dimensional reconstruction method based on air-ground fusion for three-dimensional reconstruction. Figure 1 An application scenario schematic diagram of an embodiment of the present application is shown. As shown in the diagram, in a three-dimensional scene, a plurality of aerial images of different perspectives obtained by continuously shooting in the air by a UAV 1 according to a UAV shooting route, and a plurality of ground images obtained by shooting on the ground by a ground scanning device 2 according to a ground end scanning path, wherein the ground scanning device 2 can be provided with a laser radar and a camera, the laser radar collects laser point cloud data of the three-dimensional scene, and the camera acquires ground images of the three-dimensional scene. The ground scanning device 2 can be a handheld device or a non-handheld device. If the ground scanning device 2 is a handheld device, the user holds the ground scanning device 2 to move in the three-dimensional scene to collect laser point cloud data and ground images of the three-dimensional scene. If the ground scanning device 2 is a non-handheld device, the ground scanning device 2 can move autonomously in the three-dimensional scene to collect laser point cloud data and ground images of the three-dimensional scene.

[0060] A terminal device 3 is in communication connection with the UAV 1 and the ground scanning device 2 respectively, receives aerial images shot by the UAV 1, and laser point cloud data and ground images scanned by the ground scanning device 2, and obtains a three-dimensional reconstruction result (fusion model) after three-dimensional reconstruction based on the above data. The terminal device 3 can directly display the three-dimensional reconstruction result, or send the three-dimensional reconstruction result to other display devices for display. The terminal device 3 can be a mobile phone, a tablet computer, a notebook computer, a desktop computer, a server, etc. The terminal device 3 and the ground scanning device 2 can be connected through wired communication or wireless communication.

[0061] The UAV 1 can be mounted with a tilt photography camera, such as a five-lens, three-lens or double-lens tilt photography camera, and be mounted with an RTK, and has a Position and Orientation System (POS) information recording function. The ground scanning device 2 can also be mounted with an RTK to record image pose information. The tilt photography of the UAV refers to an aerial photography technology for collecting images of ground objects from vertical and multiple tilt angles (usually five directions of front, back, left and right) through a tilt photography camera mounted on the UAV.

[0062] In some other scenarios, the laser radar and the camera can also be separate devices rather than integrated in the ground scanning device 2, and they send the data collected respectively to the terminal device 3, and the terminal device 3 performs data fusion processing to obtain the three-dimensional reconstruction result.

[0063] The three-dimensional scene of the embodiments of the present application mainly refers to a large-scale outdoor scene.

[0064] Figure 2 A flowchart of a three-dimensional reconstruction method provided by the embodiments of the present application is shown, and the method is performed by a terminal device. The terminal device can be the terminal device 3 in Figure 1 As shown in Figure 2 , the method includes the following steps:

[0065] S110, performing unmanned aerial vehicle route planning and ground end scanning planning on the scanning area to obtain an unmanned aerial vehicle shooting route and a ground end scanning path.

[0066] The unmanned aerial vehicle shooting route includes unmanned aerial vehicle take-off and landing points and an unmanned aerial vehicle shooting path, and the ground end scanning path includes a scanning path of the ground scanning device, and the scanning path of the ground scanning device passes through the unmanned aerial vehicle take-off and landing points.

[0067] S110 can specifically include the following steps:

[0068] S111, dividing the scanning area into a plurality of grids;

[0069] Figure 3 A grid diagram of the divided scanning area is shown. As shown in Figure 3 , each grid can be a square grid with the same side length, for example, 30 m. Setting a larger grid side length can make a certain number of sparse points covered in the captured image, which can improve the matching accuracy in the subsequent feature point matching process, thereby improving the alignment accuracy of the air-ground data.

[0070] S112, determining the vertex of each grid as an unmanned aerial vehicle take-off and landing point;

[0071] As shown in Figure 3 , each grid vertex shown by the black dot in the figure is determined as an unmanned aerial vehicle take-off and landing point P, and there are 36 unmanned aerial vehicle take-off and landing points P in the figure.

[0072] S113, determining a starting shooting height and an ending shooting height of the unmanned aerial vehicle to obtain an unmanned aerial vehicle shooting path R, wherein the unmanned aerial vehicle shooting path R is a path in the vertical direction, the starting point H1 is a position in the vertical direction at a distance from the unmanned aerial vehicle take-off and landing point P to the starting shooting height, and the ending point H2 is a position in the vertical direction at a distance from the unmanned aerial vehicle take-off and landing point P to the ending shooting height;

[0073] The difference between the initial shooting height and the shooting height of the ground scanning device is set within a preset range to ensure that multiple aerial images captured by the drone and multiple ground images captured by the ground scanning device have overlapping areas. The preset range can be set according to the actual application scenario, for example, [-10cm, 10cm], or other ranges. For example, the initial shooting height can be set to 1.5m. Based on the average height of a person holding the ground scanning device (1.7m), the shooting height of the ground scanning device is approximately 1.5m. Setting the drone's initial shooting height to 1.5m ensures that the initial aerial images captured by the drone and the ground images captured by the ground scanning device have the same field of view, resulting in an overlapping area for subsequent fusion. If shooting begins immediately after drone takeoff, or if the initial shooting height is set too low, the initial images captured by the drone will be almost entirely ground images rather than building images, which are of little use for 3D reconstruction and lead to wasted resources. The aforementioned "overlapping area" refers to a region in the same physical space that appears in both an aerial image and a ground image; such a region is called an overlapping area.

[0074] The ending shooting height can be determined based on the usual oblique photography height, for example, 80-120 meters in urban areas, and 150-200 meters in rural areas or open areas.

[0075] Figure 4 A schematic diagram of the drone's shooting path is shown. (For example...) Figure 4 As shown in the diagram, the arrows indicate the drone's flight direction, which is vertical ascent. For example, at each drone takeoff and landing point P, the drone takes off from point P, and upon reaching the starting point H1, its onboard camera begins taking pictures. It then continues to ascend vertically and take pictures continuously, ceasing filming upon reaching the endpoint H2, thus completing the filming work at that drone takeoff and landing point P. The thick line R represents the drone's filming path.

[0076] S114, determine the scanning path of the ground scanning equipment based on the UAV take-off and landing point, so that the scanning path of the ground scanning equipment passes through the UAV take-off and landing point.

[0077] The scanning path of the ground scanning equipment passing through the UAV take-off and landing point means that the points passed by the ground scanning equipment include points with the same coordinates as the UAV take-off and landing point, or points with a difference of no more than a preset threshold (the difference is not large), for example, the difference is only within tens of centimeters.

[0078] In the above process, the scanning area can be automatically designed using traditional oblique photogrammetry flight path planning methods, ensuring an overlap rate greater than 80%. The overlap rate of an oblique photogrammetry flight path refers to the proportion of the same area between adjacent images. It ensures that the captured area is completely and without omission, and provides sufficient viewpoint information for subsequent 3D model reconstruction. The overlap rate is usually measured in two directions: forward overlap rate and lateral overlap rate. The aforementioned overlap rate greater than 80% refers to the forward overlap rate.

[0079] If there is no overlap between aerial and ground images, it will be difficult to match aerial images from a purely aerial perspective and ground images from a ground perspective due to the difference in viewpoints. This application, by performing the aforementioned UAV flight path planning and ground-end scanning planning on the scanning area, creates an overlapping area between the take-off and landing images in the aerial images captured by the UAV and the ground images captured by the ground scanning equipment. This is equivalent to creating a common area in space, which is simultaneously covered by ground laser point clouds / images and UAV images. This enables the establishment of a connection between the two types of images, allowing for feature point extraction and matching between the aerial and ground images. This provides a reliable control basis for high-precision matching of aerial and ground models, improving the alignment accuracy of aerial and ground image fusion.

[0080] S120: Acquire multiple aerial images from different perspectives obtained by the UAV continuously taking aerial photos along the UAV's shooting route, and multiple ground images obtained by the ground scanning device taking ground photos along the ground end scanning path.

[0081] After planning the drone's shooting route on the S110, the drone can be manually controlled to take off at each take-off and landing point, and then controlled to shoot according to the planned drone shooting route. The images captured by the drone include multiple images from different perspectives. The perspectives of the images are different depending on the take-off and landing point and the shooting height.

[0082] On the ground, operators can use handheld ground scanning equipment to scan along the ground scanning path, or send the ground scanning path to a ground scanning device with autonomous mobility, which will then scan along the ground scanning path.

[0083] After the drone and ground scanning equipment have completed their shooting, the terminal equipment acquires multiple aerial images from different perspectives obtained by the drone following its shooting route, and multiple ground images obtained by the ground scanning equipment following its ground scanning path.

[0084] The coordinate system of the ground image is the same as that of the aerial image. For example, if the UAV uses the WGS84 coordinate system to take pictures, the ground scanning device also uses the WGS84 coordinate system to take pictures; if the UAV uses the CGCS2000 coordinate system to take pictures, the ground scanning device also uses the CGCS2000 coordinate system to take pictures. In summary, the coordinate system of the ground image can be consistent with the coordinate system of the aerial image. By unifying the coordinate system, the spatial reference is unified from the data acquisition source, avoiding complex and serious precision loss of coordinate system conversion work in the later stage, and ensuring the consistency of the aerial model and the ground model in the spatial reference.

[0085] In S130, the aerial model is reconstructed according to the plurality of aerial images to obtain an aerial sparse point cloud and first pose information of each aerial image, and the ground model is reconstructed according to the plurality of ground images to obtain a ground sparse point cloud and second pose information of each ground image.

[0086] This step respectively performs aerial oblique photography reconstruction and ground image reconstruction.

[0087] In the aerial model reconstruction according to the plurality of aerial images, the aerial sparse point cloud and the first pose information of each aerial image are obtained, including:

[0088] In S a1, position and attitude POS data of each aerial image are obtained.

[0089] The RTK carried by the UAV can provide initial positioning information, based on which the initial position (three-dimensional coordinates) and attitude, i.e., POS data, can be determined. The POS data provides an initial value for subsequent feature matching, which can greatly reduce the amount of calculation, avoid matching errors, and improve the absolute accuracy of the entire model.

[0090] In S a2, feature points and feature descriptors of the feature points in each aerial image are determined, wherein the feature points include corner points, edge intersection points, centers of high-contrast regions, etc.

[0091] This step is a feature extraction process, which can use Scale-invariant feature transform (SIFT), SURF or faster feature extraction algorithm combining FAST key point detection and BRIEF descriptor (Oriented FAST and Rotated BRIEF, ORB) to extract unique, stable and repeatable feature points from each image. By scanning the entire image, the corner points, edge intersection points, and centers of high-contrast regions are found. For each feature point found, a digital vector called a feature descriptor is calculated, which encodes the unique information of the pixel region around the point. Through the above process, each image gets a set of feature points and their corresponding descriptors.

[0092] Step a3, determining a plurality of pairs of aerial image pairs from the plurality of aerial images according to the POS data;

[0093] Each pair of aerial image pairs includes two aerial images, and the two aerial images have an overlapping region. The "overlapping region" here refers to a region in the same physical space that appears in two or more aerial images. Such a region is the overlapping region here.

[0094] This step determines a plurality of pairs of aerial image pairs with overlapping regions from a plurality of aerial images based on the pose information (POS data) of the images. Such images usually contain feature points belonging to the same physical point in a three-dimensional scene because they have overlapping regions.

[0095] Step a4, for each pair of aerial image pairs, determining corresponding feature points in the two aerial images to obtain a plurality of pairs of matching feature points composed of corresponding feature points;

[0096] Corresponding feature points refer to pixel points appearing in different images that represent the same physical point in the real world.

[0097] Steps a3 and a4 are feature matching processes, aiming to find the same physical point captured in different images. Due to the difference in shooting angle, lighting, scale and occlusion, the same point may look very different in two images. By using POS data (initial position and attitude) to intelligently select the image pairs most likely to have overlapping areas for comparison, the amount of calculation is reduced. Then, for the aerial image pairs with overlapping areas (such as Image A and Image B), the descriptor of a feature point in Image A is taken out, and then the most similar one (which can be calculated by Euclidean distance or Hamming distance, usually the nearest neighbor) is found among all the descriptors in Image B. The most similar point in Image B and the feature point in Image A are determined as the same name feature point. The initial matching may contain errors (false matches), and robust estimation algorithms such as the Random Sample Consensus (RANSAC) method can be used. A small number of matching point pairs are randomly selected, a hypothesis model of a Fundamental Matrix or a Homography is calculated, and then all other matching point pairs are tested with this model to see how many points conform to this model (i.e. "inliers"). Repeat this process several times, finally adopt the model with the most "inliers", and eliminate all "outliers" (false matches) that do not conform to the model. Finally, a purified and accurate matching feature point set is obtained.

[0098] Step a5, according to the matching feature points, determine the relative pose information between the two aerial images to which the matching feature points belong, according to the relative pose information between the two aerial images, determine the initial first pose information and the initial aerial sparse point cloud of each aerial image, and optimize the initial first pose information and the initial aerial sparse point cloud through global bundle adjustment, to obtain the first pose information and the aerial sparse point cloud.

[0099] This step includes the processes of incremental reconstruction (Incremental SfM), determination of initial rotation and translation, and global bundle adjustment. Global Bundle Adjustment (Global BA) is a core technology in computer vision, photogrammetry and robotics (especially in the fields of SLAM and SfM), which is a nonlinear optimization process that can simultaneously optimize all camera parameters (pose - position and orientation) and all three-dimensional point coordinates in the scene. In the global bundle adjustment optimization process, optimization is achieved by minimizing the difference (called re-projection error) between the observed data (usually two-dimensional feature points in the image) and the projected points predicted according to the current three-dimensional structure and camera parameters. Finally, the accurate position and orientation (pose information) of the camera when shooting each image, and the accurate coordinates of the three-dimensional points are obtained.

[0100] This step recovers the accurate pose information and sparse point cloud (sparse 3D point cloud) of each image from the matching points of 2D images. First, two images with the highest overlap, the most matching feature points and uniform distribution can be selected as the starting point of the calculation. At this time, their POS data is used as the initial value. Then, according to the matching feature points of the two images, the initial relative rotation and translation between them is calculated accurately, for example, the essential matrix can be solved by five-point method or other algorithms. Knowing the accurate relative position of the two images, for each pair of matching feature points, its coordinates (X, Y, Z) in 3D space can be calculated by triangulation principle, thereby obtaining the first batch of sparse points. Next, find the third image with enough (for example, more than 20) 2D matching points (i.e. the feature points obtained in step a2) of the known 3D points (i.e. the first batch of sparse points) of the current model. By using the PnP (Perspective-n-Point) algorithm or similar algorithm, the accurate camera position and pose of the third image are solved by using the corresponding relationship between 3D and 2D. The same process as the third image is used to match other new images with the existing model, triangulate new 3D points, add them to the sparse point cloud, and repeat the process to continuously register new images to the growing model and expand the sparse 3D point cloud until all images that meet the conditions are added. Finally, through global bundle adjustment, the pose information of all images and the coordinates of all sparse points are optimized to minimize the reprojection error, and the optimized first pose information and aerial sparse point cloud are obtained.

[0101] The above is the aerial oblique photography reconstruction process, and the final aerial model includes the first pose information and the aerial sparse point cloud. The matching information of the feature points between the aerial images is used in the reconstruction process.

[0102] For the reconstruction of the ground model, the high-precision and low-drift trajectory generated by laser SLAM can be used as a strong constraint for visual SLAM or visual SfM (structure from motion), thereby generating a model with both laser-level geometric precision and rich visual texture and feature information. Specifically, the ground images are images taken by the camera of the ground scanning device. According to multiple ground images, the ground model is reconstructed to obtain a ground sparse point cloud and second pose information of each ground image, including:

[0103] Step b1, determining a laser trajectory according to the point cloud data collected by the laser radar, wherein the laser trajectory includes a sequence of poses of the laser radar;

[0104] In this step, by processing the original point cloud sequence collected by the laser radar, for each frame of laser point cloud, feature points are extracted, and the relative motion (pose transformation) of the radar itself is estimated by calculating the matching relationship of the feature points between consecutive frames (using ICP, NDT, etc. algorithms). At the same time, the point clouds of these frames are registered into a global map, and loop detection and map optimization can be performed to correct the cumulative error. Finally, a set of high-precision laser radar pose sequences T_li (i=1, 2,..., N) is obtained, i.e. the laser trajectory. The trajectory is very accurate in the local coordinate system and has little drift.

[0105] Step b2, using the pre-marked extrinsic parameters between the laser radar and the camera, the laser trajectory is transferred to the image to obtain an image trajectory, wherein the image trajectory includes an initial estimated pose of each ground image.

[0106] In this step, the extrinsic parameters between the laser radar and the camera, which can be a 4x4 rigid transformation matrix representing the rotation and translation relationship of the radar coordinate system to the camera coordinate system, are used. For each laser pose T_li, the corresponding camera pose T_ci can be obtained by simple coordinate transformation:

[0107] T_ci = T_li * T_lc;

[0108] Wherein, T_lc is the extrinsic parameter between the laser radar and the camera mentioned above.

[0109] Finally, an initial, high-precision pose estimate is calculated for each ground image.

[0110] Step b3, based on the image trajectory, the pose information of the ground image is adjusted and optimized through feature point extraction and matching to obtain second pose information and a sparse ground point cloud.

[0111] In this step, visual SfM optimization is performed based on the image track obtained in the previous step. First, feature point extraction (such as SIFT, ORB, etc.) is performed on all ground images and a matching relationship is established. Then, global bundle adjustment optimization is performed to optimize the accurate poses (T ci) of all images (i.e., camera poses) and the spatial positions of all three-dimensional feature points (sparse points). During the optimization process, the reprojection error constraint is satisfied, that is, the camera poses and three-dimensional points after optimization should make the error between the positions of the three-dimensional points re-projecting onto the image and the positions of the originally extracted feature points minimized. In this step, laser track prior constraints are also used, that is, during the optimization process, the camera pose cannot deviate too far from the initial pose in the image track. This constraint can be added as a penalty term to the optimization objective function, which can be in the form of: Σ || T ci_optimized - T ci_initial ||², where T ci_initial is the initial pose in the image track and T ci_optimized is the optimized camera pose. After global optimization, the optimal camera pose (second pose information) and high-precision three-dimensional point cloud (ground sparse point cloud) are obtained.

[0112] The above is the ground image reconstruction process, and the final ground model includes second pose information and ground sparse point cloud. The matching information of the feature points in the ground image is used in the reconstruction process.

[0113] The pose information (first pose information, second pose information, relative pose information, and other pose information, etc.) in the embodiments of the present application refers to external parameters, including rotation matrix R and translation vector T, which describe the position (determined by translation) and attitude (determined by rotation) of the image in three-dimensional space.

[0114] In S140, based on the aerial sparse point cloud and the ground sparse point cloud, the target aerial images and the target ground images in the overlapping region of the plurality of aerial images and the plurality of ground images are matched to obtain the relative pose information between the target aerial images and the target ground images.

[0115] As mentioned earlier, images with pure aerial perspective and images with ground perspective are difficult to match due to the difference in perspective, so the embodiments of the present application use the take-off and landing part of the unmanned aerial vehicle image and the nearby ground image to extract and match feature points, thereby improving the success rate of matching and improving the alignment accuracy of the coordinate systems of the air and the ground. This step only performs image matching in the overlapping region. For example, there are 1000 aerial images and 500 ground images, of which there are 300 images in the overlapping region. In this step, the aerial images and the ground images in the 300 images are matched.

[0116] The step can adopt a matching method based on scale and perspective weight to perform sparse matching on the air-ground overlapping region. Due to different shooting distances and resolutions of the aerial image and the ground image, the scale (size) of the same object in different images is greatly different, and therefore the step performs feature point extraction and matching from the images after scale unification.

[0117] S140 specifically can include the following steps:

[0118] S141, determining target aerial images and target ground images with overlapping regions according to the first pose information of the target aerial images and the second pose information of the target ground images.

[0119] In the step, the aerial images and the ground images with the same or similar positions and the same or similar poses can be determined as the target aerial images and the target ground images with overlapping regions.

[0120] S142, determining homonymous sparse points located in the target aerial images and the target ground images with overlapping regions based on the aerial sparse point cloud and the ground sparse point cloud to obtain a plurality of pairs of sparse point pairs composed of homonymous sparse points, wherein each pair of sparse point pairs includes one aerial sparse point and one ground sparse point.

[0121] In the step, the three-dimensional coordinates of the aerial sparse points located in the target aerial images with overlapping regions can be obtained based on the aerial sparse point cloud, and the three-dimensional coordinates of the ground sparse points located in the target ground images with overlapping regions can be obtained based on the ground sparse point cloud; and then the aerial sparse points and the ground sparse points with similar three-dimensional coordinates are determined as homonymous sparse points. The homonymous sparse points are similar to the homonymous feature points, and the difference lies in that the homonymous feature points are 2D points and the homonymous sparse points are 3D points, and the homonymous sparse points represent the same physical point in the real world.

[0122] S143, for each pair of sparse point pairs, determining the target aerial image corresponding to the aerial sparse point in the pair of sparse point pairs from the plurality of aerial images to obtain a plurality of target aerial images, and determining the target ground image corresponding to the ground sparse point in the pair of sparse point pairs from the plurality of ground images to obtain a plurality of target ground images.

[0123] In the step, the image corresponding to the sparse point refers to the image in which the sparse point can be observed, and the number of images includes one or more. The more images corresponding to the sparse point, the more stable the sparse point is.

[0124] S144, combining the plurality of target aerial images and the plurality of target ground images two by two to obtain a plurality of pairs of air-ground image pairs, wherein each pair of air-ground image pairs includes one target aerial image and one target ground image.

[0125] For example, for a pair of sparse points, the determined target aerial image is 2, and the determined target ground image is 3, then this step can obtain 6 pairs of aerial-ground image pairs through two-by-two combination between the aerial and ground.

[0126] S145, for each pair of aerial-ground image pairs, performing multi-scale scaling on the target aerial image to obtain a plurality of aerial scaled images of different scales, and performing multi-scale scaling on the target ground image to obtain a plurality of ground scaled images of different scales;

[0127] The multi-scale scaling can be performed in a pyramid-like manner to scale into images of multiple scales, and can be performed according to a plurality of preset scales. By setting a reasonable number of layers, it can be ensured that there are images of the same scale in the scaled target aerial image and the target ground image.

[0128] S146, determining aerial scaled images and ground scaled images of the same scale from the plurality of aerial scaled images and the plurality of ground scaled images;

[0129] S147, performing feature point extraction and matching based on the perspective weight on the aerial scaled images and the ground scaled images of the same scale to obtain relative pose information between the target aerial image and the target ground image;

[0130] Among them, the perspective weight of the feature point extracted from the aerial scaled image with lower shooting height is greater than the perspective weight of the feature point extracted from the aerial scaled image with higher shooting height. The image perspective of the image with lower shooting height is more flat, and the reference is stronger, so a higher weight is given to the feature point in such an image.

[0131] S150, according to the first pose information, the second pose information and the relative pose information, adjusting and optimizing the pose information and the sparse point position of each image in the target aerial image and the target ground image in the overlapping area by using global bundle adjustment with control point constraint to obtain a fusion model.

[0132] This step only optimizes the image information in the overlapping area. For example, there are 1000 aerial images and 500 ground images, and there are 300 images in the overlapping area. This step optimizes the aerial images and the ground images in the 300 images.

[0133] In the aerial and ground fusion process of the embodiment of the present application, the images in the overlapping area need to be optimized. Therefore, the sparse points outside the overlapping area can be fixed as control points, and the control point constraint is added in the optimization process to improve the stability of the optimization.

[0134] S150 can specifically include the following steps:

[0135] S151, determining the control points from all sparse points included in the aerial sparse point cloud and the ground sparse point cloud;

[0136] In this step, the number of images corresponding to all sparse points included in the aerial sparse point cloud and the ground sparse point cloud is obtained; then the sparse points with the number of corresponding images greater than a first preset value and the number of images without overlapping regions in the corresponding images less than or equal to a second preset value are determined as control points. The first preset value can be 4, and the second preset value can be 2.

[0137] S152, according to the first pose information, the second pose information and the relative pose information, using the position of the control points as a constraint, adjusting and optimizing the pose information and sparse point position of each image in the target aerial image and the target ground image with overlapping regions by global bundle adjustment, to obtain a fusion model.

[0138] Since the control points are points outside the overlapping region, or although in the overlapping region, the corresponding images are more stable, the control points are used as a constraint when global bundle adjustment is performed, so that the position of the control points is not changed, thereby improving the reconstruction stability, and the images without overlapping regions on the ground will not be optimized incorrectly, and similarly, the images without overlapping regions in the air will not be incorrect.

[0139] The fusion model includes a plurality of optimized sparse points and their positions, and the pose information of each optimized image. The position of the control points in the fusion model remains unchanged from the position of the control points before optimization.

[0140] The present application plans the UAV flight path and the ground scanning path for the scanning area, so that the scanning path of the ground scanning device passes through the UAV take-off and landing point, so that there is an overlapping region between the image of the take-off and landing part in the aerial image taken by the UAV and the ground image taken by the ground scanning device. The feature points in the overlapping region can be extracted and matched between the aerial and ground images, and the images in the overlapping region are added to the global bundle adjustment and optimization process, which improves the alignment accuracy of the coordinate system during aerial and ground fusion, improves the consistency of the aerial and ground poses, and has higher accuracy compared to simple alignment of only the aerial image and the ground image. And when global bundle adjustment and optimization is performed, the control points are used as a constraint to improve the reconstruction stability, so that the images without overlapping regions on the ground will not be optimized incorrectly, and similarly, the images without overlapping regions in the air will not be incorrect, reducing the geometric distortion and misplacement problems in the fusion area, and further improving the accuracy.

[0141] After the fusion model containing sparse points and image poses is obtained by the present application, subsequent dense reconstruction and texture mapping operations can be further performed. Since the fusion model obtained by the embodiments of the present application has high accuracy of sparse point positions and image poses, that is, the accuracy of sparse reconstruction is high, and the subsequent dense reconstruction and texture mapping depend on the accuracy of sparse reconstruction, therefore, the accuracy of subsequent dense reconstruction, 3dgs, etc. can be improved by the present application. The fused sparse three-dimensional model has large coverage and high detail accuracy, and is suitable for subsequent city modeling, digital twin, smart city, etc.

[0142] Figure 5 The structure schematic diagram of the terminal device provided by the embodiments of the present application is shown, and the specific implementation of the terminal device is not limited by the embodiments of the present application.

[0143] As shown in Figure 5 The terminal device 3 can include a processor 302 and a memory 304.

[0144] The memory 304 is configured to store a computer program 306. The memory 304 can include a high-speed RAM memory, and can also include a non-volatile memory such as at least one disk memory. The computer program 306 can include computer executable instructions.

[0145] The processor 302 is configured to execute the computer program 306 to implement the three-dimensional reconstruction method embodiments described above.

[0146] The processor 302 can be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the terminal device 3 can be the same type of processor, such as one or more CPUs; or can be different types of processors, such as one or more CPUs and one or more ASICs.

[0147] The embodiments of the present application provide a computer readable storage medium, and the storage medium stores a computer program. The computer program is executed by a processor to implement the three-dimensional reconstruction method embodiments described above.

[0148] The embodiments of the present application provide a computer program, which can be executed by a processor to implement the three-dimensional reconstruction method embodiments described above.

[0149] The embodiments of the present application provide a computer program product, which includes a computer program. The computer program is executed by a processor to implement the three-dimensional reconstruction method embodiments described above.

[0150] In several embodiments provided by the present application, any function if realized in the form of a software function module / unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the technical solutions of the present application can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server terminal device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media capable of storing computer program codes.

[0151] The algorithms and displays presented herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with these teachings, based on the description as set forth herein. In terms of input / output, the structure of such systems and other apparatus can be apparent to those skilled in the art from the description herein. Also, the present embodiments are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the present application as described herein, and any references below to specific languages are provided for disclosure of enablement of the best mode of the present application.

[0152] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps not listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In the claims, the word 'first','second', and 'third', etc. does not imply any order. These words are used to name the elements. The steps of the methods described herein do not have to be performed in the order described, unless otherwise specified.

[0153] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A three-dimensional reconstruction method, characterized in that, The method includes: The scanning area is planned by UAV flight path and ground scanning to obtain UAV shooting flight path and ground scanning path. The UAV shooting flight path includes the UAV take-off and landing point and the UAV shooting path. The ground scanning path includes the scanning path of the ground scanning device and the scanning path of the ground scanning device passes through the UAV take-off and landing point. The system acquires multiple aerial images taken by the UAV from different perspectives along the UAV's shooting route, and multiple ground images taken by the ground scanning device along the ground end scanning path, wherein the coordinate system of the ground images is the same as that of the aerial images. Aerial model reconstruction is performed based on multiple aerial images to obtain aerial sparse point clouds and the first pose information of each aerial image; and ground model reconstruction is performed based on multiple ground images to obtain ground sparse point clouds and the second pose information of each ground image. Based on the sparse point cloud in the air and the sparse point cloud on the ground, the target aerial image and the target ground image with overlapping areas in multiple aerial images and multiple ground images are matched to obtain the relative pose information between the target aerial image and the target ground image. Based on the first pose information, the second pose information, and the relative pose information, the pose information and sparse point positions of each image in the target aerial image and the target ground image with overlapping regions are adjusted and optimized using global binding with control point constraints to obtain a fusion model. Based on the aerial sparse point cloud and the ground sparse point cloud, a target aerial image and a target ground image with overlapping regions are matched among multiple aerial images and multiple ground images to obtain the relative pose information between the target aerial image and the target ground image, including: Based on the first pose information of multiple aerial images of the target and the second pose information of multiple ground images of the target, the aerial images and ground images of the target with overlapping regions are determined. Based on the aerial sparse point cloud and the ground sparse point cloud, sparse points with the same name in the target aerial image and the target ground image located in the overlapping area are determined, resulting in multiple pairs of sparse point pairs composed of sparse points with the same name, wherein each pair of sparse point pairs includes one aerial sparse point and one ground sparse point. For each pair of sparse points, the target aerial image corresponding to the sparse point in the sparse point pair is determined from multiple aerial images to obtain multiple target aerial images; and the target ground image corresponding to the sparse point in the sparse point pair is determined from multiple ground images to obtain multiple target ground images. Multiple aerial images of the target and multiple ground images of the target are combined in pairs to obtain multiple pairs of air-ground image pairs, wherein each pair of air-ground image pairs includes one aerial image of the target and one ground image of the target; For each pair of air-ground images, the target air image is scaled at multiple scales to obtain multiple scaled air images at different scales, and the target ground image is scaled at multiple scales to obtain multiple scaled ground images at different scales. Determine aerial and ground zoom images of the same scale from multiple aerial zoom images and multiple ground zoom images; Feature points are extracted and matched from the aerial and ground scaled images of the same scale to obtain the relative pose information between the target aerial image and the target ground image.

2. The method according to claim 1, characterized in that, The aerial model reconstruction based on multiple aerial images, to obtain an aerial sparse point cloud and the first pose information of each aerial image, includes: Acquire position and orientation (POS) data for each aerial image; Determine the feature points and feature descriptors of each feature point in each aerial image, wherein the feature points include corner points, edge intersections, and the center of high-contrast regions; Multiple pairs of aerial images are determined from multiple aerial images based on the POS data, wherein each pair of aerial images includes two aerial images with overlapping areas. For each pair of aerial images, identify the same-named feature points in the two aerial images to obtain multiple pairs of matching feature points composed of the same-named feature points; The relative pose information between the two aerial images to which the matching feature points belong is determined based on the matching feature points. The initial first pose information and the initial aerial sparse point cloud of each aerial image are determined based on the relative pose information between the two aerial images. The initial first pose information and the initial aerial sparse point cloud are then adjusted and optimized through global binding to obtain the first pose information and the aerial sparse point cloud.

3. The method according to claim 1, characterized in that, The ground scanning device is equipped with a lidar and a camera, and the ground image is an image captured by the camera. The step of reconstructing a ground model based on multiple ground images to obtain a sparse point cloud of the ground and second pose information for each ground image includes: Based on the point cloud data collected by the lidar, the laser trajectory is determined, wherein the laser trajectory includes the pose sequence of the lidar; Using pre-labeled extrinsic parameters between the lidar and the camera, the laser trajectory is transferred to the image to obtain an image trajectory, wherein the image trajectory includes the initial estimated pose for each ground image; Based on the image trajectory, the pose information of the ground image is globally bundled, adjusted, and optimized through feature point extraction and matching to obtain the second pose information and the ground sparse point cloud.

4. The method according to claim 1, characterized in that, The step of extracting and matching feature points from the aerial and ground scaled images of the same scale to obtain the relative pose information between the target aerial image and the target ground image includes: Feature point extraction and matching based on viewpoint weights are performed on the aerial zoomed image and the ground zoomed image of the same scale to obtain the relative pose information between the target aerial image and the target ground image. The viewpoint weight of the feature points extracted from the aerial zoomed image at a lower shooting height is greater than the viewpoint weight of the feature points extracted from the aerial zoomed image at a higher shooting height.

5. The method according to claim 4, characterized in that, The step of determining the corresponding sparse points in the target aerial image and the target ground image located in the overlapping region based on the aerial sparse point cloud and the ground sparse point cloud includes: Based on the aerial sparse point cloud, the three-dimensional coordinates of aerial sparse points in the target aerial image located in the overlapping region are obtained, and based on the ground sparse point cloud, the three-dimensional coordinates of ground sparse points in the target ground image located in the overlapping region are obtained. Sparse points in the air and sparse points on the ground with similar three-dimensional coordinates are identified as sparse points with the same name.

6. The method according to claim 1, characterized in that, The step of adjusting and optimizing the pose information and sparse point positions of each image in the target aerial image and target ground image with overlapping regions based on the first pose information, the second pose information, and the relative pose information, using global binding with control point constraints, to obtain a fusion model, includes: Control points are determined from all sparse points included in the aerial sparse point cloud and the ground sparse point cloud; Based on the first pose information, the second pose information, and the relative pose information, and using the position of the control point as a constraint, the pose information and sparse point positions of each image in the target aerial image and the target ground image with overlapping regions are adjusted and optimized using global binding to obtain a fusion model. In the fusion model, the position of the control point remains unchanged from the position of the control point before optimization.

7. The method according to claim 6, characterized in that, The step of determining control points from all sparse points included in the aerial sparse point cloud and the ground sparse point cloud includes: Obtain the number of images corresponding to all sparse points included in the aerial sparse point cloud and the ground sparse point cloud; The sparse points whose corresponding image number is greater than a first preset value and whose corresponding image number does not contain the overlapping region is less than or equal to a second preset value are determined as the control points.

8. The method according to claim 6, characterized in that, The fusion model includes multiple optimized sparse points and their positions, and optimized pose information for each image.

9. The method according to claim 1, characterized in that, The process of planning the drone flight path and the ground-based scanning path for the scanned area, resulting in the drone's imaging flight path and the ground-based scanning path, includes: The scanning area is divided into multiple grids; Each vertex of the grid is designated as the take-off and landing point of the UAV; The start and end shooting heights of the drone are determined to obtain the drone shooting path. The drone shooting path is a vertical path, with its starting point at a position in the vertical direction at the start shooting height from the drone's take-off and landing point, and its ending point at a position in the vertical direction at the end shooting height from the drone's take-off and landing point. The difference between the start shooting height and the shooting height of the ground scanning device is within a preset range, so that the multiple aerial images and multiple ground images contain overlapping areas of the target aerial image and the target ground image. The scanning path of the ground scanning device is determined based on the take-off and landing point of the UAV, so that the scanning path of the ground scanning device passes through the take-off and landing point of the UAV.

10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the three-dimensional reconstruction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for reconstructing map based on air map data and storage medium

    CN116630556A