Three-dimensional reconstruction method and device based on aerial images
By constructing a 3D reconstruction method and device based on aerial images, image data is collected to generate an initial point cloud and iteratively optimized into a Gaussian sputtering model. This solves the problems of long processing time and high computing resources in existing technologies, and realizes efficient, real-time, and high-precision 3D reconstruction, which is applicable to fields such as surveying and mapping, urban planning, and emergency rescue.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU JINGAN TECH CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-19
AI Technical Summary
Existing UAV aerial 3D reconstruction technology has long processing time, high computing resource requirements, and large model size, making it difficult to meet the needs of fast, real-time, and high-precision scene modeling, especially in emergency scenarios where efficient modeling cannot be achieved.
By acquiring multiple aerial images and their position and attitude data, an initial 3D point cloud is generated and transformed into a Gaussian primitive. Through iterative optimization, a Gaussian sputtering 3D model is formed, realizing the close linkage between image acquisition, point cloud generation and Gaussian modeling. A reconstruction strategy covering image acquisition, point cloud generation and Gaussian modeling is constructed. Combined with the linkage mechanism of data fusion and iterative modeling, high-precision and highly continuous 3D reconstruction results are generated.
It significantly improves the efficiency and accuracy of 3D reconstruction from UAV aerial photography, supports rapid modeling in complex scenarios, provides reliable high-precision 3D data support, and is suitable for applications such as surveying, urban planning, emergency rescue, and environmental monitoring.
Smart Images

Figure CN121582512B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a method and apparatus for three-dimensional reconstruction based on aerial images. Background Technology
[0002] With the rapid development of drone technology, 3D reconstruction technology based on drone aerial imagery is becoming an important tool in various fields such as surveying and mapping, urban planning, disaster emergency rescue, environmental monitoring, agricultural management, and infrastructure inspection. This technology enables drones equipped with high-resolution cameras to quickly acquire large-scale surface images and, combined with image processing, computer vision, and 3D modeling algorithms, achieves high-precision 3D reconstruction of terrain, buildings, and various target objects, thus providing an intuitive and accurate data foundation for decision support, planning and design, and emergency response.
[0003] However, existing UAV aerial 3D reconstruction technologies generally suffer from problems such as long processing times, high computational resource requirements, and large model sizes, making it difficult to meet the needs of rapid, real-time, and high-precision scene modeling. Existing methods for 3D reconstruction of aerial images have significant limitations, and efficient 3D reconstruction systems optimized for the characteristics of UAV aerial photography are insufficient to meet the application requirements of emergency response, rapid modeling, and real-time airborne rendering. Summary of the Invention
[0004] This application provides a method and apparatus for 3D reconstruction based on aerial images, which can dynamically construct a regional 3D scene based on multi-source aerial image data acquired by UAVs, achieving high-precision 3D reconstruction. First, multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data, are acquired to form a complete set of observation information. After data preprocessing and feature matching, the reconstruction module generates an initial 3D point cloud based on the acquired images and their pose information, establishing a spatial geometric model of the region. Based on this, the initial 3D point cloud is transformed into a Gaussian primitive, and a Gaussian sputtering 3D model is formed through iterative optimization, further improving the continuity and accuracy of the scene representation. Finally, the Gaussian sputtering 3D model is used to generate a complete 3D reconstruction result, supporting high-quality scene presentation from multiple angles and new viewpoints. This application achieves close linkage between image acquisition, point cloud generation, and Gaussian sputtering modeling by constructing a data mapping and iterative optimization mechanism between aerial images, point clouds, and Gaussian models. This method can quickly and accurately generate high-quality 3D models in complex scenes, providing reliable data support for applications such as surveying, urban planning, emergency rescue, and environmental monitoring, and significantly improving the efficiency and accuracy of UAV aerial 3D reconstruction.
[0005] Firstly, this application provides a three-dimensional reconstruction method based on aerial images, including:
[0006] Collect multiple aerial images of the area to be reconstructed, along with their corresponding location and attitude data;
[0007] An initial 3D point cloud is generated based on multiple aerial images and their corresponding position and attitude data.
[0008] The initial 3D point cloud is converted into a Gaussian primitive, and the Gaussian primitive is iterated to obtain a Gaussian sputtering 3D model. The 3D reconstruction result is generated using the Gaussian sputtering 3D model.
[0009] Secondly, this application provides a three-dimensional reconstruction device based on aerial images, comprising:
[0010] The acquisition module is used to acquire multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data.
[0011] The generation module is used to generate an initial three-dimensional point cloud based on multiple aerial images and their corresponding position and attitude data;
[0012] The reconstruction module is used to convert the initial 3D point cloud into a Gaussian primitive, iterate the Gaussian primitive to obtain a Gaussian sputtering 3D model, and use the Gaussian sputtering 3D model to generate a 3D reconstruction result.
[0013] Thirdly, this application provides a three-dimensional reconstruction device based on aerial images, comprising:
[0014] One or more processors;
[0015] A memory that stores one or more programs that, when executed by one or more processors, enable the one or more processors to implement the three-dimensional reconstruction method based on aerial images as described in the first aspect.
[0016] Fourthly, this application provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the three-dimensional reconstruction method based on aerial images as described in the first aspect.
[0017] In this application, a multi-stage collaborative reconstruction strategy covering image acquisition, point cloud generation, and Gaussian modeling is constructed. Combined with a data fusion and iterative modeling linkage mechanism, automated and high-precision execution of 3D reconstruction based on aerial images is achieved. First, multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data, are acquired to form a complete set of observation information. Simultaneously, the shooting parameters and environmental condition data of candidate image frames are acquired to construct a complete reconstruction resource map. Based on the convergence of multi-source data, the reconstruction module generates a structured instruction set containing core reconstruction elements according to image pose, scene coverage, and reconstruction priority strategies. This instruction set includes the coordinate information of the initial 3D point cloud, parameters for converting the point cloud into Gaussian primitives, the number of iterations, and the control sequence for generating the Gaussian sputtering 3D model. The iteration sequence is parsed by the reconstruction algorithm to obtain the optimization path and is pushed to the computing unit as a modeling instruction. Upon receiving the reconstruction instruction, the initial point cloud is transformed into Gaussian primitives and iteratively optimized according to the specified sequence to generate a high-precision, highly continuous Gaussian sputtering 3D model, ultimately outputting the complete 3D reconstruction result. This method achieves efficient linkage between image acquisition, point cloud generation, and Gaussian sputtering modeling. Through data fusion and iterative optimization mechanisms, it ensures the accuracy, integrity, and stability of the reconstruction process, effectively improving the efficiency and visualization quality of UAV aerial 3D reconstruction in complex scenarios. It provides reliable 3D data support for applications such as surveying, urban planning, and emergency rescue, and has high practicality and engineering promotion value. Attached Figure Description
[0018] Figure 1 This is a flowchart of a three-dimensional reconstruction method based on aerial images provided in an embodiment of this application;
[0019] Figure 2 This is a flowchart of the aerial data acquisition method provided in the embodiments of this application;
[0020] Figure 3 This is a flowchart of the initial 3D point cloud generation method provided in the embodiments of this application;
[0021] Figure 4 This is a flowchart of the matching relationship graph generation method provided in the embodiments of this application;
[0022] Figure 5 This is a flowchart of the Gaussian sputtering three-dimensional model iteration method provided in the embodiments of this application;
[0023] Figure 6 This is a flowchart of the Gaussian primitive quantity control method provided in the embodiments of this application;
[0024] Figure 7 This is a flowchart of the Gaussian primitive splitting method provided in the embodiments of this application;
[0025] Figure 8 This is a flowchart of the method for determining the upper limit of Gaussian primitives provided in the embodiments of this application;
[0026] Figure 9 This is a flowchart of the Gaussian primitive compression method provided in the embodiments of this application;
[0027] Figure 10 This is a schematic diagram of the structure of the three-dimensional reconstruction device based on aerial images provided in the embodiments of this application;
[0028] Figure 11 This is a schematic diagram of the structure of a three-dimensional reconstruction device based on aerial images provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as being processed sequentially, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. A process can be terminated when its operation is completed, but it may also have additional steps not included in the drawings. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0030] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0031] With the rapid development of UAV technology, 3D reconstruction based on UAV aerial images has become an important technical means in fields such as surveying, urban planning, and emergency rescue. The PSDK development kit supports developers in integrating customized payload devices and applications onto UAVs, facilitating the multi-functionality of UAV platforms. 3D Gaussian sputtering technology uses a large number of 3D Gaussian primitives to represent the geometry and appearance of a scene, achieving high-quality synthesis of new perspectives and providing an effective tool for high-precision 3D modeling.
[0032] However, existing UAV aerial 3D reconstruction technologies still have significant limitations in practical applications. While UAV aerial photography offers the advantage of rapidly covering large areas, flight time is limited, emergency scenarios place stringent requirements on rapid model generation, and onboard computing resources and network bandwidth are constrained. Traditional photogrammetry methods are time-consuming and produce generally poor texture quality; NeRF methods are slow to train and cannot render in real-time; the original 3DGS methods have long training times, and fixed-resolution training does not fully consider the characteristics of aerial images, leading to an overabundance of Gaussian primitives, making them difficult to run on airborne platforms; furthermore, existing PSDK applications lack deep integration with advanced 3D reconstruction algorithms, hindering efficient modeling.
[0033] In summary, there is currently a lack of PSDK and 3DGS system integration solutions optimized for the characteristics of UAV aerial photography. The training speed is slow and the model size is large, making it difficult to meet the needs of rapid, accurate and efficient 3D reconstruction in emergency scenarios. This seriously restricts the promotion and use of UAV 3D modeling technology in high-time-sensitive application scenarios.
[0034] To address the aforementioned challenges, this application constructs a UAV aerial 3D reconstruction system based on automatic navigation and multi-source sensing capabilities. This system enables fully automated processing from image acquisition to high-precision 3D model generation without human intervention. By integrating an onboard sensing platform, the system can acquire multiple aerial images of the area to be reconstructed in real time, along with their corresponding position and attitude data. Combined with environmental information and dynamic data such as flight parameters, it generates a multi-dimensional reconstruction resource map. Based on this, the reconstruction module utilizes point cloud generation, conversion of the initial 3D point cloud into Gaussian primitives, and an iterative optimization mechanism to automatically generate a Gaussian sputtered 3D model and output the final 3D reconstruction result.
[0035] The reconstruction method proposed in this application not only significantly improves the speed and accuracy of 3D model generation, but also effectively utilizes onboard computing resources, reduces reliance on manual operation, and supports rapid modeling in complex terrains and multiple scenarios. Compared with traditional photogrammetry or fixed-resolution training methods, the system possesses higher real-time performance, continuity, and data coordination capabilities. It can achieve efficient execution of UAV aerial 3D reconstruction while ensuring scene coverage and model quality, providing reliable high-precision 3D data support for applications such as surveying and mapping, urban planning, emergency rescue, and environmental monitoring.
[0036] The 3D reconstruction method based on aerial images provided in this embodiment can be executed by a 3D reconstruction device based on aerial images. The 3D reconstruction device based on aerial images can be implemented by software and / or hardware. The 3D reconstruction device based on aerial images can be composed of two or more physical entities, or it can be composed of a single physical entity.
[0037] The 3D reconstruction device based on aerial imagery is equipped with at least one type of operating system, including but not limited to Android, Linux, and Windows. The device can install at least one application on the operating system; this application can be a built-in application or an application downloaded from a third-party device or server. In this embodiment, the 3D reconstruction device based on aerial imagery can at least execute an application for a 3D reconstruction method based on aerial imagery.
[0038] Figure 1 A flowchart of a 3D reconstruction method based on aerial images provided in an embodiment of this application is given. (Reference) Figure 1 The 3D reconstruction method based on aerial images specifically includes:
[0039] S110: Collect multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data.
[0040] In some embodiments, the area to be reconstructed refers to the surface area or target scene that needs to be reconstructed in three dimensions; aerial images refer to two-dimensional image data covering the area acquired from the air by a drone or other aerial platform, which may be RGB images, multispectral images or high-resolution photographic images; location data refers to the geospatial coordinate information corresponding to each aerial image, including longitude, latitude and altitude, used to determine the spatial position of the image in the actual scene; attitude data refers to the attitude information of the flight platform when each aerial image was taken, including yaw angle, pitch angle and roll angle, used to determine the shooting direction and tilt state of the image.
[0041] In one embodiment, aerial images can be acquired by using a drone equipped with a high-resolution camera to fly autonomously or take pictures along a preset route, while simultaneously recording the spatial position and attitude information corresponding to the images through GNSS, IMU or other positioning sensors.
[0042] In one embodiment, the acquired aerial images and their position and attitude data can be used for subsequent image registration, feature matching, and 3D reconstruction processing to ensure that the reconstruction model can accurately restore the 3D structure and spatial distribution of the area to be reconstructed.
[0043] In one embodiment, the acquired aerial images can be downsampled to reduce the amount of data transmitted.
[0044] In one embodiment, to improve reconstruction accuracy, overlapping coverage planning, exposure control, and attitude correction can be performed on the acquired aerial images to obtain a high-quality, multi-angle image dataset.
[0045] In one embodiment, distortion correction can be performed on the aerial image, and the formula for distortion correction of the aerial image is as follows:
[0046]
[0047] in, and The coordinates after distortion correction. and The radial distortion coefficient is... and For normalized camera coordinates, Let be the square of the normalized radius from the point to the optical center. , and The calculation formula is as follows:
[0048]
[0049] in, and These are pixel coordinates, representing the pixel position of a point in the aerial image in the horizontal and vertical directions. and Here are the coordinates of the camera's principal point, representing the position of the center of the camera's imaging plane in the pixel coordinate system. and This refers to the camera's focal length in both the horizontal and vertical directions.
[0050] Optionally, Figure 2 A flowchart of the aerial data acquisition method provided in an embodiment of this application is given. (Reference) Figure 2 The aerial data acquisition method specifically includes:
[0051] S1101. Obtain the geographical range of the area to be reconstructed and determine the flight altitude and heading overlap rate of the UAV.
[0052] For example, the area to be reconstructed refers to the surface area or target scene that needs to be reconstructed in three dimensions. Its geographical range refers to the longitude, latitude boundaries and area information of the area. The flight altitude refers to the vertical distance of the UAV relative to the ground when performing aerial photography missions, which is used to control the shooting angle and image resolution. The heading overlap rate refers to the overlap ratio of adjacent flight paths or continuous images in the flight direction, which is used to ensure that there is a sufficient common viewing area between multi-view images, so as to facilitate subsequent feature matching and three-dimensional reconstruction.
[0053] In one embodiment, the flight altitude and heading overlap ratio can be determined by calculating the optimal flight altitude and overlap ratio based on the terrain complexity of the area to be reconstructed, the required resolution, and camera parameters, in order to balance the image coverage area and the accuracy of 3D reconstruction.
[0054] In one embodiment, a reasonable heading overlap rate can ensure that there are enough common viewpoints between adjacent images, enhance the reliability of feature matching, and thus improve the accuracy of the initial 3D point cloud generation.
[0055] In one embodiment, the geographical range and flight parameters can be used for UAV route planning and mission execution control, providing basic information for acquiring high-quality aerial images and corresponding position and attitude data.
[0056] S1102. Determine the ground sampling distance of the UAV based on the flight altitude, and calculate the waypoint spacing of the UAV based on the ground sampling distance and the heading overlap rate.
[0057] For example, flight altitude refers to the vertical distance of the UAV relative to the ground, used to control the spatial resolution of aerial images; ground sampling distance refers to the actual spatial spacing of the image on the ground at the flight altitude, i.e., the actual ground distance represented by each pixel, used to measure image resolution and detail capture capability; heading overlap rate refers to the overlap ratio of adjacent flight paths or continuous images in the flight direction, used to ensure sufficient common viewing area between images for feature matching and 3D reconstruction; waypoint spacing refers to the horizontal distance between adjacent waypoints along the UAV's flight path, used to plan flight paths and shooting positions.
[0058] In one embodiment, the ground sampling distance can be determined by calculating the actual ground distance corresponding to each pixel based on the drone's flight altitude, camera focal length, and image sensor resolution using similar triangle relationships.
[0059] In one embodiment, the formula for calculating the ground sampling distance can be:
[0060]
[0061] in, Ground sampling distance, For flight altitude, For the physical width of the drone camera sensor, For the focal length of the drone camera, This represents the number of pixels in the horizontal direction of the aerial image.
[0062] In one embodiment, the waypoint spacing can be calculated by multiplying the ground sampling distance by a coefficient to obtain the UAV waypoint spacing while ensuring the heading overlap rate, thereby achieving uniform image coverage and avoiding redundant acquisition.
[0063] In one embodiment, the formula for calculating waypoint spacing can be:
[0064]
[0065] in, The distance between waypoints, This represents the heading overlap rate.
[0066] In one embodiment, a properly calculated waypoint spacing can ensure that the aerial image has enough common viewpoints within the reconstruction area, thereby improving the accuracy of the initial 3D point cloud generation and the completeness of the 3D reconstruction results.
[0067] S1103. Generate multiple waypoints based on the waypoint spacing, control the UAV to fly according to the multiple waypoints, and collect multiple aerial images of the area to be reconstructed and their corresponding position data and attitude data at the locations of the multiple waypoints.
[0068] For example, waypoint spacing refers to the horizontal distance between adjacent waypoints along the flight path, used for planning UAV flight routes; waypoint refers to the three-dimensional spatial location that the UAV needs to reach during flight, and each waypoint may contain target shooting location and shooting direction information; UAV refers to a flight platform that can fly autonomously or semi-autonomously in the air and carry a camera to collect images; aerial images refer to two-dimensional image data collected at waypoint locations, used to cover the area to be reconstructed; location data refers to the geospatial coordinate information corresponding to each aerial image, including longitude, latitude, and altitude; attitude data refers to the yaw angle, pitch angle, and roll angle of the UAV when collecting images at waypoints, used to determine the image shooting direction and tilt state.
[0069] In one embodiment, waypoints can be generated by generating a sequence of waypoints covering the entire area along a preset route, based on the geographical extent of the area to be reconstructed and the planned waypoint spacing, while considering the parallel or staggered layout of the routes to ensure uniform image coverage.
[0070] In one embodiment, the method of controlling the flight and image acquisition of the UAV can be: through a flight control system or autonomous navigation algorithm, the UAV flies in the order of waypoints, while triggering the camera to acquire images and synchronously recording the position information and attitude data of each image.
[0071] In one embodiment, the acquired aerial images and corresponding position and attitude data can be used for subsequent multi-view image registration, feature matching, and initial 3D point cloud generation, thereby ensuring the accuracy and completeness of the 3D reconstruction results.
[0072] S120. Generate an initial three-dimensional point cloud based on the multiple aerial images and their corresponding position and attitude data.
[0073] In some embodiments, aerial images refer to two-dimensional image data covering the area to be reconstructed, acquired by a drone or other aviation platform; location data refers to the geospatial coordinate information corresponding to each aerial image, including longitude, latitude, and altitude, used to determine the spatial location of the image in the actual scene; attitude data refers to the attitude information of the flight platform when each aerial image was taken, including yaw angle, pitch angle, and roll angle, used to determine the image shooting direction and tilt state; the initial 3D point cloud refers to the set of spatial points generated by multi-view image registration, feature matching, and 3D reconstruction algorithms, where each point contains 3D coordinate information and optional attributes such as color and texture, used to describe the 3D shape and spatial structure of the area to be reconstructed.
[0074] In one embodiment, the initial 3D point cloud can be generated by: identifying common viewpoints in multiple aerial images through feature extraction and matching algorithms, performing spatial triangulation by combining the position and attitude data corresponding to the images, obtaining the 3D coordinates of each matching point, and then combining them to form the initial point cloud.
[0075] In one embodiment, to improve the density and accuracy of the initial 3D point cloud, multi-view image fusion, stereo matching optimization, and error correction processing can be used to ensure that the point cloud can fully reflect the geometric structure of the target area.
[0076] In one embodiment, the generated initial 3D point cloud can serve as the basis for subsequent point cloud optimization, mesh reconstruction, and texture mapping, providing accurate spatial information for high-precision 3D reconstruction.
[0077] Optionally, Figure 3 A flowchart of the initial 3D point cloud generation method provided in the embodiments of this application is given. (Reference) Figure 3 The initial 3D point cloud generation method specifically includes:
[0078] S1201. Generate a matching relationship diagram based on the multiple aerial images and their corresponding location data.
[0079] For example, aerial images refer to two-dimensional image data covering the area to be reconstructed, collected by a drone at multiple waypoints; location data refers to the geospatial coordinates of the shooting location corresponding to each aerial image, including longitude, latitude, and altitude information; the matching relationship graph refers to the association graph structure constructed by analyzing the overlap relationship and relative position between multiple aerial images to describe the feature matching between images, and is used to guide the subsequent feature extraction, feature matching, and 3D point cloud generation process.
[0080] In one embodiment, the method for generating a matching relationship graph may be: calculating the spatial distance and viewing angle difference between images based on the location data corresponding to the aerial images, determining whether the image pair meets the feature matching conditions by judging whether there is a visual overlap area between the images, and further constructing a relationship graph containing image nodes and their matchable edges.
[0081] In one embodiment, the matching graph may include image nodes, matching edges, and edge weight information, wherein the edge weights can be used to represent the overlapping area, viewpoint similarity, or number of potential co-viewpoints between images, so as to improve the accuracy and efficiency of subsequent feature matching.
[0082] In one embodiment, the generated matching graph can effectively reduce redundant image matching operations, providing an efficient image matching path and registration basis for subsequent multi-view geometric calculations, initial 3D point cloud generation, and 3D structure optimization.
[0083] Optionally, Figure 4 A flowchart of the matching relationship graph generation method provided in an embodiment of this application is given. (Reference) Figure 4 The matching graph generation method specifically includes:
[0084] S12011. Combine the multiple aerial images in pairs to obtain multiple pairs of first image combinations, wherein the first image combination includes a first image and a second image.
[0085] For example, multiple aerial images are combined in pairs to obtain multiple pairs of first image combinations. A first image combination refers to an image pair consisting of a first image and a second image, used for subsequent image overlap determination and feature matching analysis. Here, an aerial image refers to two-dimensional image data covering the area to be reconstructed, collected by a UAV at multiple waypoints; a first image refers to any image selected from multiple aerial images; a second image refers to another image paired with the first image to form a combination; and a first image combination refers to a basic image pairing unit used to evaluate whether there is visual overlap and a matchable region between the images.
[0086] In one embodiment, the aerial images can be combined in pairs by: performing full or partial permutations of multiple aerial images according to their index or shooting order to generate an image combination set covering all possible pairings, thus establishing complete candidate image pairs for subsequent overlapping area judgment and feature matching.
[0087] In one embodiment, the generated first image combination can be used to quickly filter image pairs with potential shared viewing areas, thereby reducing invalid matches and improving the efficiency and accuracy of the initial 3D point cloud construction.
[0088] S12012. Calculate the projection distance on the ground between the shooting positions of the first image and the second image in each first image combination, and determine the first image combination whose projection distance is less than a distance threshold as the second image combination.
[0089] For example, the first image refers to the first image in an image pair consisting of one of two aerial images; the second image refers to the other image paired with the first image to form an image pair; the shooting position refers to the spatial coordinates of the drone when it acquired the aerial images, including longitude, latitude, and altitude information; the projection distance refers to the Euclidean distance between the planar coordinates obtained by projecting the shooting position onto the ground plane after ignoring the altitude dimension of the shooting position, which is used to measure the possibility of overlap between the horizontal distance and the coverage area of the two images on the ground; the distance threshold refers to a preset distance limit used to determine whether the two images may have a shared viewing area; the second image combination refers to the target image pair selected from the first image combination that meets the projection distance condition, which is used for subsequent feature matching and 3D reconstruction processing.
[0090] In one embodiment, the method for calculating the projected distance of the shooting location can be: converting the longitude and latitude of the shooting location into a plane coordinate system under a unified coordinate system, and using the distance formula to calculate the horizontal distance between the two images on the ground plane.
[0091] In one embodiment, the formula for calculating the projected distance of the shooting position can be:
[0092]
[0093] in, Projected distance from the shooting position. For the Earth's radius, and The latitude of the locations where the first and second images were taken. and The longitude of the location where the first and second images were taken.
[0094] In one embodiment, the method of filtering the second image combination based on the distance threshold can be: comparing the projection distance of each pair of first image combinations, and when the projection distance is less than a set threshold, it is considered that there is a high probability of visual field overlap between the two images, thereby marking the image pair as a valid image combination.
[0095] In one embodiment, the second image combination obtained through screening can enhance the stability and reliability of subsequent feature matching, avoid invalid matching of images that are far apart and have no overlapping areas, thereby improving the efficiency and accuracy of 3D point cloud generation.
[0096] S12013. Extract the first feature and the second feature corresponding to the first image and the second image in each second image combination, and match the first feature and the second feature to obtain the first matching set corresponding to the second image combination.
[0097] For example, a second image combination refers to a combination of two target images acquired under different imaging conditions. The first image in the combination provides baseline visual information, and the second image provides supplementary viewpoint or texture information. The first feature refers to local keypoints, edge textures, or depth-related descriptions extracted from the first image, and the second feature refers to semantic structures, local textures, or gradient patterns extracted from the second image. The first matching set refers to a preliminary matching result set obtained based on the correspondence between the first and second features. After acquiring the second image combination, the first and second features corresponding to the first and second images in each second image combination are extracted, and a first matching set corresponding to that second image combination is generated using a feature matching strategy.
[0098] In one embodiment, the extraction of the first and second features can be achieved by performing gradient direction statistics, local contrast calculation, and deep network convolutional feature encoding on the input image based on a multi-scale feature pyramid, thereby generating a set of feature vectors with stability and discriminative power.
[0099] In one embodiment, the matching of the first feature and the second feature can be achieved by using a feature distance metric function, combined with nearest neighbor matching, ratio testing, and bidirectional consistency verification strategies, to select feature point pairs that satisfy spatial consistency and feature similarity constraints.
[0100] In one embodiment, the first matching set can be constructed by aggregating the effective feature pairs obtained through matching and filtering according to spatial distribution consistency, feature strength threshold, and mismatch elimination rules, and finally forming the first matching set for subsequent reconstruction, registration, or fusion processing.
[0101] S12014. Filter the first matching set corresponding to each second image combination to obtain the second matching set corresponding to each second image combination, and generate a matching relationship graph based on the second matching set corresponding to each second image combination.
[0102] For example, the first matching set refers to the set of candidate matching point pairs generated based on the preliminary feature correspondence, and the second matching set refers to the set of stable matching point pairs obtained after consistency verification, robustness filtering, and mismatch removal. The matching relationship graph refers to the graph structure representation constructed based on the second matching set, used to describe the geometric constraints and feature correspondence between different images.
[0103] In one embodiment, the filtering of the first matching set can be performed by: verifying each feature point pair based on geometric consistency constraints, local structural consistency constraints, and feature intensity thresholds, and removing feature point pairs that do not meet the constraints, thereby obtaining a second matching set with higher stability.
[0104] In one embodiment, filtering the first matching set can be done by using RANSAC to estimate the fundamental matrix and then removing outliers from the first matching set.
[0105] In one embodiment, the matching relationship graph can be generated by: using feature point pairs in the second matching set as edge connections, treating the images participating in the matching as nodes in the graph, and constructing a matching relationship graph that reflects the registration correlation between images by setting edge weights, thus providing structured input for subsequent multi-image fusion, global optimization, or 3D reconstruction.
[0106] S1202. Generate first camera pose data corresponding to each of the multiple aerial images based on the position data and attitude data corresponding to each of the multiple aerial images.
[0107] For example, aerial images refer to images acquired by an imaging device mounted on a drone at different spatial positions and imaging angles; position data refers to the positioning information of the drone in three-dimensional space when the image is captured, such as latitude and longitude, altitude, or three-dimensional coordinates; attitude data refers to the spatial orientation information of the camera at the moment of capture, such as rotation angle, heading angle, pitch angle, and roll angle around three axes. First camera pose data refers to the comprehensive pose parameters that characterize the spatial position and orientation of the camera at the moment of capture in a unified coordinate system, which are used to drive subsequent image alignment, dense reconstruction, or geometric optimization.
[0108] In one embodiment, the first camera pose data can be generated by mapping the position data corresponding to each aerial image to a translation vector in three-dimensional space, and constructing a rotation matrix or rotation quaternion based on the corresponding attitude data to form a complete camera extrinsic representation; and unifying the original position and attitude data into the same geographic coordinate system, body coordinate system or world coordinate system through coordinate system transformation, reprojection calibration or inertial navigation fusion, so as to finally obtain structured first camera pose data.
[0109] In one embodiment, to improve the accuracy of the first camera pose data, a sensor fusion approach can be further adopted: the angular velocity and acceleration data output by the UAV inertial measurement unit are fused with the original GPS or RTK positioning data, and the pose is smoothed and error compensated by extended Kalman filtering, attitude calculation model or sliding window optimization, so that the generated first camera pose data has higher temporal continuity and geometric consistency.
[0110] S1203. Process the multiple aerial images and their corresponding first camera pose data and the matching relationship diagram to generate an initial three-dimensional point cloud.
[0111] For example, aerial images refer to image sequences acquired by a drone from different positions and viewing angles; first camera pose data refers to the set of pose parameters used to describe the spatial position and orientation of the camera in a unified coordinate system; a matching graph refers to a graph structure constructed based on the feature correspondences between different images, used to reveal the geometric correlation and matching connectivity between images. An initial 3D point cloud refers to a sparse set of 3D spatial points obtained through multi-view geometric calculations, used to characterize the basic 3D structure of the scene.
[0112] In one embodiment, the method for generating an initial 3D point cloud based on multiple aerial images and pose data from a first camera can be as follows: based on the stable feature correspondence between adjacent image pairs in the matching relationship graph, the 3D coordinates of each matching feature point are recovered using multi-view geometric constraints; by minimizing the feature reprojection error under different image perspectives, the joint optimization of extrinsic pose and 3D point coordinates is achieved, thereby obtaining a sparse 3D structure with stronger consistency.
[0113] In one embodiment, the method for generating an initial 3D point cloud based on multiple aerial images and first camera pose data can be to construct an incremental SfM optimization target, combined with GPS constraints, as shown in the following equation:
[0114]
[0115] in, For the first The rotation matrix of each camera represents the camera pose. For the first The translation vector of each camera represents its position. For the first The spatial coordinates of a 3D point represent a point in a 3D scene. For projection function, For the camera intrinsic parameter matrix, For the first The camera observed the first... The pixel coordinates of a 3D point These are weighting coefficients used to balance the importance of GPS constraints and reprojection errors. For the first GPS location measurement of each camera.
[0116] In one embodiment, to improve the accuracy of the initial 3D point cloud, local bundle adjustment or global structure optimization can be introduced: key frame images are selected based on the connectivity structure of the matching graph, an optimized graph model is constructed, and spatial offset caused by measurement noise, pose deviation and feature error is further reduced through least squares optimization or factor graph optimization, so that the initial 3D point cloud has higher geometric consistency and depth reliability.
[0117] Optionally, the step of processing the multiple aerial images and their corresponding first camera pose data and the matching relationship graph to generate an initial 3D point cloud includes:
[0118] The multiple aerial images and their corresponding first camera pose data and the matching relationship graph are processed to generate an initial 3D point cloud and second camera pose data corresponding to each of the multiple aerial images.
[0119] For example, aerial images refer to a sequence of images captured by a drone at different spatial locations and observation angles; first camera pose data refers to the preliminary pose parameters representing the spatial position and orientation of the camera during shooting in a unified coordinate system; the matching graph refers to a graph structure generated based on image feature matching, used to reveal the geometric relationships between images. The initial 3D point cloud refers to a sparse set of 3D spatial points generated through multi-view geometric calculations, used to represent the basic structure of the scene; second camera pose data refers to the optimized and adjusted camera pose parameters under the constraints of the initial point cloud, used to improve the accuracy and consistency of image alignment and 3D reconstruction.
[0120] In one embodiment, the method for generating the initial 3D point cloud and the second camera pose data can be as follows: based on the stable feature correspondence in the matching graph, the feature points are triangulated to recover the 3D coordinates and form the initial sparse point cloud; subsequently, through local or global bundle adjustment optimization, the first camera pose data and the initial 3D point cloud are jointly optimized to adjust the camera pose to minimize the reprojection error under all viewpoints, thereby obtaining more accurate second camera pose data.
[0121] In one embodiment, to improve the stability and accuracy of the generated results, a robust optimization method can be further adopted: introducing feature point weighting, mismatch elimination strategy or iterative reweighted least squares method to constrain abnormal matching points or high error perspectives, so that the final generated initial 3D point cloud and second camera pose data are more reliable in terms of spatial distribution and geometric consistency, providing high-quality input for subsequent dense point cloud generation and 3D scene reconstruction.
[0122] S130. The initial three-dimensional point cloud is converted into a Gaussian primitive, and the Gaussian primitive is iterated to obtain a Gaussian sputtering three-dimensional model. The three-dimensional reconstruction result is generated using the Gaussian sputtering three-dimensional model.
[0123] In some embodiments, the initial 3D point cloud refers to a set of spatial points generated from multi-view aerial images and their corresponding position and attitude data, used to describe the 3D geometry and spatial structure of the area to be reconstructed; the Gaussian primitive refers to a voxel representation constructed at each point location of the initial 3D point cloud with a Gaussian distribution as the kernel, used to transform discrete point cloud data into a continuous probability field representation; iteration refers to updating and densifying the Gaussian primitive multiple times through optimization algorithms to enhance the continuity and morphological integrity of the point cloud; the Gaussian sputtered 3D model refers to the Gaussian volume representation after iterative optimization, which has spatial structural features of continuity and enhanced local details, and can be used to generate the final 3D model; the 3D reconstruction result refers to the visualized 3D model obtained by performing surface reconstruction, mesh generation, and texture mapping on the Gaussian sputtered 3D model.
[0124] In one embodiment, the initial 3D point cloud can be transformed into a Gaussian primitive by assigning a Gaussian kernel to each point and generating a continuous probability field in 3D space, thereby smoothing the discrete characteristics of the point cloud and facilitating iterative optimization.
[0125] In one embodiment, the Gaussian primitive can be iterated by updating the Gaussian parameters based on point density, neighborhood structure, and smoothness constraints using gradient descent or other optimization algorithms to achieve natural densification of the point cloud and enhancement of local geometric details.
[0126] In one embodiment, the method of generating 3D reconstruction results using Gaussian sputtering 3D models can be: by using isosurface extraction, mesh generation, or volume rendering methods, the optimized Gaussian volume representation is transformed into a visualized 3D model, while preserving the geometric shape, local details, and spatial continuity of the target region.
[0127] In one embodiment, the generated 3D reconstruction results can be used for scene visualization, measurement analysis, path planning, or subsequent 3D data processing to provide a high-precision 3D model for the drone aerial photography area.
[0128] Optionally, Figure 5A flowchart of the Gaussian sputtering three-dimensional model iteration method provided in this application embodiment is given. (Reference) Figure 5 The iterative method for the Gaussian sputtering 3D model specifically includes:
[0129] S1301. Generate a multi-level resolution sequence based on the multiple aerial images.
[0130] For example, aerial images refer to a sequence of images captured by a drone at different spatial locations and observation angles; multi-resolution sequences refer to an image pyramid generated from the original images according to different resolution levels, where each resolution level image refers to an image copy with specific pixel size and detail information, used for image processing, feature extraction, or 3D reconstruction operations at different scales.
[0131] In one embodiment, the multi-resolution sequence can be generated by downsampling the original aerial image and generating multiple resolution levels of images through Gaussian filtering or bicubic interpolation. Each level of image maintains the spatial proportions of the original image, but the number of pixels decreases progressively, thus forming a top-down resolution pyramid structure.
[0132] In one embodiment, generating a multi-level resolution sequence can be achieved by: performing a 2D-FFT on the aerial image to calculate frequency energy; calculating the overall scene frequency distribution based on the frequency energy; calculating window energies for different scaling factors S based on the overall scene frequency distribution; determining the initial resolution scaling factor based on the window energy; and generating a multi-level resolution sequence based on the initial resolution scaling factor. More specifically, the formula for calculating frequency energy is as follows:
[0133]
[0134] in, For frequency energy, For image In frequency domain coordinates The complex Fourier coefficients at the point, and For spatial domain pixel index, and For frequency domain indexing, and For image width and height, For the first Image at pixel coordinates The grayscale value or intensity value at that location.
[0135] The formula for calculating the overall frequency distribution of a scene is shown below:
[0136]
[0137] in, For the overall frequency distribution of the scene, This represents the number of images used for statistical analysis.
[0138] Window energy The calculation formula is as follows:
[0139]
[0140] The formula for calculating the initial resolution scaling factor is as follows:
[0141]
[0142] in, This is the initial resolution scaling factor. The total energy of the image.
[0143] In obtaining Next, construct a multi-level resolution sequence { },in, , In a multi-level resolution sequence, the initial iteration number for each level is:
[0144]
[0145] in, This represents the initial iteration number for each level. It is the minimum value among all frequency energies.
[0146] In one embodiment, to ensure the consistency of details and the availability of features between images of different resolutions, edge-preserving filtering, local contrast enhancement, or noise suppression strategies can be introduced during the downsampling process. This allows multi-resolution sequences to retain key feature information while improving computational efficiency at low resolution levels, providing reliable input for subsequent multi-scale feature matching, sparse point cloud reconstruction, and optimization processing.
[0147] S1302. Output the rendered images corresponding to each of the aerial images based on the Gaussian primitive and the pose data of the second cameras.
[0148] For example, the Gaussian primitive refers to a 3D scene model generated by modeling the initial 3D point cloud using a Gaussian distribution and generating a continuous voxel representation, used to describe the sparse to dense structure of the scene in 3D space; the second camera pose data refers to the optimized and adjusted camera spatial position and orientation parameters; the rendered image refers to the image generated based on the Gaussian primitive under a specific camera viewpoint, reflecting the color, brightness, and structural information of the 3D scene under the corresponding viewpoint.
[0149] In one embodiment, the output rendered image can be generated by projecting the Gaussian primitive onto the image plane corresponding to each second camera pose, and combining voxel color, transparency and lighting information for pixel-by-pixel accumulation and fusion to generate a rendered image with the same viewpoint as the original aerial image.
[0150] In one embodiment, the pixels of the rendered image The formula for calculating the rendering value can be shown below:
[0151]
[0152] in, For pixels The rendering value, This represents the number of elements sampled on the current path. For the first The cumulative transmission coefficient of each sample For the first The opacity of each sample, For the first Each sample is at a distance The luminescence value below, It is a mixed weight.
[0153] Mixed weights The calculation formula is as follows:
[0154]
[0155] in, For the first The weights of each Gaussian body For the current query point, For the first The center coordinates of a Gaussian body.
[0156] In one embodiment, to improve the quality of the rendered image, ray stepping optimization, anti-aliasing processing, and color consistency adjustment can be introduced to ensure that the rendered image retains the spatial structural details of the Gaussian primitive while visually approximating the real captured image, providing accurate input for subsequent multi-scale feature matching, error calculation, and dense 3D reconstruction.
[0157] S1303. The multiple aerial images and their corresponding rendered images are downsampled according to the multi-level resolution sequence to obtain multiple sampled aerial images and their corresponding sampled rendered images.
[0158] For example, a multi-resolution sequence refers to an image pyramid with different resolution levels generated from the original aerial image; an aerial image refers to a sequence of images acquired by a drone at different spatial locations and viewing angles; a rendered image refers to an image generated from Gaussian primitives and second camera pose data, with the same viewpoint as the aerial image. A sampled aerial image refers to a low-resolution copy of the aerial image obtained through a downsampling operation, and a sampled rendered image refers to a low-resolution rendered image copy obtained through the same downsampling operation.
[0159] In one embodiment, downsampling of aerial images and rendered images can be achieved by using bilinear interpolation, bicubic interpolation, or Gaussian filtering downsampling methods based on each resolution level in a multi-level resolution sequence to progressively reduce the original image to the target resolution while preserving the key texture and structural features of the image.
[0160] In one embodiment, to ensure the reliability of the downsampled image in feature matching and subsequent reconstruction, edge-preserving filtering or local contrast enhancement strategies can be introduced during the downsampling process. This reduces the computational load of the sampled aerial images and the sampled rendered images while maintaining the recognizability of feature information at low resolution, thereby providing efficient and stable input for multi-resolution iterative optimization and dense 3D reconstruction.
[0161] S1304. Calculate the rendering loss based on the multiple sampled aerial images and their corresponding sampled rendered images.
[0162] For example, a sampled aerial image refers to a low-resolution copy of an aerial image obtained by downsampling at multiple resolution levels; a sampled rendered image refers to a low-resolution rendered image copy obtained by the same downsampling operation; and rendering loss refers to a quantitative metric used to measure the difference between the sampled rendered image and the corresponding sampled aerial image at the pixel level or feature level, which is used to drive 3D scene reconstruction, voxel optimization, or network parameter adjustment.
[0163] In one embodiment, the rendering loss can be calculated by comparing the pixel values of each sampled aerial image with the corresponding sampled rendered image point by point based on pixel-level error metrics, such as mean squared error, mean absolute error, or weighted perceptual loss, and accumulating the results across all images and resolution levels to obtain the overall rendering loss.
[0164] In one embodiment, the formula for calculating the rendering loss can be:
[0165]
[0166] in, For the first Rendering loss of frame-rendered images For loss weighting coefficients, for Pixel error loss, A collection of image pixels. Total number of pixels For the first A frame of a real image in pixels color value, For structural similarity loss, To render the image, It is a structural similarity index.
[0167] In one embodiment, to improve the robustness and optimization effect of rendering loss calculation, edge weighting, local texture preservation, or multi-scale perceptual loss strategies can be introduced, so that the rendering loss not only reflects global brightness and color differences, but also captures local texture and structural errors, thereby providing effective guidance signals for subsequent voxel optimization, camera pose adjustment, and dense 3D reconstruction.
[0168] S1305. Based on the rendering loss, iterate through multiple Gaussian primitives to obtain a Gaussian sputtering 3D model.
[0169] For example, a Gaussian primitive refers to a 3D scene model generated by modeling an initial 3D point cloud using a Gaussian distribution and generating a continuous voxel representation, used to describe the sparse to dense spatial structure of a scene; rendering loss refers to a quantitative indicator reflecting the pixel or feature-level difference between a sampled rendered image and a sampled aerial image; a Gaussian sputtering 3D model refers to a 3D representation that more accurately fits the optical and geometric information of a real scene while preserving the spatial structure through iterative optimization of the Gaussian primitive.
[0170] In one embodiment, the method of iteratively optimizing the Gaussian primitive through rendering loss can be as follows: using the rendering loss as the optimization objective function, the center position, variance, color weight, and transparency parameters of each Gaussian voxel are adjusted by gradient descent or other optimization algorithms; in each iteration, a rendered image is generated through forward rendering, the rendering loss is calculated, and the loss information is backpropagated to update the Gaussian primitive parameters, thereby gradually approximating the optical and geometric distribution of the real scene.
[0171] In one embodiment, to improve the convergence speed and model accuracy of iterative optimization, a multi-scale resolution strategy, momentum optimization, or adaptive learning rate adjustment can be introduced to enable the Gaussian sputtering 3D model to converge quickly with a limited number of iterations, while ensuring the continuity, detail preservation, and spatial consistency of the model, providing a reliable foundation for subsequent high-precision dense 3D reconstruction and visualization applications.
[0172] Optionally, Figure 6 A flowchart of the Gaussian primitive number control method provided in this application embodiment is given. (Reference) Figure 6 The method for controlling the number of Gaussian primitives specifically includes:
[0173] S1306. During the iteration process of the Gaussian primitive, after each first preset number of iterations, the average gradient of the multiple Gaussian primitives is calculated.
[0174] For example, the Gaussian primitive refers to the continuous voxel representation generated by modeling the initial 3D point cloud using a Gaussian distribution, used to describe the sparse to dense structure of the scene; the first preset number refers to the set iteration step size threshold, used to control the frequency of gradient calculation; the gradient refers to the optimization direction information of the Gaussian primitive parameters in the current iteration state, used to guide the next iteration update.
[0175] In one embodiment, the average gradient of the Gaussian primitive can be calculated by: solving the partial derivatives of the voxel parameters based on the rendering loss between the current Gaussian primitive rendered image and the corresponding sampled aerial image, to obtain the sensitivity information of each parameter to the rendering loss; the gradient information can be used to optimize the algorithm to update the Gaussian primitive parameters, thereby enabling the model to gradually approximate the optical and geometric distribution of the real 3D scene.
[0176] In one embodiment, the formula for calculating the average gradient of the Gaussian primitive is as follows:
[0177]
[0178] in, The average gradient of the Gaussian primitive is given. The cumulative gradient of the Gaussian primitive. For the loss function on the th The center of Gaussian body gradient, To accumulate the gradient count, To prevent small constants from being divided by zero.
[0179] In one embodiment, to improve the stability and efficiency of gradient calculation, gradient accumulation, batch calculation, or regularization constraint strategies can be introduced to reduce the interference of noisy gradients on iterative convergence, while ensuring the optimization effect of Gaussian primitive in terms of spatial continuity, detail preservation, and overall consistency, providing reliable gradient information support for the subsequent generation of high-precision Gaussian sputtering 3D models.
[0180] S1307. If the average gradient of the Gaussian primitive is lower than the first gradient threshold, remove the Gaussian primitive.
[0181] For example, the Gaussian primitive refers to the continuous voxel representation generated by modeling the initial 3D point cloud using a Gaussian distribution, used to characterize the sparse to dense structure of the scene; the gradient refers to the sensitivity information of the voxel parameters to the rendering loss calculated during the iterative optimization process; the first gradient threshold refers to the set minimum gradient value. When the average gradient of the Gaussian primitive is less than this threshold, it indicates that the voxel contributes little to the optimization of the rendering loss or is close to convergence.
[0182] In one embodiment, the Gaussian primitives can be removed by comparing the gradient magnitude of each voxel with a first gradient threshold. When the gradient is lower than the threshold, the voxel is removed from the Gaussian primitives set to reduce redundant information, reduce computational complexity, and maintain the continuity and accuracy of the overall scene structure.
[0183] In one embodiment, to ensure the stability and accuracy of the removal operation, local neighborhood verification or multiple iterations can be performed on low-gradient voxels before removal to avoid premature removal of voxels that may contribute to scene reconstruction. This achieves efficient sparsification during the optimization process and provides a concise and reliable 3D representation for subsequent Gaussian sputtering 3D model generation.
[0184] S1308. If the average gradient of the Gaussian primitive is higher than the second gradient threshold and the neighborhood scale of the Gaussian primitive is greater than the scale threshold, perform a splitting process on the Gaussian primitive.
[0185] For example, the Gaussian primitive refers to a continuous voxel representation generated by modeling an initial 3D point cloud using a Gaussian distribution, used to characterize the sparse to dense structure of the scene; the gradient refers to the sensitivity of voxel parameters to rendering loss during iterative optimization; the second gradient threshold is an upper limit threshold used to judge the optimization potential of voxels; when the voxel gradient is higher than this threshold, it indicates that the voxel contributes significantly to the optimization of rendering loss; the neighborhood scale refers to a parameter describing the spatial range or voxel size covered by the Gaussian primitive; the scale threshold is a preset spatial scale limit used to determine whether voxels need to be subdivided to improve local detail representation. Splitting refers to splitting the Gaussian primitive along the spatial structure to generate multiple sub-voxels to enhance the representation accuracy of the local scene.
[0186] In one embodiment, the splitting process can be performed by dividing the voxel along the main spatial axis or uniform grid direction according to the current voxel's center position, variance, and neighborhood scale to generate several child voxels. Each child voxel inherits the parent voxel's color, transparency, and other attributes, and is assigned weights according to the local gradient, thereby enhancing local details and improving spatial resolution.
[0187] In one embodiment, the formula for performing the splitting process is as follows:
[0188]
[0189] in, For the first The new center coordinates of the Gaussian body in the x and y planes For the first The current center coordinates of the Gaussian body This refers to the Gaussian scaling or step size factor. For the first The scale of a Gaussian body For the first A new scale for Gaussian bodies. For the first The current Gaussian body scale.
[0190] In one embodiment, to ensure the stability and efficiency of the splitting process, a neighborhood consistency constraint or gradient weighting strategy can be introduced to avoid excessive splitting and redundant computation. At the same time, it can ensure that the sub-voxels after splitting can effectively reduce rendering loss during the rendering optimization process and improve the ability of the Gaussian sputtering 3D model to express details of complex scenes.
[0191] In one embodiment, the formula for calculating the neighborhood scale of the Gaussian primitive can be as follows:
[0192]
[0193] in, At the neighborhood scale, This represents the number of neighboring points, and can be a value of 3. For point of The set of nearest neighbors and Midpoint in three-dimensional space and neighboring points The coordinates.
[0194] Optionally, Figure 7 A flowchart of the Gaussian primitive splitting method provided in the embodiments of this application is given. (Reference) Figure 7 The Gaussian primitive splitting method specifically includes:
[0195] S13081. Obtain the resolution scaling factor corresponding to the current iteration, and calculate the number of target original bodies based on the resolution scaling factor.
[0196] For example, the resolution scaling factor refers to the scaling ratio of the image or voxel in the current iteration relative to the original resolution, used to control the spatial resolution at each level in the multi-resolution iterative optimization process; the target primitive number refers to the total number of Gaussian primitives expected to be used to represent the scene at the current iteration resolution, used to guide the voxel addition, subtraction, splitting or sparsification operations.
[0197] In one embodiment, the formula for calculating the resolution scaling factor can be as follows:
[0198]
[0199] in, This is the resolution scaling factor. and Discrete time points and The corresponding scale, For the target resolution point, These are the normalized interpolation weights.
[0200] In one embodiment, the number of target primitives can be calculated by using a linear or nonlinear scaling formula to determine the number of voxels required for the current iteration based on the ratio between the resolution scaling factor and the number of primitive Gaussian primitives. This ensures that fewer voxels are used in the low-resolution stage to accelerate optimization, while more voxels are added in the high-resolution stage to improve the ability to express scene details.
[0201] In one embodiment, the formula for calculating the number of target primitives is as follows:
[0202]
[0203] in, For the target number of primitive bodies, For the first The maximum number of Gaussian bodies achievable in a single iteration.
[0204] In one embodiment, to improve the balance and optimization efficiency of voxel distribution, a local density adjustment or gradient weighting strategy can be introduced to dynamically adjust the number of target primitives according to scene complexity and local feature distribution, so that the voxel distribution at each resolution level meets the rendering accuracy requirements while maintaining efficient utilization of computing resources, providing reliable control parameters for iterative optimization of Gaussian sputtering 3D models.
[0205] S13082. Calculate the densification rate based on the number of target primordia, and calculate the number of splits based on the densification rate.
[0206] For example, the target primitive number refers to the total number of Gaussian primitives expected to be used to represent the scene at the current iteration resolution; the density ratio refers to the incremental ratio of the Gaussian primitive set in the current iteration relative to the voxel density of the previous iteration, which is used to reflect the refinement of the local spatial distribution; and the number of splits refers to the number of child voxels generated by performing split operations on primitives with high gradients and high scales in the current iteration.
[0207] In one embodiment, the densification rate can be calculated by: obtaining the densification adjustment ratio required for the current iteration based on the ratio of the target primitive number to the current actual primitive number, or by combining the scene complexity index, to guide voxel splitting and incremental generation.
[0208] In one embodiment, the density rate is calculated. The formula is shown below:
[0209]
[0210] In one embodiment, the method for calculating the number of splits based on the densification rate can be: mapping the densification rate to the number of splits of each high-gradient voxel, and determining the number of sub-voxels generated after each voxel split through linear, nonlinear, or adaptive allocation strategies. This ensures that voxel density is increased in complex regions to preserve details, while maintaining a lower density in low-complexity regions to save computational resources, thereby providing a balanced and efficient voxel distribution for high-precision reconstruction of Gaussian sputtering 3D models.
[0211] In one embodiment, the number of splits is calculated based on the density rate. The formula is shown below:
[0212]
[0213] S13083. Select Gaussian primitives from the plurality of Gaussian primitives that correspond to the number of splits and perform splitting processing.
[0214] For example, Gaussian primitive refers to the continuous voxel representation generated by modeling the initial 3D point cloud with a Gaussian distribution, used to characterize the sparse to dense structure of the scene; number of splits refers to the number of child voxels generated by performing splitting operations on high-gradient or large-scale voxels in the current iteration; splitting process refers to splitting the parent voxel into multiple child voxels along the spatial axis or uniform grid direction to enhance the detail representation capability of the local scene.
[0215] In one embodiment, the method for selecting Gaussian primitives may be: sorting voxels with high gradients and large spatial coverage based on voxel gradients, neighborhood scale, or rendering loss contribution, selecting a set of voxels corresponding to the number of splits, and performing splitting processing to prioritize the refinement of the regions that have the greatest impact on rendering and geometric reconstruction.
[0216] In one embodiment, the splitting process can be performed by dividing the selected voxel along the main spatial direction or uniformly to generate sub-voxels. Each sub-voxel inherits the color, transparency, and initial variance of the parent voxel, and assigns weights according to the local gradient and density rate, so as to ensure that the sub-voxels after splitting maintain scene continuity and detail accuracy while providing a more refined and balanced 3D representation for subsequent iterative optimization.
[0217] S1309. If the average gradient of the Gaussian primitive is higher than the second gradient threshold and the neighborhood scale of the Gaussian primitive is less than or equal to the scale threshold, perform cloning processing on the Gaussian primitive.
[0218] For example, the Gaussian primitive refers to a continuous voxel representation generated by modeling the initial 3D point cloud using a Gaussian distribution, used to characterize the sparse-to-dense structure of the scene; the gradient refers to the sensitivity of voxel parameters to rendering loss during iterative optimization; the second gradient threshold is an upper limit threshold used to judge the optimization potential of a voxel; when the voxel gradient is higher than this threshold, it indicates that the voxel contributes significantly to the optimization of rendering loss; the neighborhood scale refers to a parameter describing the spatial range or voxel size covered by the Gaussian primitive; and the scale threshold refers to a preset spatial scale limit. Cloning refers to generating one or more copies of a voxel based on its gradient and neighborhood characteristics to enhance the spatial expressiveness and optimization flexibility of the region.
[0219] In one embodiment, the cloning process can be performed by generating one or more sub-voxels in the spatial location or parameter space for high-gradient small-scale voxels that meet the conditions. Each sub-voxel inherits the color, transparency, and variance attributes of the parent voxel. At the same time, the parameters of the sub-voxels can be fine-tuned to increase local detail representation and optimize the degree of freedom.
[0220] In one embodiment, to ensure the effectiveness and computational efficiency of the cloning process, gradient weighting, local neighborhood consistency constraints, or secondary voxel number restrictions can be introduced to ensure that the cloning operation can enhance the local scene representation capability without causing redundant computation, thereby providing reliable support for the iterative optimization and final high-precision reconstruction of the Gaussian sputtering 3D model.
[0221] Optionally, Figure 8 A flowchart of the method for determining the upper limit of the Gaussian primitive provided in this application is given. (Reference) Figure 8 The method for determining the upper limit of the Gaussian primitive specifically includes:
[0222] S1310. Determine the natural densityification number of the Gaussian primitives based on the average gradient of the plurality of Gaussian primitives.
[0223] For example, Gaussian primitive refers to the continuous voxel representation generated by modeling the initial 3D point cloud with a Gaussian distribution, used to characterize the sparse to dense structure of the scene; gradient refers to the sensitivity information of voxel parameters to rendering loss during iterative optimization; natural density quantity refers to the number of sub-voxels that each voxel should increase or maintain in the current iteration based on the voxel gradient adaptively, used to balance local detail preservation and computational resource consumption.
[0224] In one embodiment, the natural densitying quantity can be determined by using a linear or nonlinear mapping relationship based on the average gradient magnitude of each Gaussian primitive and the global gradient distribution, allocating more sub-voxels to high-gradient voxels to enhance local representation capabilities, while allocating fewer or no densitying to low-gradient voxels to optimize computational efficiency.
[0225] In one embodiment, the natural density quantity The calculation formula is as follows:
[0226]
[0227] In one embodiment, to improve the stability and effectiveness of the natural densitying strategy, a neighborhood gradient weighting, local complexity evaluation, or iterative adaptive adjustment mechanism can be introduced. This ensures that the natural densitying quantity can retain sufficient spatial details in key areas while avoiding excessive splitting that leads to redundant computation, thereby providing a balanced and efficient voxel distribution for high-precision reconstruction of Gaussian sputtering 3D models.
[0228] S1311. Calculate the momentum of the Gaussian primitive based on the natural density quantity, and calculate the upper limit of the number of Gaussian primitives based on the momentum, the upper limit of the number being used to limit the number of Gaussian primitives.
[0229] For example, the natural densityization quantity refers to the number of sub-voxels that should be increased or maintained for each voxel in the current iteration, which is adaptively determined based on the voxel gradient; momentum refers to the voxel growth trend calculated based on the natural densityization quantity and iteration history information, which is used to reflect the cumulative adjustment direction and magnitude of voxels in continuous iterations; the quantity limit refers to the total number of Gaussian primitives allowed in the current iteration, which is used to constrain voxel growth and prevent excessive increase in the number of voxels from leading to waste of computational resources or optimization instability.
[0230] In one embodiment, the formula for calculating the initial momentum of the Gaussian primitive can be as follows:
[0231]
[0232] in, To initialize momentum, This represents the initial reference iteration number or the initial number of samples.
[0233] Therefore, the upper bound of the initialized Gaussian primitives It can be set to:
[0234]
[0235] In one embodiment, the Gaussian primitive momentum can be calculated by weighting and accumulating the natural density quantity with the momentum from the previous iteration, and incorporating a forgetting factor. Smoothing is performed to obtain the momentum information of each voxel in the current iteration, so as to reflect its density growth trend.
[0236] In one embodiment, the method for calculating the upper limit of the number of voxels based on momentum can be: mapping voxel momentum to a global upper limit of the number of voxels, and determining the total number of Gaussian primitives allowed in the current iteration through a linear or nonlinear function, thereby controlling the global computational complexity while ensuring the ability to express local details, and avoiding excessive memory consumption or reduced iteration optimization efficiency due to an excessive number of voxels.
[0237] In one embodiment, the upper limit of the number of Gaussian primitives is calculated. The formula is shown below:
[0238]
[0239] in, For the first The upper limit of the number of Gaussian bodies in the next iteration. It is a forgetting factor.
[0240] Optionally, Figure 9 A flowchart of the Gaussian primitive compression method provided in this application embodiment is given. (Reference) Figure 9 The Gaussian primitive compression method specifically includes:
[0241] S1312. Quantize and compress the parameters in the Gaussian sputtering three-dimensional model to obtain the compressed Gaussian sputtering three-dimensional model.
[0242] For example, a Gaussian sputtering 3D model refers to a 3D scene representation generated by iteratively optimizing a Gaussian primitive, which has a continuous spatial structure and rich attributes such as color and transparency; parameters refer to data such as the center position, variance, color weight, and transparency of the Gaussian primitive used to describe the characteristics of voxels; a compressed Gaussian sputtering 3D model refers to a 3D model that reduces the amount of data stored after quantization or encoding, while maintaining the availability of scene geometric and optical features.
[0243] In one embodiment, quantization compression can be performed by mapping the floating-point parameters of the Gaussian primitive to a fixed-precision integer representation, or by reducing the precision of low-importance parameters through an adaptive quantization method based on gradient information, while retaining the key parameters of high-gradient voxels, in order to reduce model storage and transmission overhead.
[0244] In one embodiment, to ensure that the compressed model can still be used for high-precision rendering and subsequent 3D processing, error control, compression ratio limits, or multi-level quantization strategies can be introduced. This allows the compressed Gaussian sputtering 3D model to maintain the continuity of spatial structure and the visual consistency of color and transparency while reducing the amount of data, thereby providing an efficient and reliable model representation for 3D model storage, transmission, and visualization applications.
[0245] S1313. Generate a three-dimensional reconstruction result using the compressed Gaussian sputtering three-dimensional model.
[0246] For example, the compressed Gaussian sputtering 3D model refers to the Gaussian sputtering 3D model that has undergone quantization compression, which retains key information such as the scene's geometry, color, and transparency, while reducing data storage; the 3D reconstruction result refers to the visualized 3D scene representation generated based on the model, which can be used for rendering, analysis, measurement, or other 3D applications.
[0247] In one embodiment, the method for generating the 3D reconstruction result may be: forward rendering the compressed Gaussian primitive, and generating a visual image or point cloud representation through voxel projection, ray stepping or image-based rendering methods, thereby obtaining the 3D reconstruction effect of the scene.
[0248] In one embodiment, to improve the visual quality and spatial accuracy of the 3D reconstruction results, multi-resolution rendering, local detail enhancement, or color consistency correction strategies can be introduced to ensure that the reconstruction results not only reflect the spatial structural features of the Gaussian sputtering 3D model, but also visually approximate the real scene, providing a high-precision and usable 3D data foundation for subsequent measurement, analysis, or virtual reality applications.
[0249] In one embodiment, after training, the rendering contribution of each primitive is calculated. The rendering contribution refers to the weight of each Gaussian primitive's influence on image pixels or features in the overall rendering result. It can be calculated as follows:
[0250]
[0251] in, Characterizing the primitive body Optical viewpoint Visibility, Primitive body Transparency or weight, Characterizing the primitive body The number of rendering passes or the number of pixels covered. Pruning is performed on original voxels whose contribution is below the 5th percentile. Pruning refers to removing voxels that contribute very little to the overall rendering to reduce model complexity and computational overhead while maintaining the overall scene geometry and visual effects. The spherical harmonic coefficients are reduced in order from SH3 to SH2 to reduce storage and computation. This reduction operation ensures a quality loss of less than 0.5dB and approximates the lighting information by preserving the main frequency components. The compressed Gaussian sputtered 3D model is exported as a PLY or binary format. The compressed model refers to the 3D representation after pruning, quantization, and order reduction. Its file size is about 1 / 8 of the original model, which greatly reduces storage and transmission overhead while maintaining usable geometric and optical features. The compressed model is deployed to the Jetson AGX Orin airborne platform, and the CUDA rendering engine is initialized. Deployment characterization loads the model into the embedded GPU environment to achieve real-time rendering and high-performance computing. Test camera trajectories are generated and the rendering frame rate is measured. The rendering frame rate represents the number of image frames rendered per unit time, and its calculation formula is:
[0252]
[0253] in, Characterizing the first The time consumed by frame rendering.
[0254] The rendering quality was then verified by calculating PSNR, which represents the pixel-level difference between the rendered image and the reference image. The formula for PSNR is as follows:
[0255]
[0256] in, Characterize the rendered image, Refer to the reference image. These represent the image's width, height, and number of channels, respectively. The scene coverage is calculated. Coverage characterizes the proportion of the scene covered by the current 3D reconstruction result, and its calculation formula is:
[0257]
[0258] in, Represents the actual number of reconstructed points. This refers to the theoretically expected number of points. If Coverage < 0.85, low-quality areas are identified and grouped using DBSCAN clustering. Low-quality areas refer to spatial regions where reconstructed points are sparse or missing. DBSCAN clustering automatically divides low-quality point sets based on spatial proximity, providing a basis for subsequent waypoint generation. Replacement waypoints are generated for low-quality areas. Replacement waypoints refer to additional waypoints generated by the UAV along a 45° oblique viewing direction, covering four angles, to enhance reconstruction accuracy. The replacement flight path sequence optimizes the UAV's flight path to reduce time costs, and the data is uploaded to the UAV via PSDK to perform the replacement task, achieving data supplementation for low-quality areas and overall 3D reconstruction optimization.
[0259] Based on the above embodiments, Figure 10 This is a schematic diagram of the structure of a 3D reconstruction device based on aerial images provided in an embodiment of this application. (Reference) Figure 10 The three-dimensional reconstruction device based on aerial images provided in this embodiment specifically includes: an acquisition module 21, a generation module 22, and a reconstruction module 23.
[0260] The acquisition module 21 is configured to acquire multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data; the generation module 22 is configured to generate an initial 3D point cloud based on the multiple aerial images and their corresponding position and attitude data; and the reconstruction module 23 is configured to convert the initial 3D point cloud into a Gaussian primitive, iterate the Gaussian primitive to obtain a Gaussian sputtering 3D model, and use the Gaussian sputtering 3D model to generate a 3D reconstruction result.
[0261] The above-described aerial image-based 3D reconstruction device provided in this application integrates key functions such as image acquisition, data processing, point cloud generation, and Gaussian sputtering modeling to construct an automated reconstruction architecture centered on multi-view information perception, 3D point cloud generation, and dynamic model iteration. The device is composed of functional units such as an acquisition module, a generation module, and a reconstruction module, forming a closed-loop control chain from aerial image acquisition, pose data acquisition, initial 3D point cloud generation to Gaussian sputtering 3D model iterative optimization, and 3D reconstruction result output, significantly improving the spatial accuracy and processing efficiency of 3D reconstruction. Specifically, the acquisition module is configured to acquire multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data; the generation module is configured to generate an initial 3D point cloud based on the multiple aerial images and their corresponding position and attitude data; and the reconstruction module is configured to convert the initial 3D point cloud into a Gaussian primitive, iteratively process the Gaussian primitive to obtain a Gaussian sputtering 3D model, and use this model to generate a high-precision 3D reconstruction result.
[0262] The 3D reconstruction apparatus based on aerial images provided in this application embodiment can be used to execute the 3D reconstruction method based on aerial images provided in the above embodiment, and has corresponding functions and beneficial effects.
[0263] Figure 11 This is a schematic diagram of the structure of a three-dimensional reconstruction device based on aerial images provided in an embodiment of this application. (Refer to...) Figure 11 The 3D reconstruction device based on aerial imagery includes a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 and the number of memories 32 in the 3D reconstruction device based on aerial imagery can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the 3D reconstruction device based on aerial imagery can be connected via a bus or other means.
[0264] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the 3D reconstruction method based on aerial images in any embodiment of this application (e.g., the acquisition module 21, generation module 22, and reconstruction module 23 in the 3D reconstruction device based on aerial images). The memory 32 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 32 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0265] The communication device 33 is used for data transmission.
[0266] The processor 31 executes various functional applications and data processing of the device by running software programs, instructions and modules stored in the memory 32, thereby realizing the above-mentioned three-dimensional reconstruction method based on aerial images.
[0267] Input device 34 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 35 may include display devices such as a display screen.
[0268] The aerial image-based 3D reconstruction equipment provided above can be used to execute the aerial image-based 3D reconstruction method provided in the above embodiments, and has corresponding functions and beneficial effects.
[0269] This application also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a three-dimensional reconstruction method based on aerial images. The three-dimensional reconstruction method based on aerial images includes: acquiring multiple aerial images of the area to be reconstructed and their corresponding position and attitude data; generating an initial three-dimensional point cloud based on the multiple aerial images and their corresponding position and attitude data; converting the initial three-dimensional point cloud into a Gaussian primitive; iterating the Gaussian primitive to obtain a Gaussian sputtering three-dimensional model; and using the Gaussian sputtering three-dimensional model to generate a three-dimensional reconstruction result.
[0270] Storage medium—any type of memory device or storage device. The term "storage medium" is intended to include: mounting media, such as CD-ROM, floppy disk, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, etc.; non-volatile memory, such as flash memory, magnetic media (e.g., hard disk or optical storage); registers or other similar types of memory elements, etc. Storage medium may also include other types of memory or combinations thereof. Furthermore, storage medium may reside in a first computer system in which a program is executed, or it may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage medium may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.
[0271] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the above-mentioned three-dimensional reconstruction method based on aerial images, but can also perform related operations in the three-dimensional reconstruction method based on aerial images provided in any embodiment of this application.
[0272] The aerial image-based 3D reconstruction apparatus, storage medium, and aerial image-based 3D reconstruction device provided in the above embodiments can execute the aerial image-based 3D reconstruction method provided in any embodiment of this application. For technical details not described in detail in the above embodiments, please refer to the aerial image-based 3D reconstruction method provided in any embodiment of this application.
[0273] The above description is merely a preferred embodiment and the technical principles employed in this application. This application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions that can be made by those skilled in the art will not depart from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application. The scope of this application is determined by the scope of the claims.
Claims
1. A three-dimensional reconstruction method based on aerial images, characterized in that, include: Collect multiple aerial images of the area to be reconstructed, along with their corresponding location and attitude data, including: Obtain the geographical extent of the area to be reconstructed, and determine the flight altitude and heading overlap of the UAV; The ground sampling distance of the UAV is determined based on the flight altitude, and the waypoint spacing of the UAV is calculated based on the ground sampling distance and the heading overlap rate. Multiple waypoints are generated based on the waypoint spacing, and the UAV is controlled to fly according to the multiple waypoints. At the locations of the multiple waypoints, multiple aerial images of the area to be reconstructed and their corresponding position and attitude data are collected. Based on multiple aerial images and their corresponding position and attitude data, an initial 3D point cloud is generated, including: A matching relationship graph is generated based on multiple aerial images and their corresponding location data; Based on the position and attitude data corresponding to each of the multiple aerial images, generate first camera pose data corresponding to each of the multiple aerial images; The aerial images and their corresponding first camera pose data and the matching relationship diagram are processed to generate an initial three-dimensional point cloud. The initial 3D point cloud is converted into a Gaussian primitive, and the Gaussian primitive is iterated to obtain a Gaussian sputtering 3D model. The 3D reconstruction result is generated using the Gaussian sputtering 3D model.
2. The three-dimensional reconstruction method based on aerial images according to claim 1, characterized in that, The step of generating a matching relationship map based on multiple aerial images and their corresponding location data includes: Multiple aerial images are combined in pairs to obtain multiple first image combinations, each first image combination including a first image and a second image. Calculate the projection distance on the ground between the shooting positions of the first image and the second image in each first image combination, and determine the first image combination whose projection distance is less than a distance threshold as the second image combination; Extract the first feature and the second feature corresponding to the first image and the second image in each second image combination, and match the first feature and the second feature to obtain the first matching set corresponding to the second image combination; The first matching set corresponding to each second image combination is filtered to obtain the second matching set corresponding to each second image combination, and a matching relationship graph is generated based on the second matching set corresponding to each second image combination.
3. The three-dimensional reconstruction method based on aerial images according to claim 1, characterized in that, The step of processing multiple aerial images and their corresponding first camera pose data and the matching relationship graph to generate an initial 3D point cloud includes: The multiple aerial images and their corresponding first camera pose data and the matching relationship graph are processed to generate an initial 3D point cloud and second camera pose data corresponding to each of the multiple aerial images; The process of iterating through the Gaussian primitive to obtain the Gaussian sputtering 3D model includes: Generate a multi-resolution sequence based on multiple aerial images; Based on the Gaussian primitive and the pose data of the second camera, output the rendered images corresponding to the aerial images respectively; The aerial images and their corresponding rendered images are downsampled according to the multi-level resolution sequence to obtain multiple sampled aerial images and their corresponding sampled rendered images. Calculate the rendering loss based on multiple sampled aerial images and their corresponding sampled rendered images; The Gaussian sputtering 3D model is obtained by iterating through multiple Gaussian primitives based on the rendering loss.
4. The three-dimensional reconstruction method based on aerial images according to claim 3, characterized in that, The process of iterating through the Gaussian primitive to obtain the Gaussian sputtering 3D model further includes: During the iteration process of the Gaussian primitive, after each first preset number of iterations, the average gradient of multiple Gaussian primitives is calculated; If the average gradient of the Gaussian primitive is lower than a first gradient threshold, the Gaussian primitive is removed. If the average gradient of the Gaussian primitive is higher than the second gradient threshold and the neighborhood scale of the Gaussian primitive is greater than the scale threshold, the Gaussian primitive is split. If the average gradient of the Gaussian primitive is higher than the second gradient threshold and the neighborhood scale of the Gaussian primitive is less than or equal to the scale threshold, then the Gaussian primitive is cloned.
5. The three-dimensional reconstruction method based on aerial images according to claim 4, characterized in that, The process of iterating through the Gaussian primitive to obtain the Gaussian sputtering 3D model further includes: The natural densityification number of the Gaussian primitives is determined based on the average gradient of the plurality of Gaussian primitives; The momentum of the Gaussian primitive is calculated based on the natural density quantity, and an upper limit on the number of the Gaussian primitive is calculated based on the momentum, the upper limit being used to limit the number of the Gaussian primitive.
6. The three-dimensional reconstruction method based on aerial images according to claim 4, characterized in that, The splitting process performed on the Gaussian primitive includes: Obtain the resolution scaling factor corresponding to the current iteration, and calculate the number of original target bodies based on the resolution scaling factor; The density rate is calculated based on the number of target primordia, and the number of splits is calculated based on the density rate. From the plurality of Gaussian primitives, select the Gaussian primitives corresponding to the number of splits and perform splitting processing.
7. The three-dimensional reconstruction method based on aerial images according to claim 1, characterized in that, The process of iterating through the Gaussian primitive to obtain a Gaussian sputtering 3D model, and using the Gaussian sputtering 3D model to generate a 3D reconstruction result, includes: The parameters in the Gaussian sputtering 3D model are quantized and compressed to obtain the compressed Gaussian sputtering 3D model. The compressed Gaussian sputtering 3D model is used to generate a 3D reconstruction result.
8. A three-dimensional reconstruction device based on aerial images, characterized in that, include: The acquisition module is used to acquire multiple aerial images of the area to be reconstructed, along with their corresponding position and attitude data, including: Obtain the geographical extent of the area to be reconstructed, and determine the flight altitude and heading overlap of the UAV; The ground sampling distance of the UAV is determined based on the flight altitude, and the waypoint spacing of the UAV is calculated based on the ground sampling distance and the heading overlap rate. Multiple waypoints are generated based on the waypoint spacing, and the UAV is controlled to fly according to the multiple waypoints. At the locations of the multiple waypoints, multiple aerial images of the area to be reconstructed and their corresponding position and attitude data are collected. The generation module is used to generate an initial 3D point cloud based on multiple aerial images and their corresponding position and attitude data, including: A matching relationship graph is generated based on multiple aerial images and their corresponding location data; Based on the position and attitude data corresponding to each of the multiple aerial images, generate first camera pose data corresponding to each of the multiple aerial images; The aerial images and their corresponding first camera pose data and the matching relationship diagram are processed to generate an initial three-dimensional point cloud. The reconstruction module is used to convert the initial 3D point cloud into a Gaussian primitive, iterate the Gaussian primitive to obtain a Gaussian sputtering 3D model, and use the Gaussian sputtering 3D model to generate a 3D reconstruction result.