Methods, storage media, and devices for rapid 3D scene reconstruction based on UAV imagery
By employing the multi-view geometry-guided collaborative density control strategies MGG and MGR, the problems of insufficient reconstruction accuracy and high model redundancy in sparse image pair samples during joint 3D reconstruction of aerospace platforms have been solved, achieving efficient and accurate 3D scene reconstruction suitable for complex time-varying scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN ENG UNIV
- Filing Date
- 2026-06-17
- Publication Date
- 2026-07-31
AI Technical Summary
Existing aerospace platform combined with 3D reconstruction technology suffers from insufficient reconstruction accuracy of sparse image pairs, inadequate fusion of multi-source data, and high model redundancy, making it difficult to meet the reconstruction needs of complex time-varying scenes with high precision and timeliness.
The multi-view geometry-guided collaborative density control strategies MGG and MGR are adopted. Through the multi-view geometry-guided growth strategy MGG and the reduction strategy MGR, the densification and reduction operations of Gaussian elements are accurately identified and processed. Combined with multi-view geometric consistency and photometric reconstruction loss term, a highly robust and high-precision 3D scene reconstruction is achieved.
It significantly improves the precision and speed of UAV image reconstruction, reduces the requirement for the number of image pairs on the aerospace platform, compresses the model's memory usage, improves reconstruction inference efficiency, and generates high-precision, high-completeness 3D building models.
Smart Images

Figure CN122492944A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV image 3D modeling technology, specifically a method, storage medium and device for rapid 3D modeling of buildings with few views based on UAV images. Background Technology
[0002] With the comprehensive advancement of digital twins and smart city construction, and the continuous upgrading of demands in fields such as emergency rescue, land surveying, and refined urban management, high-precision and high-timeliness 3D building models have become core foundational data supporting various spatial information applications. Traditional 3D building modeling mainly relies on manual surveying and modeling methods, which have inherent drawbacks such as high cost, long cycle, and accuracy depending on the operator's experience, making it difficult to meet the urgent needs of rapid modeling of large-scale urban areas and dynamic updates of time-varying scenes. In recent years, UAV photography technology, with its advantages of mobility, low cost, short data acquisition cycle, and ability to simultaneously acquire texture information of building tops and multiple facades, has gradually become the mainstream technology for 3D building modeling and has been widely used in the 3D reconstruction of small and medium-scale scenes.
[0003] However, existing 3D reconstruction technologies based on single UAV platforms still face numerous limitations in practical applications. First, due to the limitations of UAV endurance and payload, the coverage area of a single platform operation is limited. For large urban areas or complex terrain, multiple segmented operations and data stitching are required, significantly extending the overall modeling cycle and easily leading to geometric misalignment and texture discontinuities at stitching points. Second, in densely built-up urban centers, heavily tree-covered residential areas, and mountainous terrain, single-view UAV observations inevitably generate numerous data blind spots, resulting in incomplete reconstruction of building base structures, obscured facades, and detailed features, making it difficult to meet the accuracy requirements of high-precision applications. More importantly, most existing 3D reconstruction technologies are designed for static scenes. When facing time-varying scenes with rapidly changing building forms and surrounding environments, such as construction sites and disaster sites, problems such as motion blur, ghosting, and accumulated point cloud registration errors easily occur, failing to accurately capture the dynamic changes of the scene. Meanwhile, the traditional reconstruction model of "full data collection first, offline processing later" has a large time delay. It usually takes several hours or even days from data acquisition to model generation, which is difficult to meet the application scenarios with extremely high timeliness requirements such as emergency rescue and real-time monitoring.
[0004] To compensate for the limitations of single platforms, academia and industry have begun exploring a multi-platform 3D reconstruction technology approach, attempting to fuse multi-source data such as satellite remote sensing imagery, high-altitude fixed-wing UAV imagery, and low-altitude multi-rotor UAV imagery. This leverages the complementary advantages of different platforms in terms of coverage, spatial resolution, and observation perspective to improve reconstruction results. However, current aerospace platform joint reconstruction technology is still in its early stages of development, facing numerous unresolved technical challenges. Existing fusion methods mostly involve simple data stitching and registration at the data level, failing to fully exploit the inherent correlations and complementary characteristics of data from different platforms, thus failing to effectively utilize the advantages of multi-source data. In summary, there is currently a lack of a 3D building reconstruction technology that can efficiently fuse data from multiple aerospace platforms, simultaneously ensuring reconstruction accuracy and timeliness, and is applicable to complex, time-varying scenarios. This technology cannot fully meet the urgent needs of various industries for high-quality 3D spatial data. Summary of the Invention
[0005] This invention aims to address the problems of insufficient reconstruction accuracy and high model redundancy in existing aerospace platform joint 3D reconstruction technologies.
[0006] A fast 3D scene reconstruction method for UAV images, comprising:
[0007] For a single scene to be reconstructed, each Gaussian primitive constituting the 3D Gaussian model is projected onto each sampled viewpoint to obtain the initial rendering result corresponding to each viewpoint. Real images from different camera views and the initial rendering results are used as training sequences. K camera views and their corresponding real images and rendering results are randomly sampled from the training sequences. For each viewpoint, the photometric residual between the rendered photometric value and the real photometric value is calculated at pixel (u,v). The residual matrix is obtained.
[0008] The residual matrix is normalized pixel-by-pixel to obtain a standardized loss map. Based on Generated error mask Determine the growth score guided by multi-view geometry If and only if When the growth score threshold is exceeded, a densification operation is performed.
[0009] For each viewpoint, calculate the pixel-level average absolute error loss between the rendered result and the corresponding real image under the i-th viewpoint. With structural similarity loss To construct a comprehensive photometric reconstruction loss term ;based on and Get each Gaussian element Reduced rating If a certain Gaussian element If the score is below the reduction threshold, a reduction operation will be performed.
[0010] 3D reconstruction is achieved by minimizing the loss of the rendering result corresponding to the Gaussian primitives after processing.
[0011] Furthermore, the photometric residual between the rendered photometric value calculated at pixel (u,v) and the actual photometric value. ,in These represent the camera's field of view. Download the pixel value (u,v) at pixel position (u,v) in the photometric channel c, where c∈{1,2,...C} represents the photometric channel.
[0012] Furthermore, according to Generated error mask as follows:
[0013]
[0014] Where I(·) is the indicator function, which takes a value of 1 when the condition is met, and 0 otherwise; τ is the error threshold, which satisfies the condition... The pixel set P is the key region with low reconstruction quality.
[0015] Furthermore, multi-view geometry-guided growth scoring , where I(·) is the indicator function; For the j-th viewpoint based on a preset threshold The mask; p represents a pixel in the two-dimensional image space, P is the set of pixels; K represents the number of sampling viewpoints. Represents the i-th Gaussian element The projection coverage area under the j-th sampling viewpoint.
[0016] Furthermore, the reduction score ,in For normalization function, Represents the i-th Gaussian element In the Projected coverage area under each sampling viewpoint.
[0017] Furthermore, the loss function in the 3D reconstruction process is achieved by minimizing the loss of the rendering result corresponding to the Gaussian primitives after processing, as follows:
[0018]
[0019]
[0020]
[0021]
[0022] in, and These represent the weighting coefficients for balancing the strength of each constraint. Indicates loss of photometric uniformity. This represents the alignment constraint between the pseudo-normals of the depth map and the rendered normals. Indicates geometric consistency constraints for multiple views; The balance coefficient representing the loss of structural similarity. This represents the normalized cross-correlation balance coefficient. Indicates the mean absolute error loss. Represents structural similarity loss. This represents the normalized cross-correlation loss; The pseudo-normal of the depth map. Let represent the rendering normal, and p represent the pixel at (u, v); where... Indicates reference perspective Depth map, Indicates perspective The depth map is projected onto the depth map of the neighboring viewpoint j.
[0023] Furthermore, normalized cross-correlation loss as follows:
[0024]
[0025] in, The set of pixels representing the entire image domain, where N is the total number of pixels in the image; I is a local spatial window centered at pixel p. For the first The luminous intensity of a real visible light image at each sampling viewpoint, and For the first The final corrected rendering intensity used in the calculation under each sampling perspective; and These respectively represent the corresponding variable I in the local window. The pixel mean and standard deviation within the range; and Then they represent the corresponding variables respectively. In local window The mean and standard deviation of pixels within the range.
[0026] Furthermore, the first The final corrected rendering intensity used in the calculation under each sampling perspective. in, It is the initial luminance intensity directly generated by the 3D Gaussian primitives in each iteration of forward rendering under the j-th sampling viewpoint, that is, the native rendered image without any luminance correction; The gain correction factor represents the viewpoint-dependent factor. This represents the bias correction term.
[0027] A computer storage medium storing at least one instruction, which is loaded and executed by a processor to implement the aforementioned method for rapid 3D scene reconstruction based on UAV images.
[0028] A high-efficiency scene 3D reconstruction device combining aerospace platform, the device includes a processor and a memory, the memory stores at least one instruction, the at least one instruction is loaded and executed by the processor to implement the aforementioned method for rapid scene 3D reconstruction based on UAV images.
[0029] The beneficial effects of this invention are:
[0030] This invention designs a multi-view geometrically guided collaborative density control strategy, MGG, and MGR, to achieve robust and high-precision 3D scene reconstruction based on sparse sequences of unconstrained visible light images. The MGG strategy evaluates the contribution of each Gaussian primitive to reconstruction quality by projecting it onto each sampling viewpoint and calculating the average distribution of high-error pixels within its projection region. Combined with a hard constraint mechanism for multi-view geometric consistency, densification is only performed on Gaussian primitives with importance scores exceeding a preset threshold. This ensures that newly added Gaussian primitives are precisely focused on blind spots where cross-view reconstruction is insufficient, effectively solving the problem of numerous redundant Gaussian ellipsoids caused by the native 3DGS algorithm relying solely on image gradient magnitude densification. Furthermore, it can target areas with few feature points or missing textures in visible light images. Effective geometric completion; the multi-view geometrically guided reduction strategy (MGR) constructs a comprehensive photometric reconstruction loss term weighted by mean absolute error loss and structural similarity loss, obtaining the global error background as a weight constraint. Combined with the growth score of Gaussian elements, it derives the reduction score for each Gaussian element, accurately selecting and reducing redundant Gaussian elements below a preset threshold. This significantly eliminates artifact stacking and random outliers caused by background noise while effectively protecting the geometric topology of core features, avoiding the structural holes in the model caused by the accidental deletion of key small-scale Gaussian elements in traditional single-attribute reduction methods. It significantly reduces the requirement for the number of image pairs on the aerospace platform, greatly compresses model memory usage, and improves reconstruction inference efficiency. It fully leverages the complementary advantages of different aerospace platforms in terms of coverage, spatial resolution, and observation perspective, effectively compensating for the limited coverage and data blind spots of a single UAV platform.
[0031] This invention significantly improves the accuracy and speed of UAV image reconstruction in scenarios with few views. Under conditions where the input UAV image resolution is no less than 1000×1000 and the number of images captured in a single scene does not exceed 15, the accuracy can reach LOD 2.2, and the reconstruction model accuracy can reach CD < 0.54m. Furthermore, this invention exhibits excellent robustness to images with different tilt angles, spatial resolutions, and building sizes. Attached Figure Description
[0032] Figure 1 A flowchart for building reconstruction using image pairs from an aerospace platform;
[0033] Figure 2 Here is a complete structural diagram of MGG and MGR;
[0034] Figure 3 Here is a flowchart of the depth back projection process;
[0035] Figure 4 A quality comparison chart of the reconstructed 3D Gaussian splashing results. Detailed Implementation
[0036] It should be noted that, where there is no conflict, the various embodiments disclosed in this invention can be combined with each other.
[0037] Specific implementation method one: Combining Figure 1 and Figure 2 This implementation method is described below.
[0038] The fast 3D scene reconstruction method for UAV images described in this embodiment includes the following steps:
[0039] Step 1: For a single scene to be reconstructed, obtain a multi-view image sequence captured by different camera poses, i.e., real images; project each Gaussian primitive constituting the 3D Gaussian model to each sampled viewpoint to obtain the initial rendering result corresponding to each viewpoint.
[0040] Gaussian elements are the basic units of a 3D Gaussian distribution, used to represent the color and shape of a small region in a scene. In this embodiment, they are created using spherical harmonic functions. Reconstruction involves fitting a set of optimizable Gaussian elements to multi-view images of a specific scene.
[0041] Step 2: Using real images from different camera perspectives and the initial rendering results as training sequences, randomly sample K camera perspectives from the training sequences. and its corresponding real images With rendering results .
[0042] It should be noted that the rendered result is actually a reconstructed result that needs optimization; once the reconstruction is complete, it becomes the final reconstruction result. Here, "training" refers to the iterative optimization process of parameters in a single scene.
[0043] For each viewpoint, calculate the photometric residual between the rendered photometric value and the actual photometric value at each pixel. :
[0044]
[0045] Where c∈{1,2,...C} represents the photometric channel; u and v represent the camera's field of view. The pixel coordinates below.
[0046] Optionally, the value of the photometric channel C is 3 (RGB visible light image channel).
[0047] Based on pixel photometric residuals The residual matrix, i.e. the residual plot, is obtained.
[0048] Step 3: Perform pixel-by-pixel min-max normalization on the residual matrix to construct a standardized loss map. The loss map is constructed to evaluate the reconstruction quality and to identify anomalous pixels; in this loss map, the closer the value is to 1, the worse the reconstruction quality.
[0049] Obtain a standardized loss plot Represented as:
[0050]
[0051] in, ∈ W and H represent the width and height of the image, respectively. This represents the normalized residual value at pixel (u,v); For the j-th perspective (i.e. The residual matrix under ).
[0052] Step 4: Using the multi-view geometry-guided growth (MGG) strategy, the distribution of high-error pixels within the Gaussian projection region is statistically analyzed across multiple sampling views, thereby quantifying the importance score of each Gaussian pixel. This strategy accurately identifies regions with insufficient cross-view reconstruction and performs directional geometric completion on areas with fewer feature points in the visible light image, ensuring the integrity of the target object's topological structure.
[0053] The indicator function is used to determine whether the pixels within its coverage area have high reconstruction residuals, and a multi-view geometrically guided growth score is calculated for each Gaussian cell. The score is obtained by statistically analyzing the number of high-error pixels within the projected area across all sampled views and then averaging them across all views.
[0054]
[0055] Where I(·) is an indicator function, which takes the value 1 when the condition is true and 0 otherwise; For the j-th viewpoint based on a preset threshold The mask; p represents a pixel in the two-dimensional image space, P is the set of pixels; K represents the number of sampling viewpoints. Represents the i-th Gaussian element The projection coverage area under the j-th sampling viewpoint.
[0056] Error masks are used to accurately identify abnormal pixels by introducing a threshold. right Perform binarization to generate the corresponding error mask. :
[0057]
[0058] Where I(·) is the indicator function, and τ is the preset threshold. This is a standardized loss plot. It satisfies... The pixel set P is the key area with low reconstruction quality, thus determining the error distribution.
[0059] Step 5, or the higher value in Step 4 (greater than the preset threshold) Growth score This means that the Gaussian unit consistently resides in regions of significant reconstruction error across multiple viewing angles, thus qualifying it as a core candidate for increasing density: if and only if the Gaussian unit's growth score... Exceeding the preset threshold At that time, a densification operation is performed on the newly added Gaussian elements to ensure that they can accurately focus on blind spots where cross-view reconstruction is insufficient. Since this statistical process can be completed synchronously during the forward propagation stage of rendering, the timeliness of the judgment is greatly improved. This hard constraint mechanism based on multi-view geometric consistency ensures that the model can effectively perform geometric completion for locations with few feature points or missing textures in visible light images.
[0060] Step Six: When dealing with model redundancy, the native 3DGS algorithm mainly controls the scale by removing Gaussian elements with opacity below a certain threshold or excessively large geometric scales. However, this heuristic method based on a single attribute has limited effectiveness in complex visible light scenes and struggles to fundamentally solve the artifact accumulation problem caused by background noise. To ensure the fidelity of the visible light luminosity field while significantly removing redundant individuals, this invention proposes a multi-view geometry-guided reduction strategy, MGR. In contrast to the increasing density logic of MGG, MGR aims to achieve precise pruning of the model scale by quantifying the actual contribution of each Gaussian element to the global reconstruction quality, such as... Figure 2 As shown. The algorithm first needs to quantify the reconstruction fidelity under this view and construct a comprehensive photometric reconstruction loss term. The construction process includes:
[0061] For each viewpoint, calculate the pixel-level average absolute error loss for both the rendered visible light image and the corresponding real visible light image under the i-th viewpoint. With structural similarity loss Construct and calculate the comprehensive photometric reconstruction loss term ;
[0062]
[0063] in, For balance coefficient, To address the average absolute error loss under the i-th viewpoint, This is for the structural similarity loss under the j-th viewpoint.
[0064] From the mean absolute error (MAE, i.e.) Loss) and Structural Similarity (SSIM) Loss) Weighted composition of losses The MGR strategy can obtain a global error background, thus providing the necessary weight constraints when subsequently calculating the reduction scores of individual Gaussian elements.
[0065] Meanwhile, the mean absolute error loss and structural similarity loss from all perspectives are denoted as... , .
[0066] Step 7: Based on the comprehensive photometric reconstruction loss term Combined with error mask Get each Gaussian element Reduced rating :
[0067]
[0068] in, This is a minimum-maximum normalization function used to map the original statistical values (the input of the function) to the standard probability space [0,1]. Represents the i-th Gaussian element In the Projected coverage area under each sampling viewpoint.
[0069] The rating By statistically analyzing the reconstruction state of the Gaussian element within the projection region under K sampling views and combining it with view-level error weights, the Gaussian element is quantitatively measured. The degree of contribution to reducing the overall photometric reconstruction error. If the reduction score of a certain Gaussian element... Below the preset threshold If the result is positive, it indicates that the primitive contributes very little to maintaining the accuracy of the photometric field under multi-view geometric constraints. It will be selected and a reduction operation will be performed to achieve accurate cleaning of redundant Gaussian primitives and complete the three-dimensional reconstruction.
[0070] It should be noted that each iteration consists of three steps: forward rendering, loss calculation, and backpropagation. When the iteration count reaches the growth / reduction iteration interval (the number of iterations since the last growth / reduction, typically several hundred), MGG / MGR operations are performed respectively. The reduction operation processes Gaussian primitives from the previous iteration and generally does not actively reduce newly grown Gaussian primitives. The growth score threshold is selected in the range of around 5 and can be fine-tuned based on actual results; the reduction score threshold should be selected as high as possible (around 0.9), and can be adjusted based on actual results.
[0071] Total loss function Represented as:
[0072]
[0073] in, Indicates loss of photometric uniformity. and These represent hyperparameters (weighting coefficients) that balance the strength of various constraints, used to adjust the balance between photometric uniformity and geometric regularity. This represents the alignment constraint between the pseudo-normals of the depth map and the rendered normals. This represents a geometric consistency constraint for multiple views.
[0074] For all 3D Gaussian primitives, iterative optimization of their parameters (such as position, size, opacity, and color) is achieved through total loss supervision using multi-view images, enabling the Gaussian ensemble to accurately render views consistent with the input image. During training, to avoid nonlinear deviations in pixel intensity for the same scene point under different viewing angles, a normalized cross-correlation loss function is introduced to enhance the model's robustness to luminance intensity fluctuations, resulting in a luminance consistency loss. The function is defined as:
[0075]
[0076] in, The balance coefficient representing the loss of structural similarity. This represents the normalized cross-correlation balance coefficient. Indicates the mean absolute error loss. Represents structural similarity loss; This represents the normalized cross-correlation loss.
[0077]
[0078] in, N represents the set of pixels in the entire image domain, where N is the total number of pixels in the image. I is a local spatial window centered at pixel p. For the first The luminous intensity of a real visible light image at each sampling viewpoint, and For the first The final corrected rendering intensity used in the calculation from each sampling perspective. and These respectively represent the corresponding variable I in the local window. The pixel mean and standard deviation within the range; and Then they represent the corresponding variables respectively. In local window The mean and standard deviation of pixels within the range.
[0079] By subtracting the mean and dividing by the standard deviation, the influence of local brightness shifts and contrast variations in the image is effectively eliminated, which enables the loss term to robustly capture the underlying evolutionary structural features of visible light images.
[0080] Corrected rendering intensity used in calculation By adaptively adjusting the scaling and translation between viewpoints, the cross-viewpoint discontinuity of luminance intensity is eliminated. The algorithm achieves deep decoupling between geometric structure and luminance characteristics, as shown below:
[0081]
[0082] in, The corrected luminous intensity is calculated for the loss in each iteration; It is the initial luminance intensity directly generated by the 3D Gaussian primitives in each iteration of forward rendering under the j-th sampling viewpoint, that is, the native rendered image without any luminance correction; The gain correction factor represents the viewpoint-dependent factor. This represents the bias correction term.
[0083] Formula (10) is a fixed calculation step in the forward propagation, located after the rendering output and before the loss calculation; parameters , Update during the backpropagation phase ( , The initial values are 1 and 0), and the model's main parameters are iteratively optimized synchronously. Compared to closed-form solutions, inverse optimization... , It does not require a "one-to-one pixel correspondence" prerequisite. Even with large initial geometric deviations, , It can also gradually converge along with the geometric parameters, and will not make incorrect corrections due to poor initial rendering quality. It has better compatibility with coarse initialization and low-quality data due to sparse high-order primitives in the early iterations.
[0084] Standard NCC theoretically maintains the same numerical value for global affine transformations, but this conclusion rests on a strict premise: the rendered image and the ground truth image must be geometrically perfectly aligned with each other, and pixels must correspond one-to-one. In the early stages of training, significant geometric deviations result in poor pixel correspondence, and the ground truth pixels within a local window are not strictly affine with the rendered pixels. At this point, global luminance differences amplify the estimation noise of NCC, causing the gradient direction to deviate from the true geometric optimization direction. The model then needs to "correct the geometry while adapting to the luminance," leading to slow convergence and a tendency to get trapped in local optima. Luminance differences in multi-view data (camera exposure, gain, white balance differences) are inherent systematic properties of each viewpoint, not temporary noise. We hope the model can "remember" the correction values for each viewpoint and reuse them directly during inference, rather than re-estimating them for each calculation.
[0085] Real-world multi-view images often exhibit inconsistent exposure, color temperature, and vignetting; these are "viewpoint-specific noise," not inherent properties of the scene itself. If a network (3D Gaussian model) is directly fitted to the original real image, the network may "remember" the luminance shift of each viewpoint, or even use geometric errors to adapt to the luminance differences, leading to scene geometric deformation. Therefore, using formula (10) for monotonic linear mapping of pixel values does not change the relative brightness relationship between pixels, explicitly modeling the luminance shift of each viewpoint. Corresponding to multiplicative gain, it absorbs the exposure ratio and global contrast differences between viewing angles. Corresponding to additive bias, the overall brightness baseline shift between viewpoints is absorbed. This viewpoint-specific photometric noise is handled by independent parameters. The rendering network no longer needs to "remember" the photometric differences of each viewpoint; it only needs to learn the common geometric structure and inherent materials of the scene itself, achieving decoupling between geometry and photometrics. Edge positions, texture directions, and relative strengths of brightness gradients in the original image are all fully preserved. Therefore, using... , By fitting the viewpoint-level photometric offset separately, the network only needs to learn the geometry and materials of the scene itself, achieving decoupling of geometry and photometrics. The resulting 3D structure is more accurate, and the consistency of the new viewpoint rendering is also better.
[0086] It is a local window-level normalization that addresses the brightness and contrast shifts in local areas, focusing on local structural similarity. , It is a global viewpoint-level pre-correction, which first aligns the luminance range of the two images as a whole, and then calculates the local... Global correction first smooths out the overall luminosity difference between viewpoints, and then local NCC further eliminates local luminosity fluctuations, allowing the loss function to more purely measure the degree of geometric matching and further reducing the interference of luminosity noise on geometric modeling.
[0087] To ensure the local smoothness of the reconstructed surface and prevent non-physical abrupt changes in the depth map, this scheme introduces pseudo-normals of the depth map. Alignment constraint loss with rendering normal N , is represented as:
[0088]
[0089] in, The pseudo-normal of the depth map. This represents the rendering normal, and p represents the pixel at (u,v).
[0090] Represented as:
[0091]
[0092] in, This means using the camera intrinsic parameter matrix K to back-project pixel point p along with the unbiased depth D(p) to a 3D space point in the camera coordinate system. and These represent the spatial partial derivatives of the three-dimensional point in the u and v directions of the image plane, respectively.
[0093] To eliminate potential geometric misalignment between different viewpoints, this scheme utilizes multi-view reprojection constraints. ,like Figure 3 As shown, by projecting the depth map of the reference viewpoint j' onto the neighboring viewpoint j and comparing it with the predicted depth of viewpoint j, it can be represented as follows:
[0094]
[0095] in, Indicates reference perspective Depth map, Indicates perspective The depth map is projected onto the depth map of the neighboring viewpoint j. This is achieved through... Minimal training completes 3D reconstruction.
[0096] To address the issues of insufficient reconstruction accuracy, inadequate multi-source data fusion, high model redundancy, and low reconstruction efficiency in existing aerospace platform-based joint 3D reconstruction technologies, this invention designs a multi-view geometrically guided collaborative density control strategy (MGG) and MGR, achieving robust and high-precision 3D scene reconstruction based on unconstrained visible light image sparse sequences. The MGG strategy evaluates the contribution of each Gaussian primitive to reconstruction quality by projecting it onto each sampling viewpoint and calculating the average distribution of high-error pixels within its projection region. Combined with a hard constraint mechanism for multi-view geometric consistency, densification is only performed on Gaussian primitives with importance scores exceeding a preset threshold, ensuring that newly added Gaussian primitives accurately focus on blind spots where cross-view reconstruction is insufficient. This effectively solves the problem of numerous redundant Gaussian ellipsoids caused by the native 3DGS algorithm relying solely on image gradient magnitude densification. Furthermore, it can target areas with few feature points or missing textures in visible light images. Effective geometric completion; the multi-view geometrically guided reduction strategy (MGR) constructs a comprehensive photometric reconstruction loss term weighted by mean absolute error loss and structural similarity loss, obtaining the global error background as a weight constraint. Combined with the growth score of Gaussian elements, it derives the reduction score for each Gaussian element, accurately selecting and reducing redundant Gaussian elements below a preset threshold. While significantly eliminating artifact stacking and random outliers caused by background noise, it effectively protects the geometric topology of core features, avoiding the structural voids in the model caused by the accidental deletion of key small-scale Gaussian elements in traditional single-attribute reduction methods. It significantly reduces the requirement for the number of image pairs on the aerospace platform, greatly compresses model memory usage, and improves reconstruction inference efficiency. It fully leverages the complementary advantages of different aerospace platforms in terms of coverage, spatial resolution, and observation perspective, effectively compensating for the limited coverage and data blind spots of a single UAV platform. It can quickly generate high-precision, high-completeness 3D building models, greatly reducing manpower and time costs, and providing solid spatial data support for fields such as digital twins, smart city construction, emergency rescue, land surveying, and refined urban management.
[0097] Figure 4 A quality comparison chart of the reconstructed 3D Gaussian splashing results is provided. Experimental analysis shows that, for typical 3D models of terrain features constructed from UAV images, under the conditions that the input UAV image resolution is no less than 1000×1000 and the number of images captured in a single scene does not exceed 15, the level of detail can reach LOD 2.2, and the accuracy of the reconstructed model can reach CD < 0.54m. Specific Implementation Method Two:
[0099] This embodiment is a computer storage medium that stores at least one instruction, which is loaded and executed by a processor to implement the aforementioned method for rapid 3D scene reconstruction based on UAV images.
[0100] It should be understood that the instructions include computer program products, software, or computerized methods corresponding to any method described in this invention; the instructions can be used to program computer systems or other electronic devices. Computer storage media may include readable media on which instructions are stored, and may include, but are not limited to, magnetic storage media, optical storage media; magneto-optical storage media include read-only memory (ROM), random access memory (RAM), erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers, or other types of media suitable for storing electronic instructions. Specific implementation method three:
[0102] This embodiment is a high-efficiency scene 3D reconstruction device that combines aerospace platforms. The device includes a processor and a memory. It should be understood that it includes any device including a processor and a memory described in this invention. The device may also include other units and modules that perform display, interaction, processing, control and other functions through signals or instructions.
[0103] The memory stores at least one instruction, which is loaded and executed by the processor to implement the aforementioned method for rapid 3D scene reconstruction based on UAV images.
[0104] Those skilled in the art will understand that at least one stored instruction constitutes a computer program product corresponding to a method or system. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0105] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, systems, and computer program products according to embodiments of the invention, and can also be used with corresponding devices. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0106] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0108] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0109] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0110] It should be noted that the specific embodiments are merely explanations and illustrations of the technical solution of the present invention and should not be used to limit the scope of protection. Any modifications made in accordance with the claims and specification of the present invention that are only partial should still fall within the protection scope of the present invention.
Claims
1. A method for fast 3D reconstruction of scenes from UAV images, characterized in that, include: For a single scene to be reconstructed, each Gaussian primitive constituting the 3D Gaussian model is projected to each sampling viewpoint to obtain the initial rendering result corresponding to each viewpoint; The real images and initial rendering results of different camera perspectives are taken as a training sequence, K camera perspectives and corresponding real images and rendering results are randomly sampled from the training sequence; for each perspective, the luminosity residual error between the rendering luminosity value result and the real luminosity value is calculated at the pixel (u, v) , to obtain a residual matrix; The residual matrix is normalized pixel-by-pixel to obtain a standardized loss map. Based on Generated error mask Determine the growth score guided by multi-view geometry If and only if When the growth score threshold is exceeded, a densification operation is performed. For each viewpoint, calculate the pixel-level average absolute error loss between the rendered result and the corresponding real image under the i-th viewpoint. With structural similarity loss To construct a comprehensive photometric reconstruction loss term ;based on and Get each Gaussian element Reduced rating If a certain Gaussian element If the score is below the reduction threshold, a reduction operation will be performed. 3D reconstruction is achieved by minimizing the loss of the rendering result corresponding to the Gaussian primitives after processing.
2. The method for rapid 3D scene reconstruction based on UAV images according to claim 1, characterized in that, The photometric residual between the rendered photometric value calculated at pixel (u,v) and the actual photometric value. ,in These represent the camera's field of view. Download the pixel value (u,v) at pixel position (u,v) in the photometric channel c, where c∈{1,2,...C} represents the photometric channel.
3. The method for rapid 3D scene reconstruction based on UAV images according to claim 1, characterized in that, according to Generated error mask as follows: Where I(·) is the indicator function, which takes a value of 1 when the condition is met, and 0 otherwise; τ is the error threshold, which satisfies the condition... The pixel set P is the key region with low reconstruction quality.
4. The method for rapid 3D scene reconstruction based on UAV images according to claim 3, characterized in that, Multi-view geometry-guided growth scoring , where I(·) is the indicator function; For the j-th viewpoint based on a preset threshold The mask; p represents a pixel in the two-dimensional image space, P is the set of pixels; K represents the number of sampling viewpoints. Represents the i-th Gaussian element The projection coverage area under the j-th sampling viewpoint.
5. A method for rapid 3D scene reconstruction based on UAV images according to claim 1, characterized in that, The reduction score ,in For normalization function, Represents the i-th Gaussian element In the Projected coverage area under each sampling viewpoint.
6. A method for rapid 3D scene reconstruction based on UAV images according to any one of claims 1 to 5, characterized in that, The loss function in the 3D reconstruction process is achieved by minimizing the loss of the rendering result corresponding to the Gaussian primitives after processing, as follows: in, and These represent the weighting coefficients for balancing the strength of each constraint. Indicates loss of photometric uniformity. This represents the alignment constraint between the pseudo-normals of the depth map and the rendered normals. Indicates geometric consistency constraints for multiple views; The balance coefficient representing the loss of structural similarity. This represents the normalized cross-correlation balance coefficient. Indicates the mean absolute error loss. Represents structural similarity loss. Indicates the normalized cross-correlation loss; The pseudo-normal of the depth map. Let represent the rendering normal, and p represent the pixel at (u, v); where... Indicates reference perspective Depth map, Indicates perspective The depth map is projected onto the depth map of the neighboring viewpoint j.
7. A method for rapid 3D scene reconstruction based on UAV images according to claim 6, characterized in that, Normalized cross-correlation loss as follows: in, The set of pixels representing the entire image domain, where N is the total number of pixels in the image; I is a local spatial window centered at pixel p. For the first The luminous intensity of a real visible light image at each sampling viewpoint, and For the first The final corrected rendering intensity used in the calculation under each sampling perspective; and These respectively represent the corresponding variable I in the local window. The pixel mean and standard deviation within the range; and Then they represent the corresponding variables respectively. In local window The mean and standard deviation of pixels within the range.
8. A method for rapid 3D scene reconstruction based on UAV images according to claim 6, characterized in that, No. The final corrected rendering intensity used in the calculation under each sampling perspective. in, It is the initial luminance intensity directly generated by the 3D Gaussian primitives in each iteration of forward rendering under the j-th sampling viewpoint, that is, the native rendered image without any luminance correction; The gain correction factor represents the viewpoint-dependent factor. This represents the bias correction term.
9. A computer storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to implement a method for rapid 3D scene reconstruction based on UAV images as described in any one of claims 1 to 8.
10. A high-efficiency scene 3D reconstruction device combining aerospace platforms, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement a method for rapid 3D scene reconstruction based on UAV images as described in any one of claims 1 to 8.