Unmanned aerial vehicle three-dimensional reconstruction method and system based on high-frequency detail enhancement

By employing a multi-scale frequency-spatial joint regularization mechanism and geometric consistency loss optimization, the problems of missing high-frequency details and insufficient geometric consistency in UAV 3D reconstruction are solved, achieving high-precision, real-time 3D reconstruction results, which are suitable for large-scale scenes such as urban building complexes and natural landscapes.

CN121962435APending Publication Date: 2026-05-01WUHAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing UAV 3D reconstruction methods based on the 3DGS framework suffer from problems such as loss of high-frequency details, insufficient geometric consistency, and poor robustness to low-texture and edge area constraints, especially in large-scale scenes such as urban building complexes and natural landscapes.

Method used

By combining a multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss, the high-frequency detail preservation capability and geometric reconstruction accuracy are improved by iteratively updating the 3D Gaussian field parameters.

Benefits of technology

It significantly improves the ability to preserve high-frequency details, enhances the accuracy and robustness of geometric reconstruction, and enables real-time high-quality 3D reconstruction, making it suitable for large-scale scenes such as urban building complexes, large public venues, and natural landscapes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962435A_ABST
    Figure CN121962435A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional reconstruction, and particularly discloses an unmanned aerial vehicle three-dimensional reconstruction method and system based on high-frequency detail enhancement, and the method comprises the steps: obtaining a preprocessed image and a sparse point cloud according to a multi-view image collected by an unmanned aerial vehicle; constructing loss, weighted luminosity loss, multi-level scale regularization loss and enhanced geometric consistency loss of a multi-scale frequency domain-spatial domain joint regularization mechanism based on the preprocessed image; initializing a 3D Gaussian field based on the sparse point cloud, and iteratively updating parameters of the 3D Gaussian field by taking a total loss function fusing loss of a multi-scale frequency domain-space domain joint regularization mechanism, weighted luminosity loss, multi-level scale regularization loss and enhanced geometric consistency loss as a target function; and outputting a three-dimensional reconstruction result of the unmanned aerial vehicle with enhanced high-frequency details and geometric consistency. According to the method, the problems of high-frequency detail missing, insufficient geometric consistency, low texture and poor marginal region constraint robustness of an existing unmanned aerial vehicle multi-view image three-dimensional reconstruction method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of 3D reconstruction technology, specifically relating to a UAV 3D reconstruction method and system based on high-frequency detail enhancement. Background Technology

[0002] The core objective of 3D reconstruction is to construct accurate 3D models of complex scenes from multi-view images or multi-source input data. Unmanned aerial vehicles (UAVs), with their flexible flight attitude, wide-area coverage capabilities, and low-cost data acquisition advantages, have become an important data source for multi-view 3D reconstruction. Currently, the mainstream 3D reconstruction technologies are mainly divided into three categories: traditional multi-view stereo matching (MVS), neural radiation field (NeRF), and 3D Gaussian sputtering (3DGS).

[0003] Traditional MVS methods rely on hand-designed feature descriptors such as SIFT, which are prone to mismatches in low-texture areas (such as smooth walls and open roads), leading to holes or incorrect geometry in the reconstructed model. In large-scale drone scenes, the geometric estimation error is further amplified by factors such as long shooting distance and sparse local view coverage, easily causing "floating artifacts" and "overgrowth" problems in the model. Although subsequent MVSNet achieves end-to-end depth estimation through 4D cost volume and TransMVSNet introduces a cross-attention mechanism to optimize occluded region matching, they still struggle to address the challenges of complex textures, variable lighting, and insufficient robustness in reconstructing sparse view areas in drone scenes. For example, when shooting in urban building clusters, the shadow areas between tall buildings have very little texture information, and traditional MVS methods often fail to accurately match feature points in these areas, resulting in obvious holes or distortions in the reconstructed model.

[0004] NeRF technology achieves photorealistic rendering by parameterizing scene density and radiation information through neural networks, but it suffers from two major bottlenecks: First, volumetric rendering requires millions of network forward propagations, which is too time-consuming to meet real-time requirements. For example, rendering a medium-sized city block using NeRF might take several hours or even longer, which is unacceptable for scenarios such as emergency response where rapid acquisition of 3D models is required. Second, it lacks generalization ability in large drone scenes, easily causing geometric drift that compromises model accuracy. When there are slight deviations in the drone's flight trajectory, the NeRF model may reconstruct originally continuous building structures into discontinuous fragments, severely affecting the model's usability. While optimization methods such as FastNeRF and Mip-NeRF 360 have made improvements, they have not fundamentally solved the problem of balancing efficiency and accuracy. Instant NGP, although speeding up through multi-resolution hashing, sacrifices detail accuracy and cannot meet the needs of fine modeling. Fine structures such as carvings on buildings and window outlines become blurry in models reconstructed by Instant NGP.

[0005] 3DGS technology represents scenes using dynamic Gaussian primitive sets and achieves real-time rendering through fast rasterization, making it widely used in UAV multi-view reconstruction. However, its inherent large-scale Gaussian kernel low-pass filtering suppresses high-frequency components, leading to overly smoothed textures and loss of fine geometric structures. Subsequent optimization methods, such as CompGS and Mip-Splatting, have not addressed the issue of missing high-frequency details; MVSGAussian has improved rendering effects but suffers from insufficient geometric consistency in low-texture areas; FreGS focuses on frequency features but is limited by the local plane assumption, resulting in distortion in non-planar regions; FSGS introduces depth information as a fixed constraint but does not consider the uncertainty of depth estimation. The core flaw of existing methods lies in the failure to achieve synergistic optimization of depth uncertainty, frequency features, and multi-view consistency, leading to an imbalance in constraint strength at the edge regions of the UAV scene and affecting the robustness of geometric reconstruction. For example, in edge regions such as the eaves of buildings, due to drastic depth changes and limited texture information, existing methods often fail to accurately reconstruct their shapes, resulting in blurred or even missing edges. Therefore, there is an urgent need to develop multi-dimensional synergistic optimization methods adapted to the characteristics of pure UAV image data. Summary of the Invention

[0006] The purpose of this invention is to address the problems of high-frequency detail loss, insufficient geometric consistency, and poor robustness of low texture and edge region constraints in existing 3D reconstruction methods based on the 3DGS framework for UAV multi-view image 3D reconstruction. This invention proposes a UAV 3D reconstruction method and system based on high-frequency detail enhancement.

[0007] The technical solution of the present invention is as follows: Firstly, a method for 3D reconstruction of a UAV based on high-frequency detail enhancement, comprising the following steps: The multi-view images acquired by the UAV are preprocessed and initialized to obtain preprocessed images. The preprocessed images are then processed using the structure-reconstruction-motion algorithm to obtain sparse point clouds. Based on preprocessed images, a multi-scale frequency-spatial joint regularization mechanism is constructed, along with loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss. The 3D Gaussian field is initialized based on sparse point clouds. The objective function is the total loss function that integrates the loss of multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss. The 3D Gaussian field parameters are iteratively updated to output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

[0008] As a preferred option, a multi-scale frequency-spatial joint regularization mechanism is used to perform spatial and frequency domain analysis on the preprocessed image to obtain the spatial gradient magnitude and logarithmic magnitude spectrum. The specific methods for spatial domain analysis are as follows: The Sobel gradient operator is used to calculate the gradients of the preprocessed image in the x and y directions using a 3×3 convolution kernel. The square root of the sum of the squares of the gradients in the two directions is then taken to obtain the spatial gradient magnitude. The formula for calculating the spatial gradient magnitude is as follows:

[0009] in, Indicates the magnitude of the spatial gradient. Indicates the partial sign. This indicates a pre-processed image; Using the Laplacian operator, the second derivative is calculated through a 3×3 convolution kernel to enhance sensitivity to high-frequency responses. The calculation formula for detecting subtle texture changes is as follows:

[0010] in, Indicates the second-order differential; The specific methods of frequency domain analysis are as follows: The preprocessed image is downsampled three times to obtain a three-layer image pyramid. A two-dimensional discrete Fourier transform is then performed on each layer of the image pyramid to transform the preprocessed image from the spatial domain to the frequency domain, thus obtaining the frequency domain features.

[0011] in, Indicates the first Frequency domain features corresponding to the layer image pyramid Indicates the index of frequency components in the horizontal direction of the frequency domain. Indicates the index of the frequency component in the vertical direction of the frequency domain. This indicates the image height within the image pyramid. This indicates the width of the image within the image pyramid. Represents the natural constant. To represent a complex unit, Represents pi; Indicates the first Layered image pyramid in ; Logarithmic transform of frequency domain features to construct logarithmic amplitude spectrum :

[0012] in, This represents a logarithmic function with the natural constant as its base.

[0013] As a preferred option, the loss of multi-scale frequency-spatial joint regularization is... The calculation formula is:

[0014] in, Indicates the first The frequency domain characteristics of the image generated by the layered image pyramid through real-time rendering of a 3D Gaussian field in the current iteration. Indicates the first Frequency domain characteristics of real images of layered pyramids. This represents the frequency domain difference measurement function. This indicates the total number of layers in the image pyramid. The weight parameters represent the cross-scale consistency constraints. This represents the cross-scale frequency consistency constraint term. These are scale-weighted parameters.

[0015] As a preferred method, the specific method for constructing the weighted photometric loss is as follows: The high-frequency region mask is obtained by fusing the spatial gradient magnitude and the logarithmic magnitude spectrum; Construct a spatial adaptive weight matrix based on the high-frequency region mask; Weighted photometric loss is constructed based on spatial adaptive weight matrix. :

[0016] in, This indicates the uniform height between the preprocessed image and the 3D Gaussian field rendered image. This indicates the uniform width between the preprocessed image and the 3D Gaussian field rendered image. Represents the spatial adaptive weight matrix. Represents the pixels of a 3D Gaussian field rendered image. Pixel value at that location, Represents actual image pixels The pixel value at that location.

[0017] Preferably, multi-level scale regularization loss is used to constrain the scale of Gaussian elements in a 3D Gaussian field. The calculation formula is:

[0018] in, The weight parameters represent the regularization parameters. This represents the set of Gaussian elements corresponding to the high-frequency region. Represents the scale vector. Indicates the scale threshold. It represents the infinite norm.

[0019] As a preferred approach, the enhanced geometric consistency loss is constructed by fusing depth uncertainty and normal uncertainty, specifically as follows: In terms of depth uncertainty modeling, dense depth maps are generated based on image matching costs between sparse point clouds and preprocessed images. The Sobel operator is used to calculate the depth gradient magnitude of the dense depth map, and the depth gradient magnitude is mapped to a depth uncertainty score in the 0-1 interval. :

[0020] in, Represents the Sobel operator. This represents the Sigmoid function. Represents a dense depth map; In terms of normal uncertainty modeling, the dot product of the normal vectors of each pixel in the preprocessed image with its neighboring pixels is calculated. The average of the dot products of each pixel in the preprocessed image with its neighboring pixels is then used to obtain the normal consistency score. This score is then combined with the gradient information from the rendering distance map to construct a composite normal uncertainty score. :

[0021] in, Indicates the normal uniformity score. Represents the weight parameters. This represents the rendering distance map generated by real-time rendering of the 3D Gaussian field in the current iteration round; Based on depth uncertainty score and normal uncertainty fraction We construct an enhanced geometric consistency loss.

[0022] As a preferred option, enhance geometric consistency loss The calculation formula is:

[0023] in, This represents the geometric consistency mask, which takes a value of 1 if geometrically consistent, and 0 otherwise. Represents pixels in the preprocessed image The score of the depth of uncertainty at that point. Represents pixels in the preprocessed image Uncertainty fraction of the normal at the location, This represents the reprojection error when each pixel of the current image is projected into its neighborhood. As a preferred option, the total loss function The specific formula is:

[0024] in, Indicates weighted photometric loss. This represents the multi-scale frequency-spatial joint regularization loss. This represents the loss of multi-level scale regularization. This represents the loss for enhancing geometric consistency. The weights represent the weighted photometric loss. The weighted multi-scale frequency-spatial joint regularization loss is represented. The weights represent the multi-level scale regularization loss. This represents the weight of the loss for enhancing geometric consistency.

[0025] The beneficial effects of this invention are: 1. Significantly Improved High-Frequency Detail Preservation: This invention utilizes multi-scale frequency-spatial joint regularization to compensate for high-frequency information loss in the frequency domain and capture local details in the spatial domain, effectively overcoming the problem of insufficient sensitivity to high-frequency components in traditional methods. In actual testing, for high-frequency details such as architectural carvings and road markings, the model reconstructed by this invention can clearly present their shape and texture, while these details are relatively blurry or even lost in models reconstructed by traditional 3DGS methods.

[0026] 2. Enhanced Geometric Reconstruction Accuracy and Robustness: This invention avoids excessive smoothing in detailed areas through a high-frequency sensing scale regularization mechanism. The geometric constraints fused with multi-dimensional uncertainty effectively reduce floating artifacts and geometric drift in low-texture and edge regions. In a test scenario of urban building clusters, the model reconstructed by this method exhibits a coherent building structure without any "floating" building fragments, while the model reconstructed by the comparative method shows significant geometric errors at building edges and in low-texture areas.

[0027] 3. Achieving a balance between efficiency and accuracy: Leveraging the rapid rasterization capabilities of the 3DGS framework, real-time rendering is maintained. In relevant tests, reconstruction efficiency and modeling accuracy were balanced, meeting the needs of real-time applications.

[0028] 4. High practicality and wide adaptability: No additional sensor data is required; it can be achieved solely based on UAV imagery, making it suitable for 3D reconstruction of large-scale UAV scenes such as urban building complexes, large public venues, and natural landscapes. In various test scenarios, whether it's a complex urban center or an open natural scenic area, this method can stably output high-quality 3D models, providing strong support for fields such as urban digital twins, cultural relic protection, and disaster monitoring.

[0029] Secondly, a UAV 3D reconstruction system based on high-frequency detail enhancement includes: The first module is used to preprocess and initialize the multi-view images acquired by the UAV to obtain preprocessed images, and then use the motion recovery structure algorithm to process the preprocessed images to obtain sparse point clouds. The second module is used to construct a multi-scale frequency-spatial joint regularization mechanism based on preprocessed images, including loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss. The third module is used to initialize a 3D Gaussian field based on sparse point clouds. It takes the total loss function, which integrates the loss of multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss, as the objective function, and iteratively updates the 3D Gaussian field parameters to output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

[0030] Thirdly, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method as described in the first aspect. Attached Figure Description

[0031] Figure 1 The diagram shows a flowchart of a UAV 3D reconstruction method based on high-frequency detail enhancement.

[0032] Figure 2 The diagram shows a flowchart of a UAV 3D reconstruction method based on high-frequency detail enhancement. Detailed Implementation

[0033] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary and are intended to illustrate the principles and spirit of the invention, and are not intended to limit the scope of the invention.

[0034] Example 1: like Figure 1 and Figure 2 As shown, a UAV 3D reconstruction method based on high-frequency detail enhancement includes the following steps: S1. Preprocess and initialize the multi-view images acquired by the UAV to obtain preprocessed images, and use the motion reconstruction structure algorithm to process the preprocessed images to obtain sparse point clouds; S2. Based on preprocessed images, construct a multi-scale frequency-spatial joint regularization mechanism loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss; S3. Initialize a 3D Gaussian field based on sparse point cloud. Use the total loss function, which integrates loss from multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss, as the objective function to iteratively update the 3D Gaussian field parameters and output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

[0035] In this embodiment, step S1 specifically involves: In the distortion correction stage, using the intrinsic parameter file provided by the camera manufacturer, the `undistort` function in the OpenCV library is used to eliminate radial and tangential distortion of the image, ensuring the accuracy of subsequent feature extraction and matching. For example, for images captured by a DJI Mini 3 Pro drone, after processing with this function, the originally curved edges of buildings become straight, laying the foundation for subsequent geometric analysis. Exposure normalization employs a histogram equalization method to adjust images captured under different lighting conditions to a similar brightness distribution, avoiding feature extraction errors caused by uneven lighting.

[0036] The preprocessed images were then processed using Structure of Motion (SfM) tools such as COLMAP. COLMAP performs feature point detection and description on each image, then finds corresponding points between different images through feature matching, thereby estimating the camera pose (rotation matrix and translation vector) for each viewpoint and generating a sparse point cloud. The generated sparse point cloud serves as the basis for initializing the 3D Gaussian field, and its density and distribution directly affect the quality of the subsequent Gaussian field. Simultaneously, the camera intrinsic parameters (focal length, principal point coordinates, etc.) output by COLMAP, along with extrinsic parameters, establish multi-view geometric constraints. These parameters are crucial for subsequent cross-view geometric consistency calculations.

[0037] In this embodiment, the method for constructing the loss of the multi-scale frequency-spatial joint regularization mechanism is as follows: To fully capture high-frequency detail features, this step constructs a multi-scale frequency and spatial joint regularization mechanism. For spatial feature extraction, the Sobel gradient operator is used, employing a 3×3 convolution kernel to calculate the gradients in the x and y directions of the image. The gradient magnitude is then obtained by taking the square root of the sum of the squared gradients in both directions, as shown in the following formula:

[0038] in, Indicates the magnitude of the spatial gradient. Indicates the partial sign. This indicates a pre-processed image; This allows for precise capture of edge geometric directional features; for example, building corners and road edges will show significantly high values ​​on the gradient magnitude map. Simultaneously, the Laplacian operator is used, employing a 3×3 convolution kernel to calculate the second derivative, further enhancing sensitivity to high-frequency responses and enabling the detection of even subtle texture changes. The specific formula is as follows:

[0039] in, It represents the second-order differential.

[0040] For frequency domain feature extraction, a three-layer image pyramid is first constructed (each layer has a downsampling factor of 2), allowing for the analysis of the image's frequency domain features at different scales. A two-dimensional discrete Fourier transform (FFT) is then performed on each layer of the pyramid to transform the image from the spatial domain to the frequency domain, yielding the frequency domain features.

[0041] in, Indicates the first Frequency domain features corresponding to the layer image pyramid Indicates the index of frequency components in the horizontal direction of the frequency domain. Indicates the index of the frequency component in the vertical direction of the frequency domain. This indicates the image height within the image pyramid. This indicates the width of the image within the image pyramid. Represents the natural constant. To represent a complex unit, Represents pi; Indicates the first Layered image pyramid in ; Different locations in the frequency domain correspond to different frequency components of the image, with high-frequency components mainly concentrated in the edge regions of the frequency domain. Then, a logarithmic amplitude spectrum is constructed using a logarithmic transformation. The specific calculation formula is as follows:

[0042] in, Represents a logarithmic function with the natural constant as its base; Logarithmic transformation can compress the dynamic range of high-frequency components while expanding the dynamic range of low-frequency components, which facilitates the identification and processing of subsequent high-frequency information.

[0043] In the loss calculation stage, a cross-scale frequency consistency optimization objective function is constructed. This function integrates the differences between the predicted and true frequency domain representations of each pyramid level, adjusting the contribution of different levels through scale-weighted parameters. Generally, higher-level pyramids (corresponding to low-resolution images) have smaller weighted parameters, while lower-level pyramids (corresponding to high-resolution images) have larger weighted parameters to highlight the importance of detailed information. Simultaneously, a consistency constraint term is added to ensure the coherence of cross-scale features, enabling frequency domain features at different scales to match each other. To avoid gradient explosion in the early stages of training, a frequency loss dynamic activation strategy based on the Sigmoid function is introduced. After 5000 iterations, the frequency domain constraints are gradually activated, allowing the model to learn the basic geometric structure in the early stages and then refine high-frequency details in the later stages. The specific calculation formula is as follows:

[0044] in, The dynamic activation weights represent the frequency loss. The maximum activation weight represents the frequency loss. This represents the Sigmoid function. Indicates the number of training iterations. This represents the number of iterations for the activation threshold of the frequency constraint. This represents the smoothing parameter.

[0045] Multi-scale frequency-spatial joint regularization loss The calculation formula is:

[0046] in, Indicates the first The frequency domain characteristics of the image generated by the layered image pyramid through real-time rendering of a 3D Gaussian field in the current iteration. Indicates the first Frequency domain characteristics of real images of layered pyramids. This represents the frequency domain difference measurement function. This indicates the total number of layers in the image pyramid. The weight parameters represent the cross-scale consistency constraints. This represents the cross-scale frequency consistency constraint term. These are scale-weighted parameters.

[0047] In this embodiment, the method for constructing the weighted photometric loss and multi-level scale regularization loss is as follows: A high-frequency region mask is generated by fusing the spatial gradient magnitude and the frequency domain logarithmic amplitude spectrum. By setting a dual-threshold filtering mechanism, when the spatial gradient magnitude is greater than 0.7 times the maximum gradient value and the frequency domain logarithmic amplitude spectrum is greater than 0.6 times the maximum logarithmic amplitude value, the region is determined to be a high-frequency region, such as building eaves, road markings, and building carvings.

[0048] For these high-frequency regions, strict scale regularization constraints are applied. Scale regularization employs a multi-level framework design, including a minimum scale constraint to prevent Gaussian degradation, ensuring that the scale of Gaussian primitives (GRPs) is not less than a minimum value to prevent them from becoming too small and losing their representational power; a maximum scale constraint to control overgrowth, limiting the scale of GRPs to a maximum value to prevent excessive expansion and blurring of details; and an axis proportion constraint to maintain geometric rationality, ensuring that the scale proportions of GRPs in the three axes do not exceed a set threshold, allowing GRPs to conform to actual geometry. The influence of each constraint is balanced by different weight parameters, with larger constraint weights set for high-frequency regions to prioritize the preservation of details in these regions. The specific calculation formula is as follows:

[0049] in, The weight parameters represent the regularization parameters. This represents the set of Gaussian elements corresponding to the high-frequency region. Represents the scale vector. Indicates the scale threshold. It represents the infinite norm.

[0050] Simultaneously, a spatial variation weight matrix is ​​constructed based on the high-frequency region mask. An enhancement factor greater than 1 is applied to the luminosity loss in the high-frequency region, dynamically increasing the optimization priority of the detailed region. This ensures that the fine structure can be fully adjusted during the model optimization process, thereby guaranteeing its reconstruction accuracy. The specific calculation formula is as follows:

[0051] in, This indicates the uniform height between the preprocessed image and the rendered image. This indicates the uniform width between the preprocessed image and the rendered image. Represents the spatial adaptive weight matrix. Represents the pixels of the rendered image Pixel value at that location, Represents actual image pixels The pixel value at that location.

[0052] In this embodiment, the method for constructing the enhanced geometric consistency loss is as follows: To improve the robustness of constraints in low-texture and edge regions, this step constructs a geometric consistency constraint mechanism that integrates both depth and normal uncertainties. For depth uncertainty modeling, the sparse depth map generated by SfM is densified using the ACMP framework, combining SfM sparse point cloud data with image matching costs to generate a dense depth map. Then, the Sobel operator is used to calculate the depth gradient, with larger gradient values ​​in regions with drastic depth changes (such as building edges). The gradient magnitude is mapped to an uncertainty score in the 0-1 range using the Sigmoid function; a larger gradient indicates a less reliable depth estimate and a higher uncertainty score. The specific calculation formula is as follows:

[0053] in, Represents the Sobel operator. This represents the Sigmoid function. Represents a dense depth map; For example, at the corner of a building, the depth gradient is larger, resulting in a higher uncertainty score. Therefore, the weight of this area will be appropriately reduced in subsequent constraints.

[0054] In modeling normal uncertainty, the dot product of the normal vectors of each pixel in the preprocessed image with its neighboring pixels is calculated. The closer the dot product value is to 1, the better the normal consistency. The average of the dot products of each pixel with its neighboring pixels is used to obtain the normal consistency score. Simultaneously, combined with the gradient information from the rendered distance map, a composite normal uncertainty score is constructed to comprehensively reflect the smoothness and reliability of the surface geometry. The specific calculation formula is as follows:

[0055] in, Indicates the normal uniformity score. Represents the weight parameters. This indicates the rendering distance map; areas with large changes in normals, such as rough walls, will have a higher uncertainty score.

[0056] Based on depth uncertainty score and normal uncertainty fraction An enhanced geometric consistency loss is constructed. In calculating this loss, the depth information of the current view is reprojected to adjacent views using the camera's intrinsic and extrinsic parameters. Based on the camera's rotation matrix and translation vector, the projected positions of 3D points in adjacent views are calculated, and then the reprojection error is calculated. An error threshold is set, and a consistency mask is generated, retaining only pixels with smaller errors in the loss calculation to reduce outlier interference. Depth and normal uncertainty scores are used as weights to perform a weighted summation of the reprojection errors, achieving dynamic adjustment of the multi-view constraint strength—reducing constraint weights in regions of high uncertainty to prevent incorrect constraints from guiding the model in the wrong direction; and strengthening constraints in reliable regions to ensure accurate model convergence. The specific calculation formula is as follows:

[0057] in, This represents the geometric consistency mask, with a value of 1 if geometrically consistent and 0 otherwise. Represents pixels The score of the depth of uncertainty at that point. Represents pixels Uncertainty fraction of the normal at the location, This indicates the reprojection error.

[0058] In this embodiment, step S3 specifically includes the following steps: A 3D Gaussian field is initialized based on sparse point clouds generated by SfM. Each Gaussian primitive contains core parameters such as position, covariance matrix (decomposed into rotation and scaling matrices), opacity, and spherical harmonic coefficients. The position parameter directly corresponds to the coordinates of the sparse point cloud, ensuring that the initial distribution of the Gaussian field matches the approximate structure of the scene. The covariance matrix is ​​initialized as a diagonal matrix with a fixed scaling factor, giving the Gaussian primitives a certain spatial coverage in the initial stage. The opacity is set to 0.5 to control the visibility of the Gaussian primitives. The spherical harmonic order is set to 3, corresponding to 16 coefficients used to represent the color information of the Gaussian primitives; the specific calculation formula is as follows:

[0059] in, The probability density function represents an anisotropic Gaussian distribution. Indicates the center coordinates, Represents the covariance matrix. Indicates transpose; When constructing the total loss function, four core loss terms are integrated: weighted photometric loss, used to ensure visual consistency between the rendered image and the real image; loss from a multi-scale frequency-spatial joint regularization mechanism, used to protect high-frequency details; multi-level scale regularization loss, used to constrain the scale of Gaussian units; and enhanced geometric consistency loss, used to improve the accuracy of geometric structures. The contribution of each loss term is adjusted by preset weight parameters. Generally, the photometric loss has the largest weight because visual effect is an important evaluation indicator of the model, followed by the scale regularization loss to ensure the rationality of Gaussian units. The specific calculation formula is as follows:

[0060] in, Indicates weighted photometric loss. This represents the multi-scale frequency-spatial joint regularization loss. This represents the loss of multi-level scale regularization. This represents the loss for enhancing geometric consistency. The weights represent the weighted photometric loss. The weighted multi-scale frequency-spatial joint regularization loss is represented. The weights represent the multi-level scale regularization loss. This represents the weight of the loss for enhancing geometric consistency.

[0061] The Adam optimizer was used to iteratively optimize the Gaussian field parameters. After 20,000 iterations, a cosine annealing strategy was applied to decay the parameters, resulting in a total of 30,000 iterations to complete model training. The total loss value was monitored in real time during training, and the final output was a high-precision 3D model containing Gaussian parameters and a colored point cloud. This model can be directly used for subsequent visualization, analysis, and other applications.

[0062] The UAV 3D reconstruction method based on high-frequency detail enhancement proposed in this invention aims to solve three core problems of existing 3DGS-based reconstruction methods in UAV multi-view image 3D reconstruction: loss of high-frequency details, insufficient geometric consistency, and poor robustness to low-texture and edge region constraints. This method achieves high-precision, high-fidelity, and real-time 3D reconstruction in UAV scenarios. Specifically, it aims to ensure that the reconstructed model clearly presents high-frequency details such as architectural carvings and road markings; maintain the coherence and accuracy of the overall geometric structure of the model, reducing "floating artifacts" and geometric drift; and achieve stable and reliable reconstruction results even on low-texture smooth walls, open roads, and building edges, while meeting the requirements of real-time applications. This invention is applicable to high-precision 3D modeling of large-scale scenes such as urban building complexes, large public venues, and natural landscapes captured by UAVs, and can strongly support the application needs of high-fidelity 3D models in fields such as autonomous driving, robot autonomous navigation, virtual reality (VR), and augmented reality (AR). In the construction of digital twins for cities, it can provide accurate three-dimensional basic data for urban planning and architectural design; in the field of cultural relic protection, it can provide high-precision digital archiving and restoration assistance for ancient buildings and sites; in disaster monitoring, it can quickly obtain three-dimensional change information of disaster areas, providing a basis for emergency decision-making.

[0063] Example 2: Based on Embodiment 1, this embodiment of the invention provides a UAV 3D reconstruction system based on high-frequency detail enhancement, which can be used to implement the UAV 3D reconstruction method based on high-frequency detail enhancement as described in the foregoing embodiments. The system includes: The first module is used to preprocess and initialize the multi-view images acquired by the UAV to obtain preprocessed images, and then use the motion recovery structure algorithm to process the preprocessed images to obtain sparse point clouds. The second module is used to construct a multi-scale frequency-spatial joint regularization mechanism based on preprocessed images, including loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss. The third module is used to initialize a 3D Gaussian field based on sparse point clouds. It takes the total loss function, which integrates the loss of multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss, as the objective function, and iteratively updates the 3D Gaussian field parameters to output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

[0064] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0065] In an exemplary embodiment, the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the UAV 3D reconstruction method based on high-frequency detail enhancement as described in Embodiment 1 above.

[0066] In an exemplary embodiment, the readable storage medium may be a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the UAV 3D reconstruction method based on high-frequency detail enhancement as described in Embodiment 1 above.

[0067] In an exemplary embodiment, the computer program product includes a computer program that, when executed by a processor, implements the UAV 3D reconstruction method based on high-frequency detail enhancement as described in Embodiment 1 above.

[0068] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0069] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0070] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0071] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0072] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0073] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A UAV 3D reconstruction method based on high-frequency detail enhancement, characterized in that, Includes the following steps: The multi-view images acquired by the UAV are preprocessed and initialized to obtain preprocessed images. The preprocessed images are then processed using the structure-reconstruction-motion algorithm to obtain sparse point clouds. Based on preprocessed images, a multi-scale frequency-spatial joint regularization mechanism is constructed, along with loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss. The 3D Gaussian field is initialized based on sparse point clouds. The objective function is the total loss function that integrates the loss of multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss. The 3D Gaussian field parameters are iteratively updated to output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

2. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 1, characterized in that, A multi-scale frequency-spatial joint regularization mechanism is used to perform spatial and frequency domain analysis on preprocessed images to obtain spatial gradient magnitude and logarithmic magnitude spectrum. The specific methods for spatial domain analysis are as follows: The Sobel gradient operator is used to calculate the gradients of the preprocessed image in the x and y directions using a 3×3 convolution kernel. The square root of the sum of the squares of the gradients in the two directions is then taken to obtain the spatial gradient magnitude. The formula for calculating the spatial gradient magnitude is as follows: in, Indicates the magnitude of the spatial gradient. Indicates the partial sign. This indicates a pre-processed image; Using the Laplacian operator, the second derivative is calculated through a 3×3 convolution kernel to enhance sensitivity to high-frequency responses. The calculation formula for detecting subtle texture changes is as follows: in, Indicates the second-order differential; The specific methods of frequency domain analysis are as follows: The preprocessed image is downsampled three times to obtain a three-layer image pyramid. A two-dimensional discrete Fourier transform is then performed on each layer of the image pyramid to transform the preprocessed image from the spatial domain to the frequency domain, thus obtaining the frequency domain features. in, Indicates the first Frequency domain features corresponding to the layer image pyramid Indicates the index of frequency components in the horizontal direction of the frequency domain. Indicates the index of the frequency component in the vertical direction of the frequency domain. This indicates the image height within the image pyramid. This indicates the width of the image within the image pyramid. Represents the natural constant. To represent a complex unit, Represents pi; Indicates the first Layered image pyramid in ; Logarithmic transform of frequency domain features to construct logarithmic amplitude spectrum : in, This represents a logarithmic function with the natural constant as its base.

3. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 2, characterized in that, Loss of multi-scale frequency-spatial joint regularization The calculation formula is: in, Indicates the first The frequency domain characteristics of the image generated by the layered image pyramid through real-time rendering of a 3D Gaussian field in the current iteration. Indicates the first Frequency domain characteristics of real images of layered pyramids. This represents the frequency domain difference measurement function. This indicates the total number of layers in the image pyramid. The weight parameters represent the cross-scale consistency constraints. This represents the cross-scale frequency consistency constraint term. These are scale-weighted parameters.

4. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 3, characterized in that, The specific method for constructing the weighted photometric loss is as follows: The high-frequency region mask is obtained by fusing the spatial gradient magnitude and the logarithmic magnitude spectrum; Construct a spatial adaptive weight matrix based on the high-frequency region mask; Weighted photometric loss is constructed based on a spatially adaptive weight matrix. : in, This indicates the uniform height between the preprocessed image and the 3D Gaussian field rendered image. This indicates the uniform width between the preprocessed image and the 3D Gaussian field rendered image. Represents the spatial adaptive weight matrix. Represents the pixels of a 3D Gaussian field rendered image. Pixel value at that location, Represents actual image pixels The pixel value at that location.

5. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 1, characterized in that, Multilevel scale regularization loss is used to constrain the scale of Gaussian elements in a 3D Gaussian field. The calculation formula is: in, The weight parameters represent the regularization parameters. This represents the set of Gaussian elements corresponding to the high-frequency region. Represents the scale vector. Indicates the scale threshold. It represents the infinite norm.

6. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 1, characterized in that, The enhanced geometric consistency loss is constructed by fusing depth uncertainty and normal uncertainty, specifically as follows: In terms of depth uncertainty modeling, dense depth maps are generated based on image matching costs between sparse point clouds and preprocessed images. The Sobel operator is used to calculate the depth gradient magnitude of the dense depth map, and the depth gradient magnitude is mapped to a depth uncertainty score in the 0-1 interval. : in, Represents the Sobel operator. This represents the Sigmoid function. Represents a dense depth map; In terms of normal uncertainty modeling, the dot product of the normal vectors of each pixel in the preprocessed image with its neighboring pixels is calculated. The average of the dot products of each pixel in the preprocessed image with its neighboring pixels is then used to obtain the normal consistency score. This score is then combined with the gradient information from the rendering distance map to construct a composite normal uncertainty score. : in, Indicates the normal uniformity score. Represents the weight parameters. This represents the rendering distance map generated by real-time rendering of the 3D Gaussian field in the current iteration round; Based on depth uncertainty score and normal uncertainty fraction We construct an enhanced geometric consistency loss.

7. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 6, characterized in that, Enhanced geometric consistency loss The calculation formula is: in, This represents the geometric consistency mask, which takes a value of 1 if geometrically consistent, and 0 otherwise. Represents pixels in the preprocessed image The score of the depth of uncertainty at that point. Represents pixels in the preprocessed image The uncertainty fraction of the normal at that location. This represents the reprojection error when each pixel of the current image is projected into its neighborhood.

8. The UAV 3D reconstruction method based on high-frequency detail enhancement according to claim 1, characterized in that, Total loss function The specific formula is: in, Indicates weighted photometric loss. This represents the multi-scale frequency-spatial joint regularization loss. This represents the loss of multi-level scale regularization. This represents the loss for enhancing geometric consistency. The weights represent the weighted photometric loss. The weighted multi-scale frequency-spatial joint regularization loss is represented. The weights represent the multi-level scale regularization loss. This represents the weight of the loss for enhancing geometric consistency.

9. A UAV 3D reconstruction system based on high-frequency detail enhancement, characterized in that, include: The first module is used to preprocess and initialize the multi-view images acquired by the UAV to obtain preprocessed images, and then use the motion recovery structure algorithm to process the preprocessed images to obtain sparse point clouds. The second module is used to construct a multi-scale frequency-spatial joint regularization mechanism based on preprocessed images, including loss, weighted photometric loss, multi-level scale regularization loss, and enhanced geometric consistency loss. The third module is used to initialize a 3D Gaussian field based on sparse point clouds. It takes the total loss function, which integrates the loss of multi-scale frequency-spatial joint regularization mechanism, weighted photometric loss, multi-level scale regularization loss and enhanced geometric consistency loss, as the objective function, and iteratively updates the 3D Gaussian field parameters to output UAV 3D reconstruction results with enhanced high-frequency details and geometric consistency.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.