Single-view 3d reconstruction of large-scale outdoor scenes based on 3d gaussian splatting

By introducing a 3D Gaussian splashing method that incorporates panoramic consistency supervision, semantic constraints, and radially weighted photometric loss, the problem of balancing global and local consistency in large-scale outdoor scenes is solved, achieving efficient and high-fidelity 3D reconstruction results.

CN121074278BActive Publication Date: 2026-05-01HANGZHOU MAQUAN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU MAQUAN INFORMATION TECH CO LTD
Filing Date
2025-11-06
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing single-view 3D reconstruction methods struggle to balance global and local consistency in large-scale outdoor scenes, suffer from ambiguity in depth and texture estimation, and lack high-quality datasets, thus limiting the reconstruction performance of complex outdoor scenes.

Method used

By employing a 3D Gaussian splashing-based approach, combined with panoramic consistency supervision, semantic constraint depth regularization, and radial weighted photometric loss, and by constructing a virtual panoramic camera array and semantic segmentation, and introducing radial weighted photometric loss and Gaussian clipping mechanisms, we can achieve efficient and high-fidelity large-scale outdoor scene reconstruction.

Benefits of technology

It significantly improves global topological consistency and local geometric accuracy in large-scale outdoor scenes, reduces computational and storage overhead, adapts to reconstruction needs of different scales and scenes, and has good scalability and practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074278B_ABST
    Figure CN121074278B_ABST
Patent Text Reader

Abstract

The application discloses a single-view large-scale outdoor scene three-dimensional reconstruction method based on three-dimensional Gaussian splash, which collects pseudo-aerial images and constructs a panoramic multi-modal supervised end-to-end single-view three-dimensional reconstruction model, simultaneously introduces panoramic consistency supervision, semantic constraint depth regularization, radial weighting photometric loss and Gaussian clipping mechanism, effectively overcomes the defects of traditional single-view three-dimensional reconstruction in the aspect of insufficient geometric constraint, realizes efficient and high-fidelity three-dimensional modeling of a large-scale outdoor scene under single image input, and is suitable for smart city construction, automatic driving simulation, virtual reality / augmented reality, digital twinning and various practical scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and 3D reconstruction technology, specifically relating to a single-view large-scale outdoor scene 3D reconstruction method based on 3D Gaussian splashing. Background Technology

[0002] 3D reconstruction is a core task in computer vision and graphics. Its goal is to recover 3D geometric structure and scene information from 2D images. It is of great significance in applications such as autonomous driving, augmented reality and virtual reality, digital twins, smart city construction, cultural heritage protection, and disaster simulation.

[0003] Traditional multi-view stereo reconstruction (MVS) methods rely on intensive multi-angle image acquisition and achieve high-precision geometric modeling through triangulation. However, such methods are computationally expensive and impractical in scenarios with limited data acquisition. In contrast, single-view... Figure 3 3D reconstruction relies on only a single image input and can complete 3D reconstruction under conditions of scarce data or inability to take multiple-view shots, thus making it more valuable in practical applications.

[0004] In recent years, with the development of deep learning and generative models, single-view... Figure 3 Significant progress has been made in 2D reconstruction. For example, diffusion-based methods can utilize pre-trained 2D generation capabilities to synthesize new perspectives, while NeRF (Neural Radiance Fields)-based methods achieve high realism through neural implicit modeling. However, these methods generally suffer from problems such as high computational cost and insufficient geometric consistency.

[0005] 3DGS (3D Gaussian Splatting), as an emerging explicit scene representation method, combines real-time rendering speed with high reconstruction fidelity, showing great potential in 3D reconstruction, sparse view synthesis, interactive scene editing, and robotic SLAM (Simultaneous Localization and Mapping). Existing research has extended 3DGS to single-view reconstruction tasks with some success. For example, GaussianAnything combines point cloud latent representation, variational autoencoders, and diffusion models to generate high-quality 3D objects; TGS (Triplane-Gaussian Splatting) combines triplane representation with Gaussian rendering to further model geometric details; and PixelSplat alleviates the local optima problem caused by sparse representation by predicting dense 3D probability distributions and sampling Gaussian centers. However, most of these existing methods focus on the reconstruction of small-scale objects or finite scenes, and they suffer from the following prominent problems when dealing with complex, large-scale outdoor scenes:

[0006] (1) It is difficult to balance global and local consistency: Large-scale scenes usually contain multi-scale spatial structures and complex geometric components (such as building complexes, transportation networks, and vegetation). Traditional methods cannot simultaneously ensure the coherence of the overall topology and the accuracy of local geometry.

[0007] (2) Ambiguity in depth and texture estimation: In distant areas, due to the lack of parallax information and the degradation of texture features, the depth estimation error of a single view accumulates, resulting in structural distortion.

[0008] (3) Lack of high-quality large-scale datasets: Existing public datasets are mostly focused on small objects or indoor scenes, lacking coverage of large-scale urban environments and complex outdoor terrains, which limits the further development of related research. Summary of the Invention

[0009] In view of the above, the present invention provides a method for 3D reconstruction of large-scale outdoor scenes based on 3D Gaussian splashing in a single view. This method achieves efficient and high-fidelity 3D modeling of large-scale outdoor scenes with a single image input by introducing panoramic consistency supervision, semantic constraint depth regularization, and radial weighted photometric loss and Gaussian clipping mechanism.

[0010] A method for 3D reconstruction of large-scale outdoor scenes from a single view based on 3D Gaussian splashing includes the following steps:

[0011] (1) High-resolution pseudo aerial images were collected from multiple cities and complex scenes to construct a dataset. The camera perspective was systematically designed and the intrinsic and extrinsic parameters of the camera were extracted.

[0012] (2) Constructing a panoramic multimodal supervision end-to-end single-view system Figure 3 A 3D reconstruction model is used to directly map a single input image into a pixel-aligned 3D Gaussian set. Each Gaussian element in the set corresponds one-to-one with a pixel. The Gaussian element contains Gaussian parameters including center position coordinates, rotation matrix, scaling vector, color vector, and transparency.

[0013] (3) Using the dataset, the above single view is evaluated under the panoramic consistency supervision mechanism. Figure 3 The 3D reconstruction model is trained.

[0014] (4) Input the single image to be reconstructed into the trained single-view image. Figure 3 In the 3D reconstruction model, the corresponding 3D Gaussian set is obtained. Using this 3D Gaussian set and the camera's intrinsic and extrinsic parameters, rendering is performed to obtain the corresponding rendered images from multiple viewpoints, thereby completing the 3D image reconstruction.

[0015] Furthermore, in step (1), a 3D map visualization plugin is used to acquire high-resolution pseudo aerial images. The pseudo aerial images in the dataset cover various environmental structures, including densely built-up urban core areas, complex transportation hubs, industrial areas, natural scenic areas, and urban-rural fringe areas. During the image data acquisition process, the camera's height, pitch angle, and rotation angle are uniformly designed to ensure that the scene has multi-angle coverage and multi-scale features. On this basis, the images are input into COLMAP (image-based 3D reconstruction pipeline). The SfM (Structure-from-Motion) algorithm tool is used to achieve sparse point cloud reconstruction and camera intrinsic and extrinsic parameter estimation through feature point extraction and matching. Then, the MVS (Multi-View Stereo) algorithm tool is used to generate dense point clouds to supplement geometric details and correct camera pose, thereby obtaining accurate camera intrinsic and extrinsic parameter information.

[0016] Furthermore, the single view in step (2) Figure 3The 3D reconstruction model uses a hierarchical 3D U-Net structure as the backbone network, consisting of an encoder and a decoder. The encoder extracts high-level semantic information of the image step by step through multiple convolutions and downsampling, while the decoder restores the spatial details of the image through upsampling and skipping mechanisms. During downsampling, each scale is implemented by several convolutional modules in series. Each convolutional module contains a 3D convolutional layer, followed by batch normalization (BN), ReLU (corrected linear unit) activation function, and max pooling layer. An attention mechanism is added to the central, lower-resolution bottleneck layer to capture long-range spatial correlations. During upsampling, each scale is implemented by several deconvolutional modules in series and the corresponding encoded features are concatenated layer by layer. Each deconvolutional module contains a 3D deconvolutional layer, followed by batch normalization and ReLU activation function. The decoder output is finally predicted by a 1×1 convolution to obtain the Gaussian unit corresponding to each pixel.

[0017] Further, the specific implementation of step (3) is as follows: In the model training stage, a virtual panoramic camera array is first constructed based on the camera pose of the input image. The array includes four viewpoints: front, back, left, and right. The three-dimensional Gaussian set predicted by the model is rendered using the intrinsic and extrinsic parameters of the camera. The corresponding rendered images are generated under different viewpoints of the virtual panoramic camera. Then, a panoramic composite loss is constructed for each viewpoint by weighted summation of depth loss, which is composed of light consistency, LPIPS (Learned Perceptual Image Patch Similarity), and semantic perception. The total loss function of the model is obtained by averaging the panoramic composite loss of all viewpoints. Finally, the model parameters are iteratively updated using gradient descent based on the total loss function until the loss function converges and the training is completed.

[0018] Furthermore, the photometric consistency is calculated using a pixel-level weighted squared error between the rendered image and the input image, specifically expressed as follows:

[0019]

[0020] in: I Represents the original input image. This represents the rendered image obtained by rendering the 3D Gaussian set predicted by the model. Represents the coordinates of any pixel in the image. Represents the original input image I median coordinate The corresponding pixel value, Indicates the rendered image median coordinate The corresponding pixel value, coordinates Radial weight of the pixel.

[0021] Furthermore, the semantically aware depth loss first uses a semantic segmentation model to segment the image, distinguishing foreground and background regions through masks. A boundary consistency constraint mechanism is employed at the foreground-background boundary, introducing edge detection or gradient constraints to ensure depth continuity in transition regions. Then, the depth loss between the rendered image and the input image is calculated using the following expression. :

[0022]

[0023] in: This refers to the error between the depth map corresponding to the rendered image and the depth map corresponding to the input image. a 1 and a 2 represents the depth loss weight coefficients corresponding to the background and foreground, and a 1 < a 2, N bg This represents the number of pixels in the background region of the image. N fg This represents the number of pixels in the foreground region of the image. This represents pixels with a value of 0 in the masked image, i.e., the background region. This represents pixels with a value of 1 in the masked image, i.e., for the foreground region.

[0024] Furthermore, the radial weight The calculation expression is as follows:

[0025]

[0026] in: Representing coordinates The radial distance from a pixel to the center of the image. r 1 and r 2 is the set distance threshold and r 1 < r 2, α , β , γ For the set weight value and α > β > γ .

[0027] Furthermore, a gating coefficient is introduced into the Gaussian element. During the rendering process, all Gaussian parameters in the primitives are multiplied by the gating coefficient before being used for rendering. At the same time, a sparsification regularization term of the gating coefficient is added to the total loss function to dynamically suppress redundant or low-contribution Gaussian primitives. The specific expression is as follows:

[0028]

[0029] in: For single vision Figure 3 The total loss function of the 3D reconstruction model, The mean of the panoramic composite loss across all viewpoints. coordinates The gating coefficient introduced by the Gaussian byte corresponding to the pixel. λ The weighting coefficients are set.

[0030] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method for single-view large-scale outdoor scene 3D reconstruction based on 3D Gaussian splashing.

[0031] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described method for single-view large-scale outdoor scene 3D reconstruction based on 3D Gaussian splashing.

[0032] Compared with the prior art, the present invention has the following beneficial technical effects:

[0033] 1. Overcoming the shortcomings of single-view geometric constraints: This invention introduces a virtual panoramic supervision mechanism under single-view input conditions. By constructing a multi-directional virtual camera array and applying cyclic consistency constraints, it simulates multi-view supervision conditions, effectively overcoming the limitations of traditional single-view... Figure 3 The shortcomings of 3D reconstruction in terms of insufficient geometric constraints are significantly improved, thereby enhancing the global topological consistency and structural rationality of large-scale outdoor scenes.

[0034] 2. Enhanced local geometry and boundary characterization capabilities: Through semantic segmentation-guided deep regularization, this invention can focus on constraining the boundaries and local geometric details of foreground objects in complex scenes, thereby ensuring the geometric accuracy of structural targets such as buildings, roads, and vegetation, while suppressing the uncertainty of the background area and reducing boundary ambiguity and structural errors that occur during deep reasoning.

[0035] 3. Balancing Reconstruction Accuracy and Computational Efficiency: The radial weighted photometric loss mechanism proposed in this invention enables the model to concentrate more optimization resources on key areas with dense geometric information during training, thereby improving the overall reconstruction accuracy. At the same time, combined with the Gaussian pruning mechanism, it effectively reduces low-contribution and redundant Gaussian units, reducing computational and storage overhead. Thus, while ensuring high-precision reconstruction, it achieves high efficiency in the inference and rendering processes.

[0036] 4. Excellent scalability and practicality: The method of this invention adopts a modular design, and the core components (such as panoramic supervision, semantic regularization, radial weighting, and sparsity mechanism) can be independently extended or replaced to adapt to the 3D reconstruction needs of different scales and scenarios. In terms of hardware, this invention can run efficiently on consumer-grade GPU (graphics processing unit) devices, has good deployability and application value, and is suitable for various practical scenarios such as smart city construction, autonomous driving simulation, virtual reality / augmented reality, and digital twins. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the process of the single-view large-scale outdoor scene 3D reconstruction method based on 3D Gaussian splashing of the present invention.

[0038] Figure 2 This invention provides an end-to-end single-view panoramic multimodal supervision system. Figure 3 A schematic diagram of the network structure of the 3D reconstruction model.

[0039] Figure 3 This is a schematic diagram of the depth loss calculation process for semantic awareness in this invention. Detailed Implementation

[0040] To describe the present invention in more detail, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] like Figure 1 As shown in the figure, this embodiment provides a method for 3D reconstruction of large-scale outdoor scenes from a single view based on 3D Gaussian splashing. The specific implementation process is as follows:

[0042] (1) This implementation first obtains high-resolution pseudo aerial images from multiple typical cities and complex scenes through a 3D map visualization plugin, covering a variety of environmental structures, including densely built-up urban core areas, complex transportation hubs, industrial areas, natural landscapes, and urban-rural fringe areas, thereby ensuring that the dataset has diversity and wide applicability.

[0043] During data acquisition, the camera's height, pitch angle, and rotation angle are systematically designed to ensure uniform distribution of the acquired images in both the horizontal and vertical directions, avoiding single-angle or scale deviations and guaranteeing omnidirectional scene coverage and multi-scale features. Subsequently, the acquired images are input into the structured motion reconstruction algorithm COLMAP, utilizing feature point extraction and matching to achieve sparse point cloud reconstruction and camera extrinsic parameter estimation. Furthermore, a multi-view stereo method is combined to generate a dense point cloud to supplement geometric details and correct camera pose, thereby obtaining accurate camera extrinsic and extrinsic parameter information. Through the above process, this implementation constructs a high-quality, large-scale dataset. This dataset not only contains multi-angle, multi-scale image samples but also has rigorously registered camera parameters, suitable for single-view... Figure 3 The reconstruction task provides a reliable training and validation foundation.

[0044] (2) Constructing a panoramic multimodal supervision end-to-end single-view Figure 3 3D reconstruction model.

[0045] single vision Figure 3 The core of the 3D reconstruction method lies in the end-to-end model architecture design. Its goal is to directly map a single input image into a pixel-aligned 3D Gaussian set and generate a 3D representation of the scene through a differentiable renderer. This model employs an end-to-end neural network architecture. The main body of the model uses a hierarchical encoder-decoder structure, achieving the transformation from a 2D input to a 3D Gaussian field through multi-scale feature extraction and pixel-wise Gaussian parameter regression. Each pixel location corresponds to a Gaussian primitive, which is described by center coordinates, covariance matrix, transparency coefficient, and color vector parameters.

[0046] like Figure 2 As shown, this implementation uses a layered 3D U-Net structure as the backbone network of the model. The encoding part gradually extracts high-level semantic information through multi-layer convolution and downsampling, while the decoding part restores spatial details through upsampling and jump connection mechanisms. This structure can preserve local geometric features while maintaining global topological consistency.

[0047] The overall structure includes the following modules:

[0048] U-shaped architecture: a symmetrical structure of downsampling, centering, and upsampling, with jumpers to preserve high-resolution details.

[0049] The residual block (ResBlock) is the basic computational unit: each scale consists of several ResBlocks concatenated, and each ResBlock contains normalization, activation, and convolution.

[0050] Self-Attention Module: An attention module is added to the intermediate, lower-resolution layers to capture long-range spatial correlations.

[0051] The specific parameter configuration details are as follows:

[0052] The number of channels in the encoding part is 64, 128, 256, and 512 respectively, and the decoding part upsamples layer by layer and splices the corresponding encoded features.

[0053] At the decoder output, the Gaussian meta-parameter set corresponding to each pixel is predicted by 1×1 convolution. The parameter dimension is preferably 14, including: the three-dimensional center position. Rotation matrix represented by quaternions Scaling vector Color vector ,transparency σ Gating coefficient m (Mapped to [0,1] via Sigmoid activation).

[0054] (3) Introduce a panoramic consistency supervision mechanism.

[0055] This implementation introduces a panoramic consistency supervision mechanism under single-view input conditions to compensate for the shortcomings of traditional single-view input. Figure 3 The problem of insufficient geometric constraints in 3D reconstruction is addressed specifically by: firstly, during the training phase, constructing a virtual panoramic camera array based on the camera pose of the input image for the generated 3D Gaussian field. The array includes four directions: forward, backward, left, and right, thus achieving approximately 360-degree coverage of the scene. Then, the 3D Gaussian set predicted by the model is input into a differentiable Gaussian renderer to generate corresponding rendered images from different perspectives of the virtual panoramic camera. and compared with the original input image from the corresponding viewpoint I By comparing these, a cyclic consistency loss is formed, which means that a multimodal supervision is constructed for each viewpoint, including photometric consistency (L1 loss), LPIPS, and semantic awareness depth loss.

[0056] The photometric consistency term uses pixel-level weighted squared error:

[0057]

[0058] LPIPS extracts image features using the deep learning model VGG (Visual Geometry Group Network) and then calculates the distance between these features to evaluate the perceptual similarity between images.

[0059]

[0060] The semantic perception depth loss is represented as The above components are combined at each virtual viewpoint to define a single-viewpoint panoramic composite loss:

[0061]

[0062] The overall training reconstruction loss is defined as:

[0063]

[0064] This mechanism simulates multi-view supervision under single-view input conditions, effectively constraining the scene geometry and improving global topological consistency.

[0065] (4) Introduce a deep regularization mechanism with semantic constraints.

[0066] This invention utilizes a semantic segmentation model to generate foreground and background region masks for a scene (providing a basis for region division for subsequent depth constraints), and introduces these masks into the depth estimation constraints. For foreground target regions (including structural targets such as buildings, roads, vehicles, and vegetation), higher weights are applied to maintain clear object boundaries and accurate local geometry. For background regions (usually the sky or distant environment), scene consistency is maintained through overall depth constraints. This regularization mechanism significantly alleviates the depth ambiguity problem caused by distant texture degradation and disparity loss, and improves the geometric representation ability of complex scenes.

[0067] like Figure 3 As shown, the goal of semantic awareness deep regularization in this embodiment is to tightly couple semantic information (instance / semantic mask and its confidence) with deep supervision / regularization, thereby achieving single-view... Figure 3 The specific process for improving foreground geometric accuracy, maintaining clear boundaries, and suppressing depth noise in DGS reconstruction is as follows:

[0068] First, the input image is processed using the semantic / instance segmenter SAM2. I Extracting instance masks and confidence scores, specifically: SAM2 will output... K Instance segmentation results M k and the corresponding IoU confidence score of the instance. Boundary stability confidence score The IoU (Intersection over Union) confidence score reflects the degree of matching between the mask and the real target region (or internal consistency). The larger the value, the more reliable the mask is. The boundary stability confidence score reflects whether the mask boundary is clear and whether there is ambiguity or noise.

[0069] Next, the segmentation results for each instance are filtered by confidence and a fusion mask is constructed:

[0070]

[0071] in: Indicates the first k An instance mask in pixels The value at that location is either 0 or 1. This is an indicator function that takes the value 1 if the condition is true, and 0 otherwise. This represents a pixel-by-pixel logical OR operation on all instance masks; and , where is the confidence level.

[0072] Finally, the fused semantic segmentation map is obtained. Where 0 indicates that the corresponding pixel belongs to the background area, thus obtaining the number of background pixels. 1 indicates that the corresponding pixel belongs to the foreground region, and the number of foreground pixels is... .

[0073] Then, based on the segmentation results of the semantic model, the depth loss for semantic perception is set. The details are as follows:

[0074]

[0075] in: The error between the rendered depth map and the actual depth map. a 1 and a 2 represents the depth loss weight coefficients corresponding to the background and foreground.

[0076] (5) Introduce a radial weighted photometric loss mechanism.

[0077] In large-scale outdoor scenes, the central region of an input single image typically contains richer and more reliable geometric cues (building facades, road centers, etc.), while the peripheral regions far from the image center are often distant views, occluded areas, or areas of texture degradation, making single-view depth and texture information less reliable in these regions. This invention's radially weighted photometric loss assigns weights to pixels according to their radial distance from the image center, making training more focused on the central region with stronger geometric cues, thereby improving the reconstruction accuracy of key regions and suppressing the interference of peripheral noise on optimization. During reconstruction, different weights are assigned based on the radial distance of pixels to the image center, with higher weights given to the central region to ensure the reconstruction accuracy of key structures, while constraints are appropriately relaxed in the edge regions to reduce accumulated errors.

[0078] For the input image, this embodiment defines a normalized radial distance between a pixel and the center, assuming the image resolution is... The length of the image diagonal is Let the center of the image be... For pixels Its normalized radial distance is:

[0079]

[0080] This normalized distance provides a metric for subsequent weighting, enabling pixels at different locations to receive differentiated weights based on their spatial position in the image.

[0081] In photometric consistency supervision, this implementation divides pixels into different regions by setting several radial thresholds and assigning piecewise constant weights to each region, thereby achieving differentiated constraints on the image center and edge regions. Specifically: a radially weighted photometric loss is defined, a radial distance threshold from a pixel to the image center is set, pixels are classified, and piecewise constant weights are set. Specifically, thresholds are set... r 1 and r 2 and weights The photometric loss weights are set as follows:

[0082]

[0083] The weighted radial weighted photometric loss is:

[0084]

[0085] in: To render the image, For real images, For radial weights.

[0086] This mechanism can mitigate the impact of noise in edge regions while ensuring the reconstruction accuracy of key areas. The final weighted photometric loss consists of the difference between the rendered image and the real image, and combined with radial weights, it enables fine control over the reconstruction results.

[0087] (6) Introduce Gaussian clipping mechanism.

[0088] The network model predicts a Gaussian primitive for each pixel, which typically results in a large and dense set of Gaussian primitives. However, most Gaussian primitives contribute little or are redundant in the final representation. Directly rendering all Gaussian primitives would waste computation and memory and may introduce noise. This implementation introduces Gaussian clipping, which uses gating mechanisms and sparsity regularization to dynamically suppress or remove low-contribution Gaussian primitives, improving running efficiency and reducing texture "leakage". Specifically, during the prediction process, a gating weight is assigned to each pixel's Gaussian primitive. The final actual parameters of the Gaussian element are:

[0089]

[0090] in: These are the original Gaussian parameters for direct regression by the network. For the final Gaussian parameters.

[0091] Add an L1 sparse regularization constraint to the weight in the final loss. m Approaching zero to eliminate the influence of unimportant Gaussian elements, ultimately only retaining the Gaussian set that significantly contributes to the rendering result, as specifically expressed by:

[0092]

[0093] This mechanism dynamically suppresses redundant or low-contribution Gaussian elements by using gating coefficients and sparsification regularization terms, thereby reducing computational burden and improving model efficiency.

[0094] The above description of the embodiments is provided to enable those skilled in the art to understand and apply the present invention. Those skilled in the art can readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made to the present invention by those skilled in the art based on the disclosure thereof should be within the scope of protection of the present invention.

Claims

1. A method for 3D reconstruction of large-scale outdoor scenes from a single view based on 3D Gaussian splashing, characterized in that, Includes the following steps: (1) High-resolution pseudo aerial images were collected from multiple cities and complex scenes to construct a dataset. The camera perspective was systematically designed and the intrinsic and extrinsic parameters of the camera were extracted. (2) Construct a panoramic multimodal supervised end-to-end single-view 3D reconstruction model to directly map a single input image into a pixel-aligned 3D Gaussian set. Each Gaussian element in the set corresponds to a pixel. The Gaussian element includes Gaussian parameters such as center position coordinates, rotation matrix, scaling vector, color vector, and transparency. The single-view 3D reconstruction model uses a hierarchical 3D U-Net structure as the backbone network, which consists of an encoder and a decoder. The encoder extracts high-level semantic information of the image step by step through multi-layer convolution and downsampling, while the decoder restores the spatial details of the image through upsampling and skipping mechanisms. During the downsampling process, each scale is implemented by several convolutional modules in series. Each convolutional module contains a 3D convolutional layer, followed by batch normalization, ReLU activation function, and max pooling layer. An attention mechanism is added to the central, lower-resolution bottleneck layer to capture long-range spatial correlations. During the upsampling process, each scale is implemented by several deconvolution modules in series and the corresponding encoded features are concatenated layer by layer. Each deconvolution module contains a 3D deconvolution layer, followed by batch normalization and ReLU activation function. The output of the decoder is finally predicted by a 1×1 convolution to obtain the Gaussian unit corresponding to each pixel. (3) The above single-view 3D reconstruction model is trained using the dataset under the panoramic consistency supervision mechanism. During the model training phase, a virtual panoramic camera array is first constructed based on the camera pose of the input image. The array includes four viewpoints: front, back, left, and right. The 3D Gaussian set predicted by the model is rendered using the intrinsic and extrinsic parameters of the camera. The corresponding rendered images are generated under different viewpoints of the virtual panoramic camera. Then, a panoramic composite loss is constructed for each viewpoint by weighted summation of the depth loss of photometric consistency, LPIPS, and semantic awareness. The total loss function of the model is obtained by averaging the panoramic composite loss of all viewpoints. Finally, the model parameters are iteratively updated using the gradient descent method according to the total loss function until the loss function converges and the training is completed. The photometric consistency is calculated using a pixel-level weighted squared error between the rendered image and the input image, as shown in the following expression: Where: I represents the original input image. Let I(i,j) represent the rendered image obtained by rendering the 3D Gaussian set predicted by the model, where (i,j) represents the coordinates of any pixel in the image, and I(i,j) represents the pixel value corresponding to coordinate (i,j) in the original input image I. Indicates the rendered image The pixel value corresponding to coordinate (i,j) is defined by W(i,j), which is the radial weight of the pixel at coordinate (i,j). r(i,j) represents the radial distance from the pixel at coordinate (i,j) to the center of the image. r1 and r2 are set distance thresholds, and r1 < r2. α, β, and γ are set weight values, and α > β > γ. The semantically aware depth loss first uses a semantic segmentation model to segment the image, using masks to distinguish foreground and background regions. A boundary consistency constraint mechanism is employed at the foreground-background boundary, introducing edge detection or gradient constraints to ensure depth continuity in transition regions. Then, the depth loss L between the rendered image and the input image is calculated using the following expression. depth : Where: ΔD 2 The depth map of the rendered image is the difference between the depth map of the input image and the depth map of the rendered image. a1 and a2 are the depth loss weight coefficients for the background and foreground, respectively, with a1 < a2. N bg N represents the number of pixels in the background region of the image. fg M represents the number of pixels in the foreground region of the image. fused =0 represents pixels with a value of 0 in the masked image, i.e., for the background region, M fused =1 indicates a pixel with a value of 1 in the masked image, i.e., for the foreground region; The Gaussian primitives also incorporate a gating coefficient m(x,y)∈[0,1]. During rendering, all Gaussian parameters in the primitives are multiplied by this gating coefficient before being used for rendering. Simultaneously, a sparsity regularization term for this gating coefficient is added to the total loss function to dynamically suppress redundant or low-contribution Gaussian primitives. The specific expression is as follows: Where: L final L is the total loss function for a single-view 3D reconstruction model. reconstruction λ is the mean of the panoramic composite loss for all views, m(i,j) is the gating coefficient introduced by the Gaussian unit corresponding to the pixel at coordinate (i,j), and λ is the set weight coefficient. (4) Input the single image to be reconstructed into the trained single-view 3D reconstruction model to obtain the corresponding 3D Gaussian set. Use the 3D Gaussian set and the camera's intrinsic and extrinsic parameters to render the corresponding rendered images from multiple viewpoints, thereby completing the 3D image reconstruction.

2. The method for single-view large-scale outdoor scene 3D reconstruction based on 3D Gaussian splashing according to claim 1, characterized in that: In step (1), a 3D map visualization plugin is used to acquire high-resolution pseudo aerial images. The pseudo aerial images in the dataset cover a variety of environmental structures, including densely built-up urban core areas, complex transportation hubs, industrial areas, natural scenic areas, and urban-rural fringe areas. During the image data acquisition process, the camera's height, pitch angle, and rotation angle are uniformly designed to ensure that the scene has multi-angle coverage and multi-scale features. Based on this, the images are input into COLMAP. The SfM algorithm tool is used to achieve sparse point cloud reconstruction and camera intrinsic and extrinsic parameter estimation through feature point extraction and matching. Then, the MVS algorithm tool is used to generate dense point clouds to supplement geometric details and correct camera pose, thereby obtaining accurate camera intrinsic and extrinsic parameter information.

3. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: The processor is used to execute the computer program to implement the single-view large-scale outdoor scene 3D reconstruction method based on 3D Gaussian splashing as described in any one of claims 1 to 2.

4. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the single-view large-scale outdoor scene 3D reconstruction method based on 3D Gaussian splashing as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Sparse visual angle three-dimensional reconstruction method based on depth prior information

    CN118657888A

  • Gaussian splashing monocular reconstruction method based on neural network

    CN119904580A

  • Interactive three-dimensional Gaussian editing method and system based on three-dimensional geometry consistent attention priori

    CN120411438A