Large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics
By employing a block-based training and hierarchical BVH tree construction method, the problem of insufficient multimodal information integration in large-scale scene reconstruction is solved. This method achieves efficient integration and real-time rendering of RGB and semantic information, reduces video memory consumption, and supports ultra-large-scale scene reconstruction.
Patent Information
- Application Number
- CN202510727707.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies suffer from insufficient multimodal information integration, high memory consumption, and artifacts in block merging during large-scale scene reconstruction, making it difficult to support ultra-large-scale scene reconstruction and lacking real-time rendering and semantic editing capabilities.
By employing a block-based training and hierarchical BVH tree construction method, and through Gaussian pixel density block partitioning, BVH tree sparsification, and granularity adjustment, unified training and rendering of RGB and semantic information are achieved, reducing memory consumption and improving rendering efficiency.
It achieves efficient integration of RGB and semantic information in large-scale scene reconstruction, reduces video memory requirements, supports real-time rendering and consistent 3D scene manipulation, and solves the problems of video memory explosion and block merging artifacts in traditional methods.
Smart Images

Figure CN120833437A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of three-dimensional computer vision and graphics, in particular to a large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics, which is suitable for large-scale scene modeling, virtual reality content generation and automatic driving environment perception. BACKGROUND
[0002] In recent years, with the rapid development of metaverse and digital twin technology, efficient reconstruction and semantic understanding of large-scale three-dimensional scenes have become a core challenge of digitalizing the real world. The neural radiance field (NeRF) technology proposed in 2020 first introduced implicit neural representation into three-dimensional reconstruction, and realized high-quality new view synthesis through multi-view images. However, traditional NeRF-based methods have significant limitations. They rely on volume rendering of multi-layer perceptron (MLP), and single-scene training takes several hours to several days. They also have scene size limitations, making it difficult to support large-scale open scenes (such as city-level reconstruction). The memory occupancy increases exponentially with the complexity of the scene, and they also lack semantic editing functions. Although subsequent work such as Semantic-NeRF (2021) supports semantic rendering, it lacks direct semantic editing capabilities for three-dimensional implicit scenes. Existing methods such as Object-NeRF require separate MLP modeling for each object, resulting in computational redundancy and difficulty in generalization.
[0003] To overcome the above limitations, researchers have begun to explore new paradigms of explicit three-dimensional representation and semantic fusion. 3D Gaussian Splatting (3DGS) proposed in 2023 parameterizes the scene through explicit Gaussian distribution, realizing real-time rendering (>100FPS) and dynamic level of detail (LOD) control. Further, frameworks such as HUGS (2024) integrate semantic and dynamic object modeling into 3DGS, demonstrating the feasibility of multi-modal joint optimization. However, existing technologies still face key challenges. In large-scale scene reconstruction, traditional 3DGS faces memory explosion and block merging artifacts on ultra-large-scale datasets (such as several kilometers of city scenes). Therefore, there is an urgent need for an RGB and semantic supervision unified framework that supports ultra-large-scale scene reconstruction, maintains real-time rendering efficiency, and realizes consistent three-dimensional scene manipulation. SUMMARY
[0004] The present application aims to provide a large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics, to solve the problem of insufficient multi-modal information integration in large-scale reconstruction in the prior art.
[0005] To achieve the above-mentioned purpose, the technical solutions adopted by the present application are as follows:
[0006] According to a first aspect of the present specification, a large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics is provided, which comprises the following steps:
[0007] (1) input multi-view RGB images of a target scene and perform data preprocessing, divide the target scene by point cloud density, and standardize the image, depth, and semantic data in each block;
[0008] (2) generalization training stage, coarse training of the target scene, training 3D Gaussian primitives of the entire scene based on global point cloud without using Gaussian densification, obtaining a Gaussian primitive set;
[0009] (3) block-independent training and hierarchical BVH tree construction stage, the following operations are performed on each block:
[0010] (3.1) training Gaussian primitives in a certain block using Gaussian densification;
[0011] (3.2) using an AABB type bounding box for all Gaussian primitives in the block, constructing a BVH tree using median division method, taking the Gaussian primitives generated during coarse training as leaf nodes, and simultaneously standardizing the rotation direction of all Gaussian primitives to be consistent;
[0012] (3.3) re-densification training of the block to correct the standardized rotation coordinate axis of each Gaussian primitive;
[0013] (4) hierarchical structure merging stage, merging the leaf nodes of all blocks to generate internal nodes of the BVH tree, the internal nodes have all the attributes of Gaussian primitives, and the mean, covariance, SH coefficient and opacity of each internal node are obtained by minimizing the weighted 3D KL divergence;
[0014] (5) rendering stage, optimizing and improving local detail quality, specifically:
[0015] (5.1) selecting rendering nodes in all nodes of the BVH tree, including: defining a standard random variable and calculating to obtain a target granularity; if the granularity of a node under a given view is less than the target granularity, but the granularity of its parent node is greater than the target granularity, the node is selected; if the node does not meet the target granularity requirement in the current view after switching the view, a new node is generated between the node and its parent node through interpolation;
[0016] (5.2) sparsifying the generated BVH tree;
[0017] (5.3) scanning all nodes of the BVH tree, if the distance from a Gaussian primitive belonging to a certain block to another block is less than the distance to the block, the Gaussian primitive is deleted;
[0018] (5.4) output RGB rendering image and semantic rendering image.
[0019] Further, step (1) is specifically:
[0020] (1.1) input multi-view RGB images of the target scene, and obtain pose information;
[0021] (1.2) generate a depth map, a semantic mask, and a global point cloud Wherein p i is the i-th point, and N is the number of points in the global point cloud;
[0022] (1.3) correct the positional deviation of points under different views by point triangulation and bundle adjustment method;
[0023] (1.4) create a sky box, place a number of 3D Gaussian primitives on a sphere with several times the diameter of the scene, and render the sky to simulate an infinite scene;
[0024] (1.5) based on the density of the global point cloud , actively block to obtain a block set Wherein C k represents the k-th block, and N' represents the total number of blocks obtained by blocking.
[0025] Further, in step (2), based on the global point cloud , the 3D Gaussian primitives of the entire scene are trained without using Gaussian densification, and the obtained Gaussian primitive set is denoted as Wherein represents the j-th Gaussian primitive, is a mean vector representing position, is a covariance matrix representing shape, is an SH coefficient vector representing color, s j is a semantic label, α j is the opacity, and M is the number of Gaussian primitives.
[0026] Further, the loss functions of the coarse training stage and the sub-block independent training and hierarchical BVH tree construction stage all adopt a multi-modal loss function of joint RGB reconstruction loss and semantic cross-entropy loss , λ S is a weight coefficient, wherein represents a rendered image, represents ground truth, λ SSIM is a hyperparameter, and SSIM(·) is a structural similarity loss function, wherein is a semantic pseudo ground truth, and Sk is the trained semantic, S is the total number of semantic labels.
[0027] Further, in step (4), the calculation formulas of the mean, covariance, SH coefficient and opacity of the internal node are as follows:
[0028]
[0029] wherein, l represents the lth layer of the BVH tree, μ (l+1) ,∑ (l+1) ,SH (l+1) and α (l+1) respectively represent the mean, covariance, SH coefficient and opacity of a certain internal node of the l+1th layer, E is the total number of all child nodes of the internal node, and respectively represent the mean, covariance, SH coefficient and opacity of the i th child node of the internal node, w i is the normalized weight.
[0030] Further, the new node is generated between the node and its parent node by interpolation, specifically:
[0031] ① Two Gaussian primitives with the same properties as the parent node are generated, and the properties of the two Gaussian primitives, the mean μ and the SH coefficient SH are interpolated respectively, and the interpolation weight t is calculated, wherein τ ∈ is the target granularity, ∈(n), ∈(p) are the granularities of the node n and its parent node respectively, and the mean and the SH coefficient of the node n are directly multiplied by the interpolation weight to obtain the mean and the SH coefficient of the new node new;
[0032] ② Since the covariance matrix ∑ can be decomposed into the combination of the scaling matrix S and the rotation matrix R, ∑=(RS)(RS) T , the scaling matrix S and the rotation matrix R are multiplied by the interpolation weight t n respectively, and the covariance matrix of the new node new is calculated;
[0033] ③ The opacity of the two new nodes is calculated by the formula α=t n α n +(1-t n )α′, wherein α n is the opacity of the node n, α p is the opacity of the parent node.
[0034] Further, the generated BVH tree is sparsified, specifically:
[0035] ①Mark the leaf nodes, traverse all training images with 3 pixels as the initial granularity to generate a candidate node set Cut1, and select the bottom layer nodes from Cut1 to form a set Cut2;
[0036] ②Delete the redundant node layers between the leaf nodes and Cut2;
[0037] ③Increase the target granularity successively, and after each granularity update, repeat the operations of Cut set generation, bottom layer node screening and intermediate node deletion until the target granularity is expanded to 50% of the original image size.
[0038] According to a second aspect of the present specification, an electronic device is provided, comprising a memory and a processor, the memory is coupled with the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to realize the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics as described in the first aspect.
[0039] According to a third aspect of the present specification, a computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to realize the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics as described in the first aspect.
[0040] According to a fourth aspect of the present specification, a computer program product is provided, comprising computer programs / instructions, which are executed by a processor to realize the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics as described in the first aspect.
[0041] The beneficial effects of the present application are: the present application is based on the large-scale three-dimensional scene reconstruction model capable of simultaneously training scene RGB and semantics, which can not only perform large-scale scene reconstruction, but also obtain multi-modal information such as scene, semantics and depth image during training; the method mainly includes preprocessing and dividing blocks, coarse training processing, independent training of each block and generating a hierarchical Gaussian distribution BVH tree, post-training processing, hierarchical structure merging, rendering and other stages. In the preprocessing and block division, the present method generates semantic, depth and point cloud information using existing tools, and then divides the blocks according to the input image density; in the coarse training processing stage, the present method does not perform branch pruning of Gaussian primitives; in the process of block training, the present application can perform branch pruning of Gaussian primitives; in the post-processing training of correcting rotation, the present method does not perform branch pruning of Gaussian primitives; in the rendering, the present method adjusts the target granularity to reduce the number of Gaussian primitives for rendering without affecting the visual effect, greatly reducing the consumption of video memory computing power. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0043] Figure 1 A flow chart of a large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics provided for an embodiment;
[0044] Figure 2 A three-dimensional scene reconstruction graph provided for an embodiment;
[0045] Figure 3 A three-dimensional scene semantic graph provided for an embodiment;
[0046] Figure 4 A structural schematic diagram of an electronic device shown for an exemplary embodiment. DETAILED DESCRIPTION
[0047] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the drawings.
[0048] It should be clear that the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0049] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0050] The present application mainly based on computer vision, image processing and other theories and technologies, proposes a large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics. The present application converts the input RGB image into point cloud, semantic and depth information, uses block to simultaneously train and supervise the three, and finally outputs the large-scale three-dimensional scene reconstruction result of the three, which can be applied to automatic driving simulation platform construction and other scenes. As shown in the figure, the method comprises the following steps: Figure 1
[0051] (1) Block pre-processing and parameter initialization stage, the target scene is divided into blocks by point cloud density, and the image, depth and semantic data in each block are standardized, including:
[0052] (1.1) Input multi-view RGB images of the target scene, use calibration algorithm to calibrate the camera (Calibrated Cameras) to obtain the pose, if the dataset comes with the pose information of the camera, directly use the pose information of the dataset, no additional calibration is needed;
[0053] (1.2) Use the depth map generation tool Depth-Anything-V2 to generate depth map, use the semantic information generation tool InverseForm to generate semantic mask, and use the point cloud generation tool (colmap) to generate global point cloud Wherein p i is the ith point, N is the number of points in the global point cloud, and colmap uses the 'hierarchical_mapper' parameter to parallelize reconstruction to adapt to subsequent block operation while reducing memory consumption;
[0054] (1.3) In order to improve the rendering quality, the position of the point cloud generated by colmap is refined, and the point cloud is preprocessed, which is to correct the position deviation of the points under different views through point triangulation and bundle adjustment method;
[0055] (1.4) Create a sky box, place 100k 3D Gaussian primitives on a sphere with ten times the diameter of the scene, render the sky to simulate an infinite scene, while improving rendering efficiency, reducing resource consumption and creating seamless visual effects;
[0056] (1.5) Based on the density of global point cloud Active block to get block set Where C k represents the kth block, N' represents the total number of blocks obtained by block, such as dense scene data collected when walking is divided into 50*50m 2 size, sparse scene data collected when driving is divided into 100*100m 2 size, which is beneficial to balance the calculation efficiency and detail retention, and avoid too fine or too coarse block.
[0057] (2) Generalization training phase, coarse optimization training for target scene, specifically:
[0058] Based on global point cloud Do not use Gaussian densification (clone, split) to train 3D Gaussian primitives of the whole scene, get Gaussian primitive set Wherein represents the jth Gaussian primitive, is the mean vector representing the position, the mean is the center of the Gaussian primitive, is a covariance matrix representing shape, is an SH coefficient vector representing color, s j is a semantic label, a j is opacity, M is the number of Gaussian primitives; loss function of the coarse training stage joint RGB reconstruction loss, i.e., rendering error (where represents a rendered image, represents ground truth, l SSIM is a hyperparameter, SSIM(·) is a structural similarity loss function) and semantic cross-entropy loss (where is a semantic pseudo ground truth, used as a real ground truth to supervise training, S k is a trained semantic, S is the total number of semantic labels), and the formula is where l S is a weight coefficient, which is set to 0.01 in the embodiment.
[0059] (3) Sub-block independent training and hierarchical BVH (Bounding Volume Hierarchy) tree construction stage, sub-block training allows distributed computing, significantly reduces the memory requirement, and the BVH tree supports dynamic detail selection (LOD), improves rendering efficiency. Specifically, the present application performs the following operations on each block C k :
[0060] (3.1) Train the Gaussian primitives in a certain block using the method of Gaussian densification (using cloning, splitting and other techniques), and the loss function adopts a multi-modal structure of RGB reconstruction loss and semantic cross-entropy loss, the formula is the same as step (2);
[0061] (3.2) Use an AABB type bounding box for all Gaussian primitives in the block, and construct a BVH tree using the median division method, and the Gaussian primitives generated during the coarse training are used as leaf nodes, and the rotation directions of all Gaussian primitives are calibrated to be consistent, to prevent the rotation coordinate axes from not matching when generating internal nodes of the BVH tree;
[0062] (3.3) Re-densify the training of the block, and adopt the same multi-modal loss function as the previous two times of training, to correct the calibrated rotation coordinate axes of each Gaussian primitive.
[0063] (4) In the hierarchical structure merging stage, to reduce the video memory requirement, the BVH leaf nodes of all blocks are fused to generate the internal nodes of the BVH tree. The internal nodes have all the properties of the 3D Gaussian primitives, namely, the mean μ, covariance Σ, SH coefficient SH, and opacity α. In order to keep the same fast rasterization process as the leaf nodes for the internal nodes and to describe the appearance of the internal nodes as much as possible, the present invention adopts the following steps:
[0064] (4.1) Minimize the weighted 3D KL divergence to obtain the mean, covariance, SH coefficient and opacity of each internal node, as follows:
[0065]
[0066] Among them, l represents the lth layer of the BVH tree, μ (l+1) ,Σ (l+1) SH (l+1) and α (l+1) They represent the mean, covariance, SH coefficient and opacity of an internal node in the l+1 layer respectively, and E is the total number of all child nodes of the internal node. and Represent the mean, covariance, SH coefficient and opacity of the i-th child node of the internal node, w i is the normalized weight, Calculated, where w i ′ is the original weight before normalization;
[0067] (4.2) Calculate the original weight w i ′, consider an independent Gaussian basis element Its contribution to a pixel position (x, y) is C i (x,y)=o i c i G(x,y), where o i for The opacity, c i for The color of , μ′, Σ′ are Gaussian basis elements The original mean and covariance of , then the Gaussian basis The contribution to the entire image is Finally, the original weight is obtained
[0068] (5) Rendering stage: Optimize and improve the quality of local details, achieve smooth transitions between different layers, and effectively reduce video memory requirements, including:
[0069] (5.1) Select a set of nodes (cuts) required for rendering from all nodes in the BVH tree, including:
[0070] (5.1.1) The granularity of node n ∈ (n) is defined as the pixel size of its projection on the screen in a given view. In order to achieve high-quality sampling and optimize the hierarchy at all levels, a standard random variable ξ ∈ [0, 1) is defined. The target granularity is given by the formula Calculated, where τ max and τ min Set the maximum and minimum values of the target granularity respectively;
[0071] (5.1.2) If the granularity of node n in a given view is smaller than the target granularity, that is, ∈(n)<τ ∈ , but the granularity of its parent node p is larger than the target granularity, that is, ∈(p)>τ ∈ , then node n is selected in the cut;
[0072] (5.1.3) If node n does not meet the target granularity requirement in the current view after switching views, a new node new is generated between node n and its parent node p by interpolation, specifically:
[0073] ① First generate two Gaussian primitives with the same attributes as the parent node (except for opacity α), and then interpolate the attributes of the two Gaussian primitives respectively. The mean μ and SH coefficient SH are calculated by interpolation weights. Calculate, where t n is the interpolation weight of node n, and the mean and SH coefficient of node n are directly multiplied by the interpolation weight to obtain the mean and SH coefficient of the new node new;
[0074] ② Since the covariance matrix Σ can be decomposed into a combination of the scaling matrix S and the rotation matrix R, Σ=(RS)(RS) T , and experiments have found that better results can be achieved by interpolating the scaling matrix S and rotating the matrix R instead of the interpolation covariance Σ. Therefore, the present invention multiplies the scaling matrix S and the rotation matrix R by the interpolation weight t n , and then calculate the covariance matrix of the new node new;
[0075] ③ The opacity of the two new nodes, i.e. the middle interpolation nodes, is determined by the formula α = t n α n +(1-t n )α′ is calculated, where t n is the interpolation weight, α n is the opacity of node n, and α′ is given by the formula Calculated, α p is the opacity of the parent node;
[0076] (5.2) To avoid excessive memory overhead, and to avoid the parent node size being only slightly larger than the child node, which will lead to the difficulty of optimization of these nodes, the generated BVH tree is sparsified, including:
[0077] (5.2.1) First, mark the leaf nodes to avoid deletion, generate a candidate node set Cut1 by traversing all training images with an initial granularity of 3 pixels, and screen the bottom layer nodes from Cut1 to form Cut2 (these nodes represent the detail retention threshold of the current granularity);
[0078] (5.2.2) Then delete the redundant node level between the leaf nodes and Cut2;
[0079] (5.2.3) Then the target granularity is multiplied successively, and after each granularity update, the Cut set generation, bottom layer node screening and intermediate node deletion operations are repeated until the target granularity is expanded to 50% of the original image size. This progressive sparsification method effectively controls the memory occupation of the tree structure while ensuring multi-level detail retention;
[0080] (5.3) Scan all nodes of the BVH tree, and if the distance from a Gaussian cell of a certain block C i to another block C j is less than the distance to the block C i (i≠j), delete the Gaussian cell;
[0081] (5.4) Output the RGB rendering image and the semantic rendering image, as shown in Figure 2 and Figure 3 .
[0082] After the above operations are completed, the RGB and semantic information images of the large-scale scene can be obtained.
[0083] Correspondingly, the application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics as described above. As Figure 4 shown, a hardware structure diagram of the device with data processing capability is provided for the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics of the embodiment of the application. In addition to the processor, memory and network interface shown in Figure 4 , the device with data processing capability in the embodiment usually includes other hardware according to the actual function of the device with data processing capability, and details are not repeated.
[0084] Correspondingly, the application further provides a computer readable storage medium, which stores computer instructions, and the instructions are executed by a processor to implement the large-scale three-dimensional scene reconstruction method capable of simultaneously training scene RGB and semantics as described above. The computer readable storage medium can be an internal storage unit of any device with data processing capability, such as a hard disk or a memory. The computer readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit of any device with data processing capability and the external storage device. The computer readable storage medium is used to store the computer program and other programs and data required by the device with data processing capability, and can also be used to temporarily store data that has been output or will be output.
[0085] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the application be limited only by the scope of the claims, including any appropriate equivalents.
[0086] It is to be understood that the application is not limited to the precise construction described and as shown in the attached drawings, and that various modifications and changes can be effected therein by those skilled in the art without departing from the scope of the application.
[0087] The above description is only the preferred embodiment of the present application, although the present application has been disclosed as above with the preferred embodiment, however, it is not intended to limit the present application. Any skilled person in the art, without departing from the scope of the technical solutions of the present application, can make many possible changes and modifications to the technical solutions of the present application disclosed above, or modify equivalent embodiments of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application, without departing from the content of the technical solutions of the present application, still belongs to the protection scope of the technical solutions of the present application.
Claims
1. A large-scale 3D scene reconstruction method capable of simultaneously training scene RGB and semantics, characterized in that, The method comprises the following steps: (1) inputting multi-view RGB images of a target scene and performing data preprocessing, dividing the target scene by point cloud density, and standardizing image, depth, and semantic data in each block; (2) in the generalization training stage, the target scene is roughly trained, and 3D Gaussian primitives of the whole scene are trained based on global point cloud without using Gaussian densification, and a Gaussian primitive set is obtained; (3) in the block-independent training and hierarchical BVH tree construction stage, the following operations are performed on each block: (3.1) training Gaussian primitives in a certain block using Gaussian densification; (3.2) using an AABB type bounding box for all Gaussian primitives in the block, constructing a BVH tree by median division, taking the Gaussian primitives generated in the rough training as leaf nodes, and simultaneously standardizing the rotation direction of all Gaussian primitives to be consistent; (3.3) re-training the block without densification to correct the rotation coordinate axes of each Gaussian primitive; (4) in the hierarchical structure merging stage, the leaf nodes of all blocks are fused to generate internal nodes of the BVH tree, the internal nodes have all the attributes of the Gaussian primitives, and the mean, covariance, SH coefficient and opacity of each internal node are obtained by minimizing the weighted 3D KL divergence; (5) in the rendering stage, the local detail quality is optimized and improved, specifically: (5.1) selecting a rendering node in all nodes of the BVH tree, including: defining a standard random variable and calculating a target granularity; if the granularity of the node under a given view is smaller than the target granularity, but the granularity of its parent node is larger than the target granularity, the node is selected; if the node does not meet the target granularity requirement in the current view after switching the view, a new node is generated between the node and its parent node by interpolation; (5.2) sparsifying the generated BVH tree; (5.3) scanning all nodes of the BVH tree, and if the distance from a Gaussian primitive belonging to a certain block to another block is smaller than the distance to the block, the Gaussian primitive is deleted; (5.4) outputting RGB rendering images and semantic rendering images.
2. The method of claim 1, wherein, Step (1) specifically comprises: (1.1) inputting multi-view RGB images of a target scene and obtaining pose information; (1.2) generating a depth map, a semantic mask, and a global point cloud wherein p i is the ith point, and N is the number of points in the global point cloud. (1.3) correcting the positional deviation of points under different views by point triangulation and beam adjustment method; (1.4) creating a sky box, placing a plurality of 3D Gaussian primitives on a sphere with a diameter several times that of the scene, and rendering the sky to simulate an infinite scene; (1.5) Active partitioning based on the density of global point cloud to obtain a block set where C k denotes the k-th block, and N' denotes the total number of blocks obtained by partitioning.
3. The method of claim 1, wherein, In step (2), based on the global point cloud Instead of training 3D Gaussians for the whole scene using Gaussian densification, we get a set of Gaussians, denoted as where denotes the j-th Gaussian, is the mean vector representing the position, is the covariance matrix representing the shape, is the SH coefficient vector representing the color, s j is the semantic label, a j is the opacity, and M is the number of Gaussians.
4. The method of claim 1, wherein, The loss function of the coarse training stage and the block independent training and hierarchical BVH tree construction stage Both adopt a multi-modal loss function of joint RGB reconstruction loss And semantic cross-entropy loss , λ S is a weight coefficient, Wherein represents a rendered image, represents ground truth, λ SSIM is a hyperparameter, and SSIM(·) is a structural similarity loss function, Wherein is a semantic pseudo ground truth, S k is the trained semantics, and S is the total number of semantic labels.
5. The method of claim 1, wherein, In step (4), the calculation formulas of the mean, covariance, SH coefficient and opacity of the internal node are as follows: where l denotes the l-th layer of the BVH tree, μ (l+1) ,∑ (l+1) ,SH (l+1) and α (l+1) respectively denote the mean, covariance, SH coefficient and opacity of a certain internal node in the (l+1)-th layer, E is the total number of all child nodes of the internal node, and respectively denote the mean, covariance, SH coefficient and opacity of the i-th child node of the internal node, w i is the normalized weight.
6. The method of claim 1, wherein, The new node generated between the node and its parent node by interpolation is specifically: ① Generate two Gaussian cells with the same attributes as the parent node, and interpolate the mean μ and SH coefficients SH of each Gaussian cell to obtain the interpolation weight of the Gaussian cell Compute where τ ∈ is the target granularity, and ∈(n), ∈(p) are the granularities of the node n and its parent node, respectively. Multiply the mean and SH coefficients of the node n directly by the interpolation weight to obtain the mean and SH coefficients of the new node new. ②Since the covariance matrix Σ can be decomposed into the combination of a scaling matrix S and a rotation matrix R, Σ = (RS)(RS) T The scaling matrix S and the rotation matrix R are multiplied by the interpolation weight t, respectively n The covariance matrix of the new node new is calculated. ③ The opacity of the two new nodes is determined by the formula α=t n α n +(1-t n )α′ is calculated, where α n is the opacity of node n, α p is the opacity of the parent node.
7. The method of claim 1, wherein, The sparsification of the generated BVH tree is specifically: ① Mark the leaf nodes, and generate a candidate node set Cut1 by traversing all training images with 3 pixels as the initial granularity, and select the bottom layer nodes from Cut1 to form a set Cut2; ② Delete the redundant node levels between the leaf nodes and Cut2; ③ Increase the target granularity by 50% each time, and after each granularity update, repeat the operations of Cut set generation, bottom layer node selection and intermediate node deletion until the target granularity is expanded to 50% of the original image size.
8. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled with the processor; wherein the memory is configured to store program data, and the processor is configured to execute the program data to implement the method for reconstructing a large-scale three-dimensional scene according to any one of claims 1-7.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method for reconstructing a large-scale three-dimensional scene according to any one of claims 1-7.
10. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the method for reconstructing a large-scale three-dimensional scene according to any one of claims 1-7.
Citation Information
Cited By
AR glasses rendering operation method and system based on spatial layering
CN122115747A
AR glasses rendering operation method and system based on spatial layering
CN122115747B