3DGS large-scale scene tiling and quality hierarchical expression method and 3DGS large-scale scene tiling and quality hierarchical expression system

By using 3DGS large-scene tile representation and quality layering, combined with layering, block division and compression techniques, the bandwidth bottleneck and latency issues in real-time rendering and streaming of large-scale 3D scenes are solved, achieving efficient adaptive rendering and transmission, suitable for urban roaming and virtual reality applications.

CN121883708APending Publication Date: 2026-04-17NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANKAI UNIV
Filing Date
2025-12-24
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies suffer from bandwidth bottlenecks and data transmission latency issues in real-time rendering and streaming of large-scale 3D scenes, especially when network conditions change, making it difficult to guarantee visual quality and efficiency.

Method used

We adopt the 3DGS large scene tile representation and quality layering method, and achieve intelligent rendering and data transmission through layered and block processing and compression technology, combined with adaptive network transmission. We use an alternating training mechanism of "base + incremental" for progressive optimization.

Benefits of technology

It effectively solves the problems of video memory consumption, network transmission bandwidth limitation and rendering efficiency in large-scene 3D Gaussian modeling, supports efficient streaming transmission and real-time rendering, and is suitable for large-scene real-time rendering applications such as city roaming and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883708A_ABST
    Figure CN121883708A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of the three-dimensional reconstruction technology, discloses a 3D GS large-scale scene tiling and quality hierarchical expression method and system, comprising: obtaining multi-view scene data and performing multi-scale downsampling to obtain a training view set; performing Gaussian training on the skeleton based on the training view set to obtain a background skeleton; non-uniform and non-overlapping partitioning is carried out on the multi-view scene data based on the scene geometric features to obtain partitioned point clouds; determining a block training set based on the background skeleton and the picture importance; on the basis of taking the block training set and the block point cloud as input, incremental hierarchical block-merging progressive training is carried out on each block, a'base + increment 'alternative training mechanism is adopted, local optimization is carried out on an increment section by taking the block as a unit, and a quality hierarchical block model is generated through cross-block global merging optimization; and performing compression based on the quality layering and partitioning model to generate a compression model organized according to partitioning and quality hierarchy. And bandwidth consumption of network transmission is self-adapted while high rendering quality is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D reconstruction technology, and more specifically to a method and system for representing large-scale 3DGS scenes as blocks and with quality layering. Background Technology

[0002] In recent years, with the rapid development of augmented reality (AR), virtual reality (VR), and cloud gaming, the demand for high-quality, seamless online 3D environments has been increasing. The success of these applications hinges on their ability to accurately represent complex 3D scenes. 3D Gaussian, as a novel 3D scene representation method, offers superior visual quality and real-time rendering capabilities compared to traditional mesh and Nerf 3D reconstruction techniques, making it suitable for interactive applications and a hot topic in research and application. However, as the scale of 3DGS scenes continues to expand, challenges in storage, computation, and transmission become increasingly prominent. Especially during network transmission, the sheer volume of scene data exceeds the bandwidth and storage capacity of different user devices, leading to bandwidth bottlenecks and data transmission delays in streaming transmission, severely impacting user experience. Achieving an adaptive transmission method to ensure the stability and efficiency of data streams under varying network conditions is crucial. The key lies in ensuring that the 3DGS model can dynamically adjust according to changes in real-time network bandwidth to adapt to the needs of different users, thereby avoiding visual quality degradation or latency caused by untimely or overloaded transmission.

[0003] Existing research has proposed a variety of optimization methods. In the paper entitled "A hierarchical 3DGS representation for real-time rendering of large datasets" published at the 2024 ACM SIGGRAPH conference, an attempt was made to divide the big data scene into spatially adjacent uniform blocks, train the blocks in parallel and achieve certain results. However, the blocks in the paper are large and do not take into account the point cloud distribution and scene geometry. The size of the 3DGS training file is still not suitable for network streaming.

[0004] Therefore, how to support real-time rendering and streaming of large-scale scenes, thereby improving user experience, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of the above problems, the present invention is proposed to provide a method and system for 3DGS large scene tile representation and quality layering to overcome or at least partially solve the above problems. By performing layered and block processing on large scenes and combining compression technology, it can provide high rendering quality while adapting to the bandwidth consumption of network transmission, and can intelligently cope with the rendering and data transmission needs under different scene, device and network conditions.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, embodiments of the present invention provide a method for 3DGS large scene tile representation and quality layering, including: Acquire multi-view scene data and perform multi-scale downsampling to obtain a training view set; The background skeleton is obtained by performing Gaussian training on the skeleton based on the training view set; Based on the scene's geometric features, the multi-view scene data is divided into non-uniform, non-overlapping blocks to obtain a block point cloud. The block training set is determined based on the background skeleton and image importance. Based on the block training set and the block point cloud as input, incremental hierarchical block segmentation-merging progressive training is performed on each block. An alternating training mechanism of "base + increment" is adopted. Local optimization is performed on the incremental segment in blocks as units. A quality hierarchical block model is generated by global merging optimization across blocks. Compression is performed based on the aforementioned quality hierarchical and block-based model to generate a compressed model organized according to blocks and quality levels.

[0007] In one embodiment, the method for obtaining the training view set is as follows: Based on the multi-view scene data, obtain the input image set and auxiliary data for image spatial alignment; Based on the input image set and the auxiliary data, a separable two-dimensional resampling operator based on the Lanczos kernel is defined on a preset scale set. Band-limited interpolation and downsampling are performed to obtain the training view set corresponding to different detail levels organized by layers.

[0008] In one embodiment, the method for obtaining the background skeleton is as follows: Obtain the corresponding scene sparse point cloud based on the training view set; The skeleton Gaussian is initialized using the sparse point cloud of the scene as the Gaussian center without introducing any densification or pruning. Throughout the training process, the position parameters of the points remain fixed, and only the appearance and shape-related attributes are optimized. The SH coefficient is gradually increased to enhance the representation ability. The anisotropic shape parameters are adjusted without changing the geometric topology of the skeleton Gaussian to obtain a stable skeleton Gaussian with global coverage, which serves as the background skeleton.

[0009] In one embodiment, the method for acquiring the segmented point cloud is as follows: Based on the scene image acquisition characteristics of the multi-view scene data, the main direction sequence is extracted from the images acquired from multiple perspectives; The cumulative path distance is calculated by accumulating the Euclidean distances between pairs of adjacent centers in the main direction sequence. The view, block length, and block direction contained in each block are determined based on the cumulative path distance, and the initial point cloud contained in each block is obtained based on the point cloud contained in each view of the total sparse point cloud. Based on the block length, the block direction, and the initial point cloud, the block range is determined, resulting in non-overlapping blocks of multiple target block lengths; The initial point cloud, which exceeds the block range, is filtered to obtain the final block point cloud.

[0010] In one embodiment, the method for determining the block range is as follows: The center coordinates of the blocks are determined based on the block length; The unit direction and unit normal vector of the block are calculated based on the block direction and the corresponding center coordinates of the block. Based on the initial point cloud contained in the block, the block center coordinates are subtracted and then projected onto the unit normal vector to obtain a series of distance values. The quantile values ​​are used to remove noise sparse points to obtain the left and right widths of the block. The block range is determined based on the left and right widths of the block.

[0011] In one embodiment, the method for determining the block training set is as follows: A set of auxiliary viewpoints is obtained based on the main direction sequence; Based on the main direction sequence, the cumulative path distance, and the target block length, the main view index set of the block is obtained; The set of all views is the union of the main view index set and the auxiliary view set. Based on the background skeleton, the corresponding image importance is calculated for each candidate view in the block candidate view set; Candidate views whose image importance is greater than a set threshold are selected as key views to form a key view set. The block training set is obtained by taking the union of the full view set and the key view set.

[0012] In one embodiment, the hierarchical, block-based, and merging progressive training specifically includes: Within each quality level, an alternating "base + increment" training mechanism is adopted, which performs local optimization on the increment and trains based on the optimization results of the previous layer. Based on the block point cloud and the block training set as input, the Gaussian properties of the base part are frozen for training, and geometric constraints are set on the incremental Gaussian of the block to gradually optimize and obtain a coarsely optimized incremental block model. The overall model is obtained by merging all incremental block models; Based on the overall model, a progressive geometric adaptation strategy that gradually tightens as the training progress is adopted to apply piecewise scaling to the gradients of various parameters of the incremental Gaussian to generate a global incremental Gaussian. Based on the global incremental Gaussian, the data is written back to each incremental block model according to the block geometric range to obtain the quality hierarchical block model.

[0013] In one embodiment, during the block incremental optimization process, an upper limit for the number of Gaussian elements in the block is set based on the number of point clouds contained in the block, the number of training images in the block training set, and the quality level, in order to control the model capacity and suppress excessive densification. And after each densification, if the total number of Gaussians exceeds the upper limit, the part with the smallest opacity is pruned from the trainable Gaussians.

[0014] In one embodiment, setting the geometric constraints specifically includes: The first projection is obtained based on the unit principal direction coordinates and the block center coordinates; The second projection is obtained based on the normal coordinates and the block center coordinates; The left and right half widths of the block are obtained based on the normal coordinates and the block center coordinates; The block-in-value is determined based on the first projection, the second projection, and the left and right half-widths. The Euclidean distance between any point in the block and the center of the current training view camera is used as the distance judgment value; Based on the intra-block judgment value and the near-far judgment value, the incremental Gaussian of the block is divided into multiple categories, and the element constraints of the corresponding category view are obtained. A normalized out-of-bounds term is obtained based on the first projection, the second projection, and the target block length; The position loss is obtained based on the current training progress, the element constraints, and the normalized out-of-bounds term. The geometric constraints are set based on the position loss as the incremental Gaussian.

[0015] In a second aspect, embodiments of the present invention provide a 3DGS large scene tile representation and quality hierarchical representation system for executing a 3DGS large scene tile representation and quality hierarchical representation method as described in any one of claims 1-9, characterized in that it includes: a skeleton Gaussian training module, a data partitioning module, a training set acquisition module, a hierarchical progressive training module, and a model compression module; The skeleton Gaussian training module is used to acquire multi-view scene data and perform multi-scale downsampling to obtain a training view set; and to train the skeleton Gaussian based on the training view set to obtain the background skeleton. The data partitioning module is used to perform non-uniform, non-overlapping segmentation of the multi-view scene data based on scene geometric features to obtain segmented point clouds. The training set acquisition module is used to determine a block training set based on the background skeleton and image importance. The hierarchical progressive training module is used to perform incremental hierarchical block-merging progressive training on each block based on the block training set and the block point cloud as input. It adopts an alternating training mechanism of "base + increment", performs local optimization on the incremental segment with the block as the unit, and generates a quality hierarchical block model through cross-block global merging optimization. The model compression module is used to compress the quality hierarchical block model to generate a compressed model organized according to blocks and quality levels.

[0016] As can be seen from the above technical solution, compared with the prior art, this invention discloses a method and system for 3DGS large-scene tile representation and quality layering. Through scene geometry-guided adaptive non-uniform tile partitioning, combined with a multi-level quality layering training mechanism, it achieves progressive optimization from the skeleton Gaussian base to incremental details. It employs a hybrid rendering strategy and dynamic Gaussian upper bound control to achieve efficient utilization of video memory during tile training. Finally, it uses customized Draco compression technology to quantize and compress Gaussian attributes, significantly reducing data transmission volume. This method effectively solves the problems of video memory consumption, network bandwidth limitations, and rendering efficiency in large-scene 3D Gaussian modeling. While ensuring visual quality, it effectively supports efficient streaming and real-time rendering of large-scale scenes, and is widely applicable to real-time rendering applications of large scenes such as city roaming and virtual reality. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0018] Figure 1 This is a flowchart of a 3DGS large scene tile representation and quality layering method provided in an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram illustrating the layered and block-based principle provided in an embodiment of the present invention.

[0020] Figure 3 This is a schematic diagram of a 3DGS large scene tile representation and quality layering system provided in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Example 1 When users are experiencing large-scale scenes such as city roaming, their current viewpoint often only occupies a portion of the scene. By dividing the scene into appropriate blocks and transmitting data on demand, unnecessary data can be avoided in areas outside the user's field of view. By dynamically predicting the loading area based on the user's perspective, loading time can be reduced and redundant transmissions can be avoided, thereby significantly reducing latency.

[0023] However, the varying needs of users under different scenarios and network conditions mean that a simple chunking strategy may not provide sufficient visual detail. For example, in low-bandwidth environments, although the predicted viewpoint region may require fewer chunks to be transmitted, the detail requirements of that region may be very high, while other regions, even if ignored, may not significantly impact the user experience. In this case, a layered strategy, by providing different levels of detail for each region, can dynamically adjust the loading of high- or low-level details based on bandwidth and user device capabilities, thereby achieving more efficient data transmission. Combining these two approaches can optimize storage and transmission resources in big data scenarios, improve user experience, ensure smooth rendering under various device and network conditions, and support high-quality adaptive transmission.

[0024] Based on the above technical background and research status, such as Figure 1 As shown, this embodiment of the invention discloses a method for 3DGS large scene tile representation and quality layering, including the following steps. For ease of description, these steps are numbered S1 to S6, and these numbers are not used to limit the sequential relationship between the various steps of this invention: S1 acquires multi-view scene data and performs multi-scale downsampling to obtain a training view set.

[0025] Furthermore, the method for obtaining the training view set is as follows: Based on multi-view scene data, the input image set and auxiliary data for image spatial alignment are obtained; Based on the input image set and auxiliary data at a preset scale set The above defines a separable two-dimensional resampling operator based on the Lanczos kernel, performs band-limited interpolation and downsampling, and obtains the training view set corresponding to different detail levels organized by layers. : ; in, Represents the camera parameter matrix. This represents the corresponding image. M This indicates the number of images.

[0026] S2 is trained using Gaussian on the skeleton based on the training view set to obtain the background skeleton.

[0027] Furthermore, the method for obtaining the background skeleton is as follows: Obtain the corresponding scene sparse point cloud based on the training view set; The skeleton Gaussian is initialized using sparse point clouds of the scene as Gaussian centers. No densification or pruning is introduced. Throughout the training process, Gaussian position gradient updates are prohibited, and the position parameters of the points are kept fixed. Only appearance and shape-related attributes are optimized, and the SH coefficient is gradually increased to enhance the representation ability. The anisotropic shape parameters are adjusted without changing the geometric topology of the skeleton Gaussian to obtain a stable skeleton Gaussian with global coverage, which serves as the background skeleton.

[0028] Furthermore, appearance and shape-related properties include: spherical harmonic coefficient, scale, rendering, and opacity.

[0029] Furthermore, the training loss used in the skeleton Gaussian training process for: ; in, I gt This represents the original image. I pred This refers to a rendered image. SSIM ( ) represents structural similarity loss, and ||1 represents L1 loss. λ SSIM This represents the weight corresponding to the SSIM loss.

[0030] Furthermore, the background skeleton provides important priors for subsequent "key image selection" and "lowest quality layered base".

[0031] S3 performs non-uniform, non-overlapping segmentation of multi-view scene data based on scene geometric features to obtain segmented point clouds.

[0032] Furthermore, the method for obtaining the segmented point cloud is as follows: Based on the scene image acquisition characteristics of multi-view scene data, the main direction sequence is extracted from the images acquired from multiple perspectives; The cumulative path distance is calculated by accumulating the Euclidean distances between pairs of adjacent centers in the main direction sequence. The view, length, and orientation of each block are determined based on the cumulative path distance, and the initial point cloud of each block is obtained based on the point cloud of each view. The segmentation range of the segments is determined based on the segmentation length, segmentation direction, and initial point cloud, resulting in non-overlapping segments with multiple target segmentation lengths. The final segmented point cloud is obtained by filtering the initial point cloud that exceeds the segmentation range.

[0033] Furthermore, the main direction sequence is Its camera center in the plane is Considering only the XY plane, for each principal viewpoint i ∈ M There exists a set of auxiliary perspectives. A ( i (Time-synchronized images), accumulating distance along the path from the main viewpoint sequence. s k (Accumulate the Euclidean distances between each pair of adjacent centers) Perform fixed-length segmentation, denoted as L for the target segment length, and divide it into K non-overlapping segments.

[0034] Furthermore, the method for determining the block range is as follows: The center coordinates of the blocks are determined based on the block length; The unit direction and unit normal vector of the block are calculated based on the corresponding block direction and block center coordinates. Based on the initial point cloud contained in the block, the block center coordinates are subtracted and then projected onto the unit normal vector to obtain a series of distance values. The quantile values ​​are used to remove noisy and sparse points to obtain the left and right widths of the block. The block range is determined based on the left and right widths of the blocks.

[0035] S4 determines the block training set based on the background skeleton and image importance.

[0036] Furthermore, the method for determining the block training set is as follows: Obtain the set of auxiliary viewpoints based on the main direction sequence; Based on the main direction sequence, path cumulative distance, and target block length, the main view index set of the block is obtained; The full view set is based on the union of the main view index set and the auxiliary view set. Calculate the image importance of each candidate view in the segmented candidate view set based on the background skeleton; Candidate views whose image importance is greater than a set threshold are selected as key views to form a key view set. The block training set is obtained by taking the union of the full view set and the key view set.

[0037] Furthermore, the main perspective index set of the j-th block for: ; in, , i k Views representing the main direction sequence; The full view set of the j-th block Main perspective index set With auxiliary perspective set A ( i Union of ) .

[0038] Furthermore, the first and last images in the main view set of segment s, sorted by path sequence, are: i min and i max The direction vector of the block (XY only) is determined by the difference vector between the camera centers of the first and last main viewpoint images: ; in, This indicates the camera center of the main view image. This indicates the camera center of the primary viewpoint image; Unit principal direction coordinates of the j-th block d j Normal coordinates p j and block center coordinates c j They are respectively: ; ; ; in, d j,y Let represent the ordinate of the unit principal direction of the j-th block. d j,x This represents the x-coordinate of the unit principal direction of the j-th block.

[0039] Furthermore, to enhance the intra-block monitoring signal, for each block, a background skeleton-based approach is used. G skel Calculate the contribution of each candidate view to the block: Start with the background skeleton G skel Filtering out Gaussians within the block range yields Gaussian-preserving Gaussians. G out ; Based on background skeleton G skel and retain Gauss Gout Each training view is rendered separately to obtain the corresponding rendered view. I full and I out Since SSIM loss can effectively capture structural differences and is to some extent insensitive to brightness changes, the importance of an image to a block is defined by the SSIM difference. :

[0040] in, I gt This refers to the original image.

[0041] Furthermore, to clarify the rule of "selecting above the threshold and not selecting below the threshold", the threshold for block j in this embodiment is set as follows: τ j The key view set obtained based on the threshold for: ; in, I cand ( j ) represents the set of candidate views, which includes all views or views that intersect with the space of block j.

[0042] Furthermore, in this embodiment, a threshold is set. τ j Use a globally fixed threshold, or adaptively set it according to the quantiles of the score set in that block: ,in, Percentile This represents the quantile statistic. q This indicates the percentile ordinal number, used to specify the percentile position.

[0043] Furthermore, based on the full view collection With key view sets By taking the union of the sets, we obtain the block training set. It provides more training images in terms of appearance / geometry.

[0044] S5 uses block training sets and block point clouds as inputs, and performs incremental hierarchical block segmentation and merging progressive training on each block. It adopts an alternating training mechanism of "base + increment", performs local optimization on the increment segment on a block-by-block basis, and generates a quality hierarchical block model through global merging optimization across blocks.

[0045] Furthermore, such as Figure 2As shown, the hierarchical block-based Gaussian training method of this invention gradually enhances the detail and accuracy of the rendering model by progressively improving the quality and resolution of the input image layer by layer: in each layer, the resolution and quality of the training image are progressively improved, starting from low-quality images and gradually transitioning to high-quality images. Based on this, the incremental Gaussian model... Training is based on the previous layer's model, while the Gaussian model of the previous layer remains frozen. This freezing operation ensures that the increment of each layer can be gradually accumulated based on the optimization results of the previous layer, thereby achieving a progressive improvement in rendering quality. In this way, the model can gradually improve rendering quality during layered training, while block-based incremental training also makes the optimization of each layer more refined, ultimately ensuring efficient improvement and quality stability of rendering effects.

[0046] Layered, segmented, and merged progressive training, specifically including: Within each quality level, an alternating "base + increment" training mechanism is adopted, which performs local optimization on the increment and trains based on the optimization results of the previous layer. Based on the block point cloud and block training set as input, the Gaussian properties of the base part are frozen for training, and geometric constraints are set on the incremental Gaussian of the blocks to gradually optimize and obtain the coarsely optimized incremental block model. The overall model is obtained by merging all incremental block models; Based on the overall model, a progressive geometric adaptation strategy that gradually tightens as training progresses is adopted to apply piecewise scaling to the gradients of various parameters of the incremental Gaussian to generate a global incremental Gaussian. The quality-layered block model is obtained by writing back the global incremental Gaussian model to each incremental block model according to the block geometric range.

[0047] Furthermore, the block point cloud and block training set are loaded as input. For incremental training of the lowest quality layer, a skeleton Gaussian is loaded, and the Gaussian within the block range is removed as the basis. For higher quality layers, the block merging optimization result of the previous layer is loaded as the basis, injected through attribute concatenation, and labeled. N base Gaussian property gradients are frozen during training in the base region.

[0048] Furthermore, during the block-based incremental optimization process, an upper limit is set on the number of Gaussian elements in each block based on the number of point clouds contained in the block, the number of training images in the block's training set, and the quality level. N max To control model capacity and suppress excessive densification in order to adapt to dynamic network transmission capacity constraints; And after each densification, if the total number of Gaussians exceeds the upper limit, the part with the smallest opacity is pruned from the trainable Gaussians.

[0049] Furthermore, after each densification strategy, if the total number N exceeds the limit, then in the trainable segment (index [ N base Delete in ascending order of opacity (N) k = N - N base Gaussian: ; ; in, G del This represents the Gaussian part from which deletion operations can be performed. S This indicates selecting the index to delete the Gaussian index. arg topk represents the k Gaussian pairs with the lowest opacity. k del This represents the range of Gaussian indices from which deletion operations can be performed; "smallest" indicates that the smallest opacity value should be selected. o [ i ] indicates the first i The opacity is a Gaussian.

[0050] Furthermore, Gaussians with low opacity contribute less to imaging and are therefore prioritized for elimination. The block-based incremental training reconstruction objective remains the same as the skeleton stage, but a positional loss is added to geometrically constrain the incremental Gaussians. The total loss can be expressed as... L : ; in, L pos Indicates position loss; To address the characteristics of "short block length and a large number of pixels within the block originating from outside the block due to outdoor scenes," geometric constraints involving layering, visibility perception, and progress adjustment are applied to the trainable block incremental Gaussian, i.e., the portion excluding the base segment. Strict constraints are imposed on near-field elements within the block; escape is allowed for distant elements outside the block; and moderate constraints are imposed on near-field elements outside the block to suppress out-of-bounds artifacts.

[0051] Furthermore, geometric constraints are set, specifically including: The first projection is obtained based on the unit principal direction coordinates and the block center coordinates; The second projection is obtained based on the normal coordinates and the block center coordinates; The left and right half-widths of the block are obtained based on the normal coordinates and the block center coordinates; The block-in-value judgment is obtained based on the first projection, the second projection, and the left and right half-widths; The Euclidean distance between any point in the block and the center of the current training view camera is used as the distance judgment value; Based on the intra-block judgment value and the near-far judgment value, the incremental Gaussian of the block is divided into multiple categories, and the element constraints of the corresponding category view are obtained. The normalized out-of-bounds term is obtained based on the first projection, the second projection, and the target block length; The position loss is obtained based on the current training progress, element constraints, and normalized out-of-bounds terms. Geometric constraints are set based on position loss as an incremental Gaussian.

[0052] Furthermore, within the chunk, a full view collection is used. The left and right widths of the cloud computing blocks representing points observed from all viewpoints. For any viewpoint... Includes the visible point X (taken as plane coordinates) Define the relative vector as the "right-side normal" of this block. p j The projection on is: ; Therefore, all projection value sets are decomposed into left and right sides. ; ; To suppress the influence of isolated points and noise, the left and right half-widths are estimated by quantiles, with quantiles set as q for the left and 1-q for the right, thus obtaining the left and right half-widths of the blocks. and : ; .

[0053] Furthermore, for block j, the unit principal direction coordinates d j Normal coordinates p j and block center coordinates c j The block length is L, and the left and right half widths are respectively... and The total width is w j For any point (the plane projection of the Gaussian center) x∈R 2 The first projection and the second projection are defined as follows: u and v : ; ; Based on the first projection u Second projection v and left and right half width and Get the judgment value within the block: ; in, β This represents the buffer coefficient, in this embodiment... β Set to 1.1; Based on any point in the block and the center of the current training view camera. c cam ∈ R 2 Euclidean distance is used as a value for judging foreground and background: ; and by dividing the diagonal Construct threshold , (In this embodiment) k =1.6), based on which the block incremental Gaussian is divided into three categories:

[0054] During training, constraints are only applied to elements visible in the current view, allowing... For visibility indication, the element constraints of category A view are: ;Element constraints of Category B view Similarly; In the loss design, a comparable scale of boundary exceedance is adopted, and soft boundaries are superimposed to avoid "edge-following jitter". First, a normalized boundary exceedance term is defined based on the first projection, the second projection, and the target block length. : ; ; ; ; Block constraints and schedule adjustment: Let the total number of iterations be T, and the current iteration be t. Let the training progress... r = t / T ∈[0,1], define the stage weights and combined loss to obtain the positional loss. L pos : ; ; ; ; ; Soft boundary terms are used to suppress near-edge oscillations while avoiding excessive penalty for the central structure within the block. Gradual weights are first relaxed and then tightened to allow the model to gain reconstruction freedom first, and then converge geometric consistency later. No constraints are imposed on the distant view to retain the necessary distant view support in outdoor scenes.

[0055] Furthermore, the block merging optimization process includes: All incremental block models in this layer are merged to obtain the overall model. The merged and optimized incremental blocks of all low-quality layers are then used as the basis. The lowest quality layer does not use a basis, and only the geometric degrees of freedom of the merged incremental Gaussian are opened. The basis properties are frozen during training. In order to achieve a balance between preserving the stable prior of the basis and adapting the incremental part to the basis, a progressive geometric adaptation strategy that gradually tightens as training progresses is adopted, and piecewise scaling is applied to the gradients of various parameters of the incremental Gaussian.

[0056] Let T be the total number of training iterations, t be the current iteration number, and let the training progress be... r = t / T ∈[0,1], the length of the base segment is N base The incremental segment index set is For any incremental Gaussian i ∈ I inc For each position x i ∈ R 3 With rotation coefficient q i Gradient application scaling function , , : ; ; ; ; In the early stages of training, the incremental Gaussian is given ample room for exploration. Relatively large values ​​are assigned to the aforementioned scale functions to help quickly align details and orientations and fully adapt to block seams. In the mid-stage, the geometry is progressively tightened, and the initial values ​​of each scale function are reduced accordingly, decreasing continuously as training progresses, in order to reduce jitter and stabilize consistency with the base. In the later stages, smaller scale function values ​​are set to keep the geometry almost fixed, shifting the optimization focus to parameters such as appearance and opacity, improving model quality without excessively perturbing the geometry, and reducing optimization instability and divergence risks.

[0057] Furthermore, based on the global incremental Gaussian, the data is written back to each incremental block model according to the block geometric range, specifically including: The metadata of each block is read, and the global incremental Gaussian is written back to each incremental block model according to the block's geometric range, forming a standardized data organization to support subsequent layer training and constructing the two-dimensional boundaries of the blocks. First, based on the function of determining whether a Gaussian point is within a block, the Gaussian point is written back to the corresponding block model. For points not within any block range, the weighted distance to each block is calculated, and the block with the smallest distance is selected and saved to the corresponding incremental block model.

[0058] Furthermore, during block-based incremental training, the densification strategy in the original 3DGS is modified to focus on the maximum gradient at Gaussian locations rather than its average value. This strategy helps to improve the densification of local details, thus better handling sparse and scattered image data scenarios. Through this strategy, the system can optimize scene training results, ensuring richer and more refined details.

[0059] Furthermore, during the incremental block training process, based on the characteristics of outdoor scenes and the block length limit, an intelligent position loss constraint is introduced to optimize visibility perception. Intra-block decision with buffered Gaussian is used, and thresholds for near and far views are set. By adjusting view visibility and out-of-bounds measurement, the weights of the training process are gradually optimized to improve the model's degrees of freedom and convergence stability.

[0060] S6 compresses based on a quality-layered block model, generating a compressed model organized according to blocks and quality levels.

[0061] Furthermore, using Draco compression technology, the quantization bits are specified for each Gaussian attribute, and the compression level is set to further reduce the data size of the quality-layered block model. The compressed model is indexed according to blocks and quality levels, supporting block-level retrieval and hierarchical progressive transmission based on view frustum or position for online systems, so as to realize on-demand loading and rendering of network streaming.

[0062] Furthermore, utilizing the Metadata API in Draco compression technology, custom Gaussian attributes are added, specifying the Gaussian attribute center position, opacity, spherical harmonic coefficient, scaling factor, quantization bit depth, and compression level. Compression levels range from 0 to 10; higher numbers result in higher compression ratios but slower decompression. Quantization bit depth affects data volume and model loss; adjusting these factors ensures a balance between rendering quality and compression efficiency. The compressed model is indexed by chunks and quality levels, supporting block-level retrieval based on view frustum or position and hierarchical progressive transmission for online systems, enabling on-demand loading and rendering via network streaming.

[0063] Furthermore, based on the obtained compression model organized according to blocks and quality levels, in practical applications, the progressive streaming of corresponding spatial blocks is triggered by the user's view frustum spatial position. The server dynamically adapts the compression level to achieve network-aware quality-level delivery. The client decompresses and restores Gaussian attribute data in real time to drive the rendering pipeline, and dynamically adjusts the loading quality according to real-time performance indicators. Finally, a continuous, interactive, and quality-adaptive 3D scene visualization is output.

[0064] Example 2 like Figure 3 As shown, based on the same inventive concept, this embodiment of the invention also provides a 3DGS large scene tile representation and quality hierarchical representation system, including: a skeleton Gaussian training module, a data partitioning module, a training set acquisition module, a hierarchical progressive training module, and a model compression module; The skeleton Gaussian training module is used to acquire multi-view scene data and perform multi-scale downsampling to obtain a training view set; the skeleton Gaussian is trained based on the training view set to obtain the background skeleton. The data partitioning module is used to perform non-uniform, non-overlapping block partitioning of multi-view scene data based on scene geometric features to obtain block point clouds. The training set acquisition module is used to determine the block training set based on the background skeleton and image importance. The hierarchical progressive training module is used to perform incremental hierarchical block-merging progressive training on each block based on the block training set and block point cloud as input. It adopts an alternating training mechanism of "base + increment", performs local optimization on the incremental segment on the block as a unit, and generates a quality hierarchical block model through cross-block global merging optimization. The model compression module is used to compress a quality-layered, block-based model, generating a compressed model organized according to blocks and quality levels.

[0065] Furthermore, in this embodiment, the functional implementation methods of each functional module correspond one-to-one with the methods described above, and will not be repeated here.

[0066] Example 3 Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in the memory, it can implement a 3DGS large scene tile representation and quality layering representation method as shown in Example 1.

[0067] Based on the same inventive concept, the present invention also provides an electronic device, which includes a processor and a memory, wherein the memory stores instructions, characterized in that the instructions are loaded and executed by the processor to implement a 3DGS large scene tile representation and quality layering representation method as in Embodiment 1.

[0068] The electronic device may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can call logical instructions in the memory to execute a 3DGS large-scene tile representation and quality layering method as described in Embodiment 1.

[0069] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0070] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0071] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for representing large-scale 3DGS scenes using tile representation and quality layering, characterized in that, include: Acquire multi-view scene data and perform multi-scale downsampling to obtain a training view set; The background skeleton is obtained by performing Gaussian training on the skeleton based on the training view set; Based on the scene's geometric features, the multi-view scene data is divided into non-uniform, non-overlapping blocks to obtain a block point cloud. The block training set is determined based on the background skeleton and image importance. Based on the block training set and the block point cloud as input, incremental hierarchical block segmentation-merging progressive training is performed on each block. An alternating training mechanism of "base + increment" is adopted. Local optimization is performed on the incremental segment in blocks as units. A quality hierarchical block model is generated by global merging optimization across blocks. Compression is performed based on the aforementioned quality hierarchical and block-based model to generate a compressed model organized according to blocks and quality levels.

2. The method for 3DGS large scene tile representation and quality layering according to claim 1, characterized in that, The method for obtaining the training view set is as follows: Based on the multi-view scene data, obtain the input image set and auxiliary data for image spatial alignment; Based on the input image set and the auxiliary data, a separable two-dimensional resampling operator based on the Lanczos kernel is defined on a preset scale set. Band-limited interpolation and downsampling are performed to obtain the training view set corresponding to different detail levels organized by layers.

3. The method for 3DGS large scene tile representation and quality layering according to claim 2, characterized in that, The method for obtaining the background skeleton is as follows: Obtain the corresponding scene sparse point cloud based on the training view set; The skeleton Gaussian is initialized using the sparse point cloud of the scene as the Gaussian center without introducing any densification or pruning. Throughout the training process, the position parameters of the points remain fixed, and only the appearance and shape-related attributes are optimized. The SH coefficient is gradually increased to enhance the representation ability. The anisotropic shape parameters are adjusted without changing the geometric topology of the skeleton Gaussian to obtain a stable skeleton Gaussian with global coverage, which serves as the background skeleton.

4. The method for 3DGS large scene tile representation and quality layering according to claim 2, characterized in that, The method for obtaining the segmented point cloud is as follows: Based on the scene image acquisition characteristics of the multi-view scene data, the main direction sequence is extracted from the images acquired from multiple perspectives; The cumulative path distance is calculated by accumulating the Euclidean distances between pairs of adjacent centers in the main direction sequence. The view, block length, and block direction contained in each block are determined based on the cumulative path distance, and the initial point cloud contained in each block is obtained based on the point cloud contained in each view of the total sparse point cloud. Based on the block length, the block direction, and the initial point cloud, the block range is determined, resulting in non-overlapping blocks of multiple target block lengths; The initial point cloud, which exceeds the block range, is filtered to obtain the final block point cloud.

5. The method for 3DGS large scene tile representation and quality layering according to claim 4, characterized in that, The method for determining the block range is as follows: The center coordinates of the blocks are determined based on the block length; The unit direction and unit normal vector of the block are calculated based on the block direction and the corresponding center coordinates of the block. Based on the initial point cloud contained in the block, the block center coordinates are subtracted and then projected onto the unit normal vector to obtain a series of distance values. The quantile values ​​are used to remove noise sparse points to obtain the left and right widths of the block. The block range is determined based on the left and right widths of the block.

6. The method for 3DGS large scene tile representation and quality layering according to claim 5, characterized in that, The method for determining the block training set is as follows: A set of auxiliary viewpoints is obtained based on the main direction sequence; Based on the main direction sequence, the cumulative path distance, and the target block length, the main view index set of the block is obtained; The set of all views is the union of the main view index set and the auxiliary view set. Based on the background skeleton, the corresponding image importance is calculated for each candidate view in the block candidate view set; Candidate views whose image importance is greater than a set threshold are selected as key views to form a key view set. The block training set is obtained by taking the union of the full view set and the key view set.

7. The method for 3DGS large scene tile representation and quality layering according to claim 5, characterized in that, The progressive training involving layered, segmented, and merged data specifically includes: Within each quality level, an alternating "base + increment" training mechanism is adopted, which performs local optimization on the increment and trains based on the optimization results of the previous layer. Based on the block point cloud and the block training set as input, the Gaussian properties of the base part are frozen for training, and geometric constraints are set on the incremental Gaussian of the block to gradually optimize and obtain a coarsely optimized incremental block model. The overall model is obtained by merging all incremental block models; Based on the overall model, a progressive geometric adaptation strategy that gradually tightens as the training progress is adopted to apply piecewise scaling to the gradients of various parameters of the incremental Gaussian to generate a global incremental Gaussian. Based on the global incremental Gaussian, the data is written back to each incremental block model according to the block geometric range to obtain the quality hierarchical block model.

8. The method for 3DGS large scene tile representation and quality layering according to claim 7, characterized in that, In the block incremental optimization process, based on the number of point clouds contained in the block, the number of training images in the training set of the block, and the quality level, an upper limit for the number of Gaussian elements in the block is set to control the model capacity and suppress excessive densification. And after each densification, if the total number of Gaussians exceeds the upper limit, the part with the smallest opacity is pruned from the trainable Gaussians.

9. A method for representing large-scale 3DGS scenes using tile representation and quality layering according to claim 7, characterized in that, Setting the geometric constraints specifically includes: The first projection is obtained based on the unit principal direction coordinates and the block center coordinates; The second projection is obtained based on the normal coordinates and the block center coordinates; The left and right half widths of the block are obtained based on the normal coordinates and the block center coordinates; The block-in-value is determined based on the first projection, the second projection, and the left and right half-widths. The Euclidean distance between any point in the block and the center of the current training view camera is used as the distance judgment value; Based on the intra-block judgment value and the near-far judgment value, the incremental Gaussian of the block is divided into multiple categories, and the element constraints of the corresponding category view are obtained. A normalized out-of-bounds term is obtained based on the first projection, the second projection, and the target block length; The position loss is obtained based on the current training progress, the element constraints, and the normalized out-of-bounds term. The geometric constraints are set based on the position loss as the incremental Gaussian.

10. A 3DGS large scene tile representation and quality layering system, used to execute a 3DGS large scene tile representation and quality layering method as described in any one of claims 1-9, characterized in that, include: The module includes a skeleton Gaussian training module, a data partitioning module, a training set acquisition module, a hierarchical progressive training module, and a model compression module. The skeleton Gaussian training module is used to acquire multi-view scene data and perform multi-scale downsampling to obtain a training view set; and to train the skeleton Gaussian based on the training view set to obtain the background skeleton. The data partitioning module is used to perform non-uniform, non-overlapping block division of the multi-view scene data based on scene geometric features to obtain a block point cloud. The training set acquisition module is used to determine a block training set based on the background skeleton and image importance. The hierarchical progressive training module is used to perform incremental hierarchical segmentation-merging progressive training on each segment based on the segmented training set and the segmented point cloud as input. It adopts an alternating training mechanism of "base + increment", performs local optimization on the incremental segment as a unit of segment, and generates a quality hierarchical segmented model through cross-segment global merging optimization. The model compression module is used to compress the quality hierarchical block model to generate a compressed model organized according to blocks and quality levels.