Layered representation compression and progressive transmission method of Gaussian splash volume video
By employing a hierarchical representation compression method for Gaussian splash volumetric video, the problems of hierarchical representation and motion adaptive modeling of dynamic scenes are solved, achieving efficient volumetric video compression and streaming transmission, and supporting multi-level quality selection and real-time decoding.
Patent Information
- Application Number
- CN202511136254.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies struggle to achieve layered representation of dynamic scenes, motion adaptive modeling, and end-to-end compression optimization within a single model, resulting in unresolved issues regarding the high bandwidth requirements and real-time decoding challenges of volumetric video streaming.
A hierarchical characterization compression method based on Gaussian splash volume video is adopted. By extracting the scene surface geometry to construct 4D Gaussian initial parameters, dividing the Gaussian keyframes into layers, performing rigid transformation and residual compensation, combining uniform noise to simulate quantization error and KDE to estimate entropy, and using H.264 compression to generate a bitstream that supports multi-level quality selection.
It achieves high-fidelity reconstruction and efficient compressed transmission of dynamic scenes in a single model, supports multi-terminal adaptation, is compatible with existing rendering pipelines, and enables real-time decoding rendering and flexible quality transitions.
Smart Images

Figure CN120915960A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of volumetric video compression and streaming, in particular to a hierarchical representation compression and progressive streaming method for Gaussian splatting volumetric video, and relates to a corresponding hierarchical representation compression system and progressive streaming system. BACKGROUND
[0002] Volumetric video provides immersive 3D experience for virtual reality, telemedicine and other scenarios with free-view navigation capability, but high bandwidth demand and real-time decoding challenges restrict its wide application. Existing dynamic neural radiance field methods have significant limitations when dealing with long sequence dynamic scenes, while 3D Gaussian and its extension methods achieve realistic rendering, but face the problem of parameter explosion when modeling long sequences (directly storing features of each frame results in a large model volume). For example, when the 3D Gaussian splatting method for static scenes is extended to dynamic scenes, the Gaussian parameters need to be preloaded for the entire sequence, which severely limits the feasibility of streaming. Existing solutions for dynamic 3D Gaussian representation mainly fall into two categories: one models Gaussian attributes as a function of time, which achieves high rate-distortion performance but ignores the streaming requirement; the other uses frame-by-frame rigid motion estimation, but the error accumulates due to the fixed reference frame for large motion scenes.
[0003] In addition, existing compression methods lack joint optimization of representation and compression: for example, the HPC method encodes dynamic scenes through residual grid, but the decoding delay is high and cannot meet real-time application requirements; 4DGC attempts joint optimization, but lacks robustness for large motion scenes.
[0004] More fundamentally, the traditional video coding framework (such as H.264) cannot be directly applied to the high-dimensional feature space of neural radiance field, and existing dynamic Gaussian compression methods generally lack hierarchical rate-distortion (RD) supervision mechanisms, resulting in loss of dynamic details after compression. Therefore, how to achieve hierarchical representation, motion adaptive modeling and end-to-end compression optimization in a single model is still a technical bottleneck that needs to be broken through in the field of volumetric video streaming. SUMMARY
[0005] The present application provides a hierarchical representation compression and progressive streaming method for Gaussian splatting volumetric video, and relates to a corresponding hierarchical representation compression system and progressive streaming system.
[0006] According to one aspect of the present application, a hierarchical representation compression method for Gaussian splatting volumetric video is provided, comprising:
[0007] extracting scene surface geometry, constructing 4D Gaussian initial parameters as the basic representation of dynamic scenes;
[0008] Based on the basic representation of the dynamic scene, an importance measure of each Gaussian initial parameter is calculated, the Gaussian is divided into L layers of progressive details, and a layered Gaussian key frame representation is obtained;
[0009] Based on the layered Gaussian key frame representation, inter-frame modeling is performed through rigid transformation and residual compensation;
[0010] The average displacement of the Gaussian primitive is calculated, and the grouping length of the inter-frame modeling is dynamically adjusted;
[0011] Uniform noise simulation quantization error is injected, and KDE is used to estimate the entropy of each attribute to guide the compression efficiency; meanwhile, a weighted loss of end-to-end optimization of distortion, code rate and time smoothness is constructed to train the inter-frame modeling;
[0012] The trained layered Gaussian parameters are encoded into 2D graphs and compressed using H.264 to complete the volumetric video compression.
[0013] According to a second aspect of the present application, a layered representation compression system of Gaussian splatting volumetric video is provided, comprising:
[0014] A parameter initialization module is configured to extract scene surface geometry, construct 4D Gaussian initial parameters as the basic representation of the dynamic scene;
[0015] A Gaussian model construction module is configured to calculate the importance measure of each Gaussian initial parameter based on the basic representation of the dynamic scene, divide the Gaussian into L layers of progressive details, and obtain a layered Gaussian key frame representation; based on the layered Gaussian key frame representation, inter-frame modeling is performed through rigid transformation and residual compensation; the average displacement of the Gaussian primitive is calculated, and the grouping length of the inter-frame modeling is dynamically adjusted;
[0016] A Gaussian model training module is configured to inject uniform noise simulation quantization error, and use KDE to estimate the entropy of each attribute to guide the compression efficiency; meanwhile, a weighted loss of end-to-end optimization of distortion, code rate and time smoothness is constructed to train the inter-frame modeling;
[0017] A dynamic scene compression module is configured to encode the trained layered Gaussian parameters into 2D graphs and compress them using H.264 to complete the volumetric video compression.
[0018] According to a third aspect of the present application, a progressive transmission method of Gaussian splatting volumetric video is provided, which is implemented by using the layered representation compression method described above.
[0019] Based on the volumetric video compression method, a bitstream supporting multi-level quality selection is generated;
[0020] The bitstream is decoded according to the required number of layers to realize progressive transmission and dynamic adaptation.
[0021] According to a fourth aspect of the present application, there is provided a progressive transmission system of Gaussian splatting volumetric video implemented by using the hierarchical representation compression method of Gaussian splatting volumetric video described above, comprising:
[0022] a bitstream generation module, which generates a bitstream supporting multi-level quality selection based on the volumetric video compression method;
[0023] a decoding module, which is used for decoding the bitstream according to the required number of layers to realize progressive transmission and dynamic adaptation.
[0024] Compared with the prior art, the present application has at least one of the following beneficial effects due to the adoption of the technical scheme described above:
[0025] The hierarchical representation compression and progressive transmission method of Gaussian splatting volumetric video provided by the present application adopts a progressive streaming framework of volumetric video, supports progressive encoding through hierarchical representation of Gauss based on importance perception, so that users can adapt to various terminals by selecting code streams of different levels. Our method can obtain multiple RD performances after a single model and one compression. Our compression uses H.264, can be compatible with existing rendering pipelines, and can be decoded and rendered in real time on mobile terminals.
[0026] The hierarchical representation compression and progressive transmission method of Gaussian splatting volumetric video provided by the present application can realize accurate non-key frame representation through rigid motion and residual deformation of Gauss.
[0027] The hierarchical representation compression and progressive transmission method of Gaussian splatting volumetric video provided by the present application can dynamically select key frames according to the actual motion of the sequence through motion-aware adaptive Gaussian grouping, thereby improving the reconstruction quality.
[0028] The hierarchical representation compression and progressive transmission method of Gaussian splatting volumetric video provided by the present application uses an end-to-end entropy optimization learning strategy, including level RD supervision and attribute-specific entropy modeling. Level RD supervision can optimize the RD performance of each Gaussian, and attribute-specific entropy modeling captures the distribution of each Gaussian attribute. The key frame uses KDE, and the residual value of the non-key frame uses Gaussian distribution to estimate, which can more accurately estimate the entropy of the attribute to make the RD performance of the model optimal. BRIEF DESCRIPTION OF DRAWINGS
[0029] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments with reference to the attached drawings:
[0030] Figure 1The workflow diagram of the hierarchical representation compression method of the Gaussian splatting volumetric video in a preferred embodiment of the present application.
[0031] Figure 2 The schematic diagram of the component modules of the hierarchical representation compression system of the Gaussian splatting volumetric video in a preferred embodiment of the present application.
[0032] Figure 3 The workflow diagram of the progressive transmission method of the Gaussian splatting volumetric video in a preferred embodiment of the present application.
[0033] Figure 4 The schematic diagram of the component modules of the progressive transmission system of the Gaussian splatting volumetric video in a preferred embodiment of the present application.
[0034] Figure 5 The overall flow schematic diagram of the hierarchical representation compression and the progressive transmission of the Gaussian splatting volumetric video in a specific application example of the present application. DETAILED DESCRIPTION
[0035] The embodiments of the present application are described in detail as follows: The embodiments are implemented on the premise of the technical solutions of the present application, and detailed implementation modes and specific operation processes are given. It should be noted that, for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application.
[0036] There are usually long sequence dynamic scene modeling parameter redundancy, low real-time decoding efficiency, and single model cannot adapt to variable bit rate requirements in existing volumetric video compression and transmission. In view of the problem, an embodiment of the present application provides a hierarchical representation compression method of Gaussian splatting volumetric video. The method is based on an end-to-end compression optimization framework of hierarchical 4D Gaussian representation, and realizes high-fidelity reconstruction and efficient compression transmission of dynamic scenes.
[0037] Specifically, as shown in the figure, Figure 1 The hierarchical representation compression method of the Gaussian splatting volumetric video provided by the embodiment can include:
[0038] S1, extracting scene surface geometry, constructing 4D Gaussian initial parameters as the basic representation of dynamic scenes;
[0039] S2, based on the basic representation of dynamic scenes, calculating the importance measure of each Gaussian primitive (i.e. each Gaussian initial parameter), dividing the Gaussian into L layers of progressive details to obtain the hierarchical Gaussian key frame representation;
[0040] S3, based on the hierarchical Gaussian key frame representation, performing inter-frame modeling through rigid transformation and residual compensation;
[0041] S4, calculate the average displacement of the Gaussian primitives, dynamically adjust the grouping length of the inter-frame modeling;
[0042] S5, inject uniform noise to simulate quantization error, and use KDE to estimate the entropy of each attribute, which is used to guide the compression efficiency; meanwhile, construct the weighted loss of end-to-end optimization of distortion, code rate and time smoothness to train the inter-frame modeling;
[0043] S6, encode the trained hierarchical Gaussian parameters into 2D maps and compress them using H.264, to complete the volumetric video compression for dynamic scenes.
[0044] In some preferred embodiments, the above S1 can further include:
[0045] S11, use NeuS2 learning method to extract scene surface geometry information, initialize Gaussian distribution parameters, including: position μ, covariance ∑, opacity α and color c, to obtain 4D Gaussian initial parameters;
[0046] S12, use 4D Gaussian initial parameters to replace traditional neural radiance field as the basic representation of dynamic scenes.
[0047] In some preferred embodiments, the above S2 can further include:
[0048] S21, calculate the importance measure ψ of each Gaussian primitive as:
[0049] ψ = α + λ ψ S
[0050] In the formula, α is the opacity, λ ψ is the weighting coefficient, used to balance the two influence factors, and S is the spatial volume.
[0051] S22, sort the importance measure ψ in descending order, and divide the Gaussians into L layers; among them, the first layer as the base layer retains the core structure of the scene, and the first layer to the L layer as the high layer gradually enhances the details.
[0052] Further, in the above steps, the position represents the coordinates of the Gaussian sphere in the three-dimensional space clock, the covariance represents the size and rotation angle, and the opacity and color are used to calculate the position and color of the Gaussian sphere projection on the plane.
[0053] In some preferred embodiments, the above S3 can further include:
[0054] Use the hierarchical Gaussian key frame representation as the reference basis for training subsequent inter-frame, capture the complex motion of each layer of Gaussian key frame through hierarchical motion modeling, and decompose the frame difference into rigid transformation and residual compensation to maintain time coherence, and perform inter-frame modeling; wherein:
[0055] Decompose frame difference into rigid transformation and residual compensation, including:
[0056] S311, update Gaussian position and orientation with previous frame Gaussian position As input, through multi-resolution hash grid Capture motion features of different scales And then through a lightweight MLP network And Calculate the translation of the current frame And rotation Update the translation of the current frame And rotation Update Gaussian position and orientation, process the overall rigid displacement of the camera or object, estimate the Gaussian rigid transformation between frames;
[0057] S312, on the basis of rigid transformation, predict local detail deformation parameters using residual compensation to obtain scaling residual Opacity residual And color residual Update Gaussian attributes (i.e. Gaussian distribution parameters) using local detail deformation parameters to compensate for non-rigid dynamic details;
[0058] S313, through the hierarchical modeling of rigid transformation and residual compensation, decouple global motion and local details, and complete the inter-frame modeling.
[0059] In some preferred embodiments, the above S4 can further include:
[0060] Set displacement threshold τ μ , according to the average translation of Gaussian Automatic division of group length; when Shorten the group length to improve the reference quality; when Lengthen the group length to optimize the compression efficiency, thereby dynamically adjusting the grouping length of inter-frame modeling.
[0061] In some preferred embodiments, the above S5 can further include:
[0062] S51, introduce uniform noise simulation quantization error, and use KDE to probabilistically model Gaussian attributes to calculate Gaussian attribute entropy rate to guide compression efficiency;
[0063] S52, for optimization of each group of Gaussian distribution parameters, construct a loss function combining color error term L color And entropy rate term L rate_key Form a weighted sum; by optimizing Gaussian distribution parameters layer by layer, the optimal bit rate and reconstruction quality of each level are obtained;
[0064] S53, optimizing the subsequent Gaussians in the group based on the key frame; wherein, only uniform noise is used to simulate quantization error in the rigid motion process, and color error L color Supervision is carried out; color loss L color , entropy rate loss L rate_inter and smoothness loss L reg are introduced in the residual compensation process to form a joint optimization target, enhance temporal consistency and reduce residual amplitude, and further optimize reconstruction quality and storage efficiency.
[0065] In some preferred embodiments, the above S6 can further include:
[0066] Differential quantization and channel tiling are performed on the Gaussian properties of each layer to form a 2D single-channel image sequence that maintains spatial continuity, and the sequence is encoded and compressed by H.264 after being aligned by group, layer and channel.
[0067] Based on the same inventive concept, an embodiment of the present application also provides a hierarchical representation compression system for Gaussian splatting volumetric video.
[0068] Specifically, as Figure 2 shown, the hierarchical representation compression system for Gaussian splatting volumetric video provided by this embodiment can include:
[0069] A parameter initialization module, which is configured to extract scene surface geometry and construct 4D Gaussian initial parameters as a basic representation of a dynamic scene;
[0070] A Gaussian model construction module, which is configured to calculate an importance measure of each Gaussian primitive based on the basic representation of the dynamic scene, divide the Gaussians into L layers of progressive details, perform inter-frame modeling on the layered Gaussians through rigid transformation and residual compensation, and calculate the average displacement of the Gaussian primitives to dynamically adjust the grouping length of the inter-frame modeling;
[0071] A Gaussian model training module, which is configured to inject uniform noise to simulate quantization error and estimate the entropy of each attribute using KDE to guide compression efficiency; and construct a weighted loss of distortion, code rate and temporal smoothness for end-to-end optimization to train the inter-frame modeling;
[0072] A dynamic scene compression module, which is configured to encode the trained layered Gaussian parameters into 2D images and compress them using H.264 to complete volumetric video compression for dynamic scenes.
[0073] The specific implementation of each functional module of the system provided by the above embodiments of the present application will be described in further detail below in conjunction with preferred embodiments.
[0074] The system provided by this preferred embodiment includes the following functional modules:
[0075] a parameter initialization module that extracts surface geometry using NeuS2, constructs 4D Gaussian initial parameters as the basic representation of dynamic scenes;
[0076] a Gaussian model construction module that calculates the importance measure of each Gaussian initial parameter based on the basic representation of dynamic scenes, divides the Gaussian into L layers of progressive details to obtain a layered Gaussian key frame representation, and performs inter-frame modeling on the layered Gaussian key frame representation through rigid transformation and residual compensation, and calculates the average displacement of Gaussian primitives to dynamically adjust the grouping length of inter-frame modeling;
[0077] a Gaussian model training module that injects uniform noise simulation quantization error during training and estimates the entropy of each attribute using KDE to guide the compression efficiency improvement; and simultaneously constructs an end-to-end optimization weighted loss of distortion, code rate and time smoothness;
[0078] a dynamic scene compression module that encodes the trained layered Gaussian parameters into 2D graphs and compresses them using H.264, and selects the number of layers on demand during decoding to achieve progressive transmission and dynamic adaptation.
[0079] In some preferred embodiments, the parameter initialization module has a specific implementation further comprising:
[0080] 4D Gaussian primitives are used to replace traditional neural radiance fields as the basic representation of dynamic scenes. The scene surface geometry information is extracted by NeuS2 to initialize the position μ, covariance ∑, opacity α, color c and other parameters of the Gaussian distribution, and to construct the initial volume representation of the dynamic scene.
[0081] In some preferred embodiments, the Gaussian model construction module has a specific implementation further comprising:
[0082] The importance measure ψ of each Gaussian primitive (fusion space volume S and opacity α) is calculated, and the calculation formula is ψ = α + λ ψ S. The Gaussians are divided into L layers in descending order of ψ. The base layer (e.g., the first layer) retains the core structure of the scene, and the higher layers (e.g., layers 2-6) progressively enhance the details to support progressive real-time rendering with on-demand decoding.
[0083] Based on the layered Gaussian key frame representation, it is used as a reference basis for training subsequent inter-frame, and through layered motion modeling, it effectively captures complex motion and decomposes frame differences into rigid transformation and residual deformation to maintain temporal coherence. Specifically, first estimate the Gaussian rigid transformation between frames, use the position of the previous frame Gaussian as input, capture motion features of different scales through a multi-resolution hash grid Then through a lightweight MLP network and Compute translation of current frame and rotation Update Gaussian position and orientation with it, handle global rigid displacement of camera or object.
[0084] Further, on top of rigid transformation, use residual compensation to predict local detail deformation parameters, including scaling residual opacity residual and color residual for updating Gaussian attributes, compensate for non-rigid dynamic details such as clothing wrinkles, object deformation, etc., reduce visual artifacts and maintain temporal consistency. By hierarchical modeling of global motion and local details through rigid transformation and residual deformation, the global motion is decoupled from the local details, where the rigid transformation handles large-scale displacement and the residual deformation captures local subtle changes, and the combination of the two achieves a compact representation of dynamic scenes, further reducing the redundancy of motion parameters.
[0085] Further, to address the problem of reference invalidation caused by motion accumulation in dynamic scenes and the redundancy and error caused by fixed grouping, a motion-aware adaptive Gaussian grouping is proposed: set a displacement threshold τ μ , automatically divide the group length according to the average translation of the Gaussian When the motion is intense shorten the group length to improve the reference quality, and when the motion is smooth extend the group to optimize the compression efficiency, so as to achieve flexible temporal modeling and efficient reconstruction.
[0086] In some preferred embodiments, the Gaussian model training module described above further includes the following specific implementation:
[0087] During training, an end-to-end entropy optimization training scheme is used to realize the joint optimization of bit rate and reconstruction quality (RD) through differentiable quantization and attribute-specific entropy modeling. Uniform noise is introduced in the training to simulate quantization error, and KDE is used to estimate the entropy of each attribute to guide the improvement of compression efficiency. For the optimization of each group of key reference frames, the loss function combines color error L color and entropy rate term L rate_key to optimize the Gaussian parameters layer by layer to ensure the optimal RD performance of each level.
[0088] Further, based on the key frames, the subsequent Gaussians in the group are optimized to improve the accuracy and compactness of each level, and only simulated quantization is used in the rigid motion process, and L color is relied on for supervision. In the residual optimization process, color loss L color , entropy rate loss L rate_inter and smoothness loss L reg are introduced to enhance temporal consistency and reduce residual amplitude, thereby improving reconstruction quality and storage efficiency.
[0089] In some preferred embodiments, the dynamic scene compression module described above further comprises the following specific implementation manners:
[0090] After the training is completed, the Gaussian attributes of each layer are differentially quantized and tiled in the channel to form a 2D single-channel image sequence that maintains spatial continuity, and are aligned in groups, layers and channels and then compressed by H.264 encoding. The bitstream supporting multi-level quality selection can be generated by encoding and compression, and further adaptive decoding on demand, smooth quality transition and real-time rendering can be realized.
[0091] It should be noted that the steps in the method provided by the present application can be implemented by using corresponding components in the system, and those skilled in the art can refer to the technical solution of the system to implement the step flow of the method, or refer to the technical solution of the method to implement the composition of the system, that is, the embodiments in the system and the embodiments in the method can be understood as preferred examples, which will not be described here.
[0092] Based on the hierarchical representation compression method and system provided in the above embodiments of the present application, an embodiment of the present application further provides a progressive transmission method of Gaussian splash volumetric video realized by using the hierarchical representation compression method or system.
[0093] Specifically, as shown in the following Figure 3 The progressive transmission method of Gaussian splash volumetric video provided by this embodiment can include the following steps:
[0094] M1, based on the hierarchical representation compression method, a bitstream supporting multi-level quality selection is generated;
[0095] M2, the bitstream is decoded according to the required number of layers to realize progressive transmission and dynamic adaptation.
[0096] In the above embodiments of the present application, decoding is the process of restoring the bitstream into Gaussian attributes, including decoding the H.264 compressed code stream into a 2D graph by H.264, and then reassembling it into a hierarchical Gaussian parameter coding. Progressive transmission means that since the Gaussian parameters are explicitly layered, the 2D graph and the bitstream are also layered, and the level can be freely selected during transmission, and the high level can be transmitted after the low level. Dynamic adaptation means that different numbers of levels can be received under different devices, and a trade-off between rendering quality and data processing amount is made.
[0097] Based on the volumetric video compression method and system provided in the above embodiments of the present application, an embodiment of the present application further provides a progressive transmission system of Gaussian splash volumetric video realized by using the hierarchical representation compression method or system.
[0098] Specifically, as shown in the following Figure 4As shown, the embodiment provides a progressive transmission system of Gaussian splatting volume video, which can include:
[0099] A bitstream generation module, which generates a bitstream supporting multi-level quality selection based on a volume video compression method;
[0100] A decoding module, which is used to decode the bitstream according to the required number of layers, to realize progressive transmission and dynamic adaptation.
[0101] The technical solutions provided by the above embodiments of the application will be further described in detail below with reference to a specific application example.
[0102] As described above, volume video provides immersive 3D experience for virtual reality, remote collaboration and other applications due to its free-view navigation capability, but its high bandwidth demand and real-time decoding challenge seriously restricts its wide application. The traditional dynamic neural radiance field method faces the problems of parameter redundancy and low compression efficiency when modeling long sequence dynamic scenes, while the existing 3D Gaussian and its extended methods achieve high-quality rendering, but lack end-to-end joint optimization mechanism in dynamic modeling and compression transmission, resulting in loss of dynamic details and poor compression rate-distortion performance, which is difficult to meet the needs of streaming transmission and adaptive decoding.
[0103] The specific application example adopts the layered representation compression method and the progressive transmission method of Gaussian splatting volume video provided by the above embodiments of the application to realize layered 4D Gaussian representation compression and progressive volume video streaming transmission.
[0104] Specifically, the layered representation compression method and the progressive transmission method of Gaussian splatting volume video involved in the specific application example have the overall process as shown in Figure 5 As shown, it includes the following steps:
[0105] Step S1: Use NeuS2 to extract surface geometry and construct 4D Gaussian initial parameters as the basic representation of dynamic scenes.
[0106] Replace the traditional neural radiance field with 4D Gaussian primitives as the basic representation of dynamic scenes. First, reconstruct high-quality key frames as references for subsequent inter-frame, extract scene surface geometry information through NeuS2, initialize parameters such as position μ, covariance ∑, opacity α, color c of Gaussian distribution, and construct the initial volume representation of dynamic scenes.
[0107] Step S2: Calculate the importance measure of each Gaussian primitive, and divide the Gaussian into L layers of progressive details.
[0108] To solve the problem of high data volume of 3DGS, the importance measure ψ of each Gaussian primitive is calculated (fusion of spatial volume S and opacity α), and the calculation formula is ψ=α+λ ψS. Sort the Gaussians by ψ in descending order, and divide them into L layers. To maintain the balance between volume and opacity, λ ψ is set to 1 x 10 5 , and the Gaussians are divided into L = 6 layers. The base layer (e.g., Layer 1) preserves the core structure of the scene, and the higher layers (e.g., Layers 2-6) progressively enhance the details. The client can dynamically select the number of decoded layers based on network conditions and computational resources, ensuring smooth playback and balancing transmission overhead and reconstruction fidelity. This hierarchical representation also supports progressive rendering, allowing efficient reconstruction of the scene at different levels of detail.
[0109] Step S3: For the layered Gaussians, inter-frame modeling is performed in two steps, including rigid transformation and residual compensation.
[0110] Based on the hierarchical Gaussian keyframe representation, it is used as a reference basis for training subsequent inter-frame, effectively capturing complex motion and decomposing frame differences into rigid transformation and residual deformation to maintain temporal coherence. Specifically, first estimate the inter-frame Gaussian rigid transformation, using the position of the previous frame Gaussians as input, through a multi-resolution hash grid to capture motion features at different scales Then, through a lightweight MLP network and calculate the translation and rotation of the current frame, i.e.:
[0111]
[0112] Update the Gaussians' positions and orientations through equations and to handle the overall rigid displacement of the camera or object.
[0113] Based on the rigid transformation, further use residual compensation to predict local detail deformation parameters, including scale residual opacity residual and color residual Update the Gaussians' attributes through equations to compensate for non-rigid dynamic details such as clothing wrinkles and object deformation, reduce visual artifacts, and maintain temporal consistency. Through hierarchical modeling of rigid transformation and residual deformation, global motion and local details are decoupled, where rigid transformation handles large-scale displacement and residual deformation captures local subtle changes, and their combination realizes compact representation of dynamic scenes, further reducing motion parameter redundancy.
[0114] Step S4: Calculate the average displacement of the Gaussians primitives, and dynamically adjust the grouping length of inter-frame modeling.
[0115] To solve the problem of reference invalidation caused by motion accumulation and the redundancy and error caused by fixed grouping in dynamic scenes, a motion-aware adaptive Gaussian grouping is proposed: set the displacement threshold τ μ , automatically divide the group length according to the Gaussian average translation :
[0116] When (big motion scene, such as fast moving people), short grouping interval (such as every 5 frames a group), avoid the rendering distortion caused by error accumulation;
[0117] When (stable scene): extend the grouping interval (such as every 20 frames a group), reduce the redundancy of motion parameter storage.
[0118] Thus, flexible time modeling and efficient reconstruction are realized.
[0119] Step S5: Inject uniform noise simulation quantization error during training, and use KDE to estimate the entropy of each attribute to guide the compression efficiency improvement; at the same time, construct the weighted loss of end-to-end optimization distortion, code rate and time smoothness.
[0120] During training, an end-to-end entropy optimization training scheme is used to realize the joint optimization of bit rate and reconstruction quality (RD) through differentiable quantization and attribute-specific entropy modeling. In the key frame optimization stage, uniform noise is introduced to enhance the robustness of quantization, where U is a uniform distribution and q is the quantization step size, and KDE (kernel density estimation) is used to model the probability of Gaussian attributes (such as rotation, scaling, color, opacity) and calculate their entropy rate. The loss function is composed of color (luminance) error term L color and entropy rate term L rate_key , and the form is weighted sum:
[0121]
[0122] Where L key represents the total loss of key frame reconstruction, λ l represents the hyperparameter balancing the error of each layer, represents the color reconstruction error of the l-th layer, λ rate_key represents the weighted hyperparameter of key frame entropy loss, represents the entropy estimated by KDE from the same level of Gaussian attributes. Through layer-by-layer optimization, each level is ensured to achieve the optimal rate-distortion (RD) performance, thereby improving the overall compression efficiency and rendering quality.
[0123] After that, the subsequent frames in the group are optimized based on the key frames. For Gaussian position and rotation parameters (i.e. rigid motion process), analog quantization is adopted and only color loss L colorSupervision without introducing entropy constraint; while in the residual deformation stage, to improve accuracy and compactness, further introduce entropy rate loss L rate_inter and temporal smoothing loss L reg , constitute a joint optimization objective:
[0124]
[0125] where L inter is the total loss of non-key frames, is the bit rate of entropy of residual attributes (scale, feature, opacity) estimated from the same level Gaussian attributes by KDE, λ rate_inter is the weighted hyperparameter of key frame entropy loss, is the L1 regularization on residual values to enhance temporal consistency and reduce residual amplitude, λ reg is the weighted hyperparameter of non-key frame entropy loss. This joint training strategy significantly improves the reconstruction quality and compression efficiency, obtaining a compact and high-fidelity layered 4D Gaussian representation, supporting efficient volumetric video compression, storage and transmission.
[0126] Step S6: Encode the trained layered Gaussian parameters into 2D maps and compress them with H.264, and when decoding, select the number of layers as needed to achieve progressive transmission and dynamic adaptation.
[0127] After training, the Gaussian attributes of each level are explicitly separated, and a differential quantization strategy is adopted: higher precision uint16 or uint32 is used for position parameters (such as μ) to reduce error sensitivity, and the remaining attributes (such as scale s, color f, opacity α) are uniformly quantized with uint8. To facilitate compression and transmission, each attribute channel is independently expanded into a 2D single-channel image that maintains spatial continuity, and is organized into a time sequence frame after alignment by group, level and channel. Subsequently, these sequences are efficiently compressed using the H.264 video encoder to generate a progressive bitstream that supports multi-level quality selection. The client can dynamically receive and decode the corresponding level data according to bandwidth and computing power, achieving smooth quality transition, flexible adaptive viewing experience and real-time rendering.
[0128] The layered representation, motion modeling, adaptive grouping and end-to-end entropy optimization are combined to realize the collaborative optimization of high-fidelity reconstruction, efficient compression and transmission. The present application first constructs a perceptually weighted layered 4D Gaussian representation. The Gaussians are divided into L layers by the importance measure of geometric volume and opacity. The base layer preserves the scene structure, and the higher layers progressively enhance the details. The number of layers can be selected on demand during decoding, realizing progressive transmission and dynamic adaptation. For dynamic scenes, a layered motion modeling strategy is proposed. The inter-frame motion is decomposed into rigid transformation and residual deformation. Combined with the motion-aware adaptive grouping mechanism, the reference frame is dynamically updated when the average displacement exceeds the threshold, balancing error accumulation and data redundancy. Through the noise disturbance of the differentiable quantization simulation compression process, the attribute entropy model is constructed by combining the kernel density estimation. The rendering distortion, code rate estimation and time smoothness constraint are integrated into a unified loss function, realizing the joint optimization of representation and compression. Finally, the layered Gaussian features are converted into progressive bit streams through standard video encoding, realizing real-time decoding and rendering on mobile devices. The present application breaks through the fixed code rate limitation of traditional volume video compression, reduces the storage and transmission cost while maintaining high-resolution details, and provides an efficient solution for free-view video streaming, AR / VR content distribution and other scenarios. Experimental results show that the rate-distortion performance of the present application is better than that of existing methods in various dynamic scenes.
[0129] The details of the above embodiments of the present application are known in the art.
[0130] The specific embodiments of the present application are described above. It should be understood that the present application is not limited to the above specific embodiments, and those skilled in the art can make various modifications or changes within the scope of the claims, which does not affect the essential content of the present application.
Claims
1. A hierarchical representation compression method for Gaussian spatters volumetric videos, characterized in that, The method comprises the following steps: extracting scene surface geometry, constructing 4D Gaussian initial parameters as the basic representation of dynamic scenes; based on the basic representation of the dynamic scene, calculating the importance measure of each Gaussian initial parameter, dividing the Gaussian into L layers of progressive details to obtain the layered Gaussian key frame representation; based on the layered Gaussian key frame representation, inter-frame modeling is performed through rigid transformation and residual compensation; calculate the average displacement of the Gaussian primitive, dynamically adjust the grouping length of inter-frame modeling; inject uniform noise to simulate quantization error, and use KDE to estimate the entropy of each attribute to guide the compression efficiency; At the same time, construct the weighted loss of end-to-end optimization of distortion, code rate and time smoothness to train the inter-frame modeling; encode the trained layered Gaussian parameters into 2D graphs and use H.264 compression to complete the volumetric video compression.
2. The hierarchical representation compression method of Gaussian-spraying volumetric videos according to claim 1, characterized in that, The method of extracting scene surface geometry, constructing 4D Gaussian initial parameters as the basic representation of dynamic scenes comprises the following steps: adopting NeuS2 learning method to extract scene surface geometry information, initializing Gaussian distribution parameters, including: position μ, covariance ∑, opacity α and color c, to obtain 4D Gaussian initial parameters; using the 4D Gaussian initial parameters to replace the traditional neural radiance field as the basic representation of dynamic scenes.
3. The hierarchical representation compression method of Gaussian-spraying volumetric videos according to claim 1, characterized in that, The method of calculating the importance measure of each Gaussian primitive and dividing the Gaussian into L layers of progressive details comprises the following steps: The importance measure ψ of each Gaussian primitive is calculated as: ψ = a + λ ψ S where a is the opacity, λ ψ is a weighting factor, and S is the spatial volume. Sort the importance measure ψ in descending order, divide the Gaussian into L layers; The first layer as the basic layer retains the core structure of the scene, and the first layer to the Lth layer as the high layer progressively enhances the details, to obtain the layered Gaussian key frame representation.
4. The hierarchical representation compression method of Gaussian-spraying volumetric videos according to claim 1, characterized in that, The method of inter-frame modeling based on the layered Gaussian key frame representation through rigid transformation and residual compensation comprises the following steps: Take the layered Gaussian key frame representation as the reference basis for training subsequent inter-frame, capture the complex motion of each layer of Gaussian key frame through layered motion modeling, and decompose the frame difference into rigid transformation and residual compensation to maintain time coherence, and perform inter-frame modeling; wherein: The method of decomposing the frame difference into rigid transformation and residual compensation comprises the following steps: Utilizing the previous frame Gaussian position As input, through a multi-resolution hash grid Capture motion features of different scales And then through a lightweight MLP network And Calculate the translation of the current frame And rotation Utilizing the translation And rotation of the current frame Update the Gaussian position and direction, process the overall rigid displacement of the camera or object, and estimate the Gaussian rigid transformation between frames; On the basis of rigid transformation, residual compensation is used to predict local detail deformation parameters to obtain scaling residuals opacity residuals and color residuals Gaussian distribution parameters are updated using the local detail deformation parameters to compensate for non-rigid dynamic details; Through the layered modeling of rigid transformation and residual compensation, the global motion and local details are decoupled, and the inter-frame modeling is completed.
5. The hierarchical representation compression method of Gaussian-spraying volumetric videos of claim 1, wherein, The method of calculating the average displacement of the Gaussian primitive and dynamically adjusting the grouping length of inter-frame modeling comprises the following steps: Setting displacement threshold τ μ , according to Gaussian average translation Automatic division of group length; when Shorten the group length to improve the reference quality; when Lengthen the group length to optimize the compression efficiency, so as to dynamically adjust the grouping length of inter-frame modeling.
6. The hierarchical representation compression method of Gaussian-spraying volumetric videos of claim 1, wherein, The method of injecting uniform noise to simulate quantization error and using KDE to estimate the entropy of each Gaussian attribute to guide the compression efficiency comprises the following steps: At the same time, construct the weighted loss of end-to-end optimization of distortion, code rate and time smoothness to train the inter-frame modeling, which comprises the following steps: Introduce uniform noise to simulate quantization error, and use KDE to probabilistically model the Gaussian attributes to calculate the entropy rate of the Gaussian attributes and guide the compression efficiency. For the optimization of each set of Gaussian distribution parameters, a loss function combining the color error term L color and the entropy rate term L rate_key is constructed, forming a weighted sum; by optimizing the Gaussian distribution parameters layer by layer, the optimal bit rate and reconstruction quality of each layer level are obtained; Optimizing the subsequent Gaussians in the group based on key frames; wherein, only uniform noise is used to simulate quantization error in the rigid motion process, and color error L color is introduced in the residual compensation process to supervise color , entropy rate loss L rate_inter and smoothness loss L reg are introduced to constitute a joint optimization objective, enhance temporal consistency and reduce residual amplitude, and further optimize reconstruction quality and storage efficiency.
7. The hierarchical representation compression method of Gaussian-spraying volumetric videos of claim 1, wherein, The method of encoding the trained layered Gaussian parameters into 2D graphs and using H.264 compression comprises the following steps: Differential quantization and channel tiling are performed on each layer of Gaussian attributes to form a 2D single-channel image sequence that maintains spatial continuity, and then the sequence is aligned by group, level and channel before being encoded and compressed by H.
264.
8. A hierarchical representation compression system for Gaussian spatters volumetric videos, characterized in that, The method comprises the following steps: A parameter initialization module is configured to extract scene surface geometry, and construct 4D Gaussian initial parameters as a basic representation of a dynamic scene. A Gaussian model construction module is configured to calculate an importance measure of each Gaussian initial parameter based on the basic representation of the dynamic scene, divide the Gaussian into L layers of progressive details to obtain a layered Gaussian key frame representation, perform inter-frame modeling on the layered Gaussian key frame representation through rigid transformation and residual compensation, and calculate average displacement of Gaussian primitives to dynamically adjust grouping length of inter-frame modeling. A Gaussian model training module is configured to inject uniform noise simulation quantization error, and estimate entropy of each attribute using KDE to guide compression efficiency. Meanwhile, a weighted loss of end-to-end optimization of distortion, code rate and time smoothness is constructed to train inter-frame modeling. A dynamic scene compression module is configured to encode the trained layered Gaussian parameters into 2D graphs and compress using H.264 to complete volumetric video compression.
9. A progressive transmission method of Gaussian splat volume video realized by using the hierarchical characterization compression method of any one of claims 1-7, characterized in that, The method comprises the following steps: Based on the layered representation compression method, a bitstream supporting multi-level quality selection is generated. The bitstream is decoded according to the required number of layers to realize progressive transmission and dynamic adaptation.
10. A progressive transmission system of Gaussian splat volume video realized by using the hierarchical characterization compression method of any one of claims 1-7, characterized in that, The method comprises the following steps: A bitstream generation module generates a bitstream supporting multi-level quality selection based on the layered representation compression method. A decoding module is configured to decode the bitstream according to the required number of layers to realize progressive transmission and dynamic adaptation.
Citation Information
Patent Citations
On-orbit spectrum calibration method of chromatic dispersion type spectral imaging system
CN104458591A
Hierarchical progressive coding framework method and system for volume video
CN118890487A
Volume video figure rendering method, system and device based on double-layer Gaussian splashing, chip and medium
CN119135950A
Plug flow volume video rendering method and device based on image grid, chip and medium
CN119206024A
Variable-code-rate 4D Gaussian compression method
CN120034657A