Variable-code-rate 4D Gaussian compression method

By adopting a 4D Gaussian compression method with variable bit rate in free-view video, using Gaussian compressed 3DGS representation and motion estimation, combined with implicit entropy model and simulated quantization, the problems of efficient compression and streaming of complex scene videos in transmission are solved, and high-quality and flexible code streams are achieved.

CN120034657AActive Publication Date: 2025-05-23SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510199863.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-23
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and transmit free-view videos of complex scenes while maintaining high quality, especially in sequences with large motion, complex backgrounds and long durations.

Method used

Using a 4D Gaussian compression method with variable code rate, the compressed variable bit rate is achieved by using a complete Gaussian compression 3DGS to represent keyframes, and motion estimation and sparse compensation are performed in non-keyframes, combined with implicit entropy model and simulation quantization.

Benefits of technology

It realizes the support of a wide range of variable bit rates while maintaining high-rate distortion performance, which is suitable for streaming, solving the large capacity and inefficiency of volume video in storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034657A_ABST
    Figure CN120034657A_ABST
Patent Text Reader

Abstract

The invention provides a 4D Gaussian compression method with a variable code rate. The 4D Gaussian compression method comprises the following steps: S1, representing a first frame (key frame) by using a complete 3DGS; s2, using a motion grid and two shared lightweight multi-layer perceptron (MLP) to carry out motion estimation on an inter-frame Gaussian primitive; s3, performing motion compensation on the inter-frame change region by using sparse compensation Gaussian; s4, reconstructing the complete Gaussian expression of the current frame by using the complete Gaussian expression of the previous frame of the buffer area, the motion grid of the current frame and sparse compensation Gaussian, and storing the complete Gaussian expression in the buffer area for reconstruction of the next frame; and S5, quantization and entropy coding are performed on the trained motion grid and sparse compensation Gaussian, the size of the model is compressed, a variable bit rate code stream is obtained, and streaming transmission is realized. The wide variable bit rate is achieved through a single model, meanwhile, the excellent rate distortion performance is kept, and the method can be used for streaming transmission so as to meet different requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of volumetric video representation and compression, and in particular to a variable bit rate 4D Gaussian compression method. Background Art

[0002] Free viewpoint video (FVV) enables immersive real-time navigation of a scene from any viewpoint, enhancing user interactivity and realism, which makes FVV ideal for applications such as entertainment, virtual reality, sports broadcasting, and telepresence. However, streaming and rendering high-quality FVV remains challenging, especially for sequences with large motion, complex backgrounds, and long duration. The main difficulty lies in developing an effective representation and compression method that enables FVV to support streaming at limited bitrates while maintaining high fidelity.

[0003] Traditional FVV reconstruction methods mainly rely on point cloud-based methods and depth-based techniques, which have difficulty in providing high-quality rendering and realism, especially in complex scenes. Neural Radiance Fields (NeRF) and its variants have shown impressive results in reconstructing FVV by learning continuous 3D scene representations, but they have limitations in supporting long sequences and streaming. Recent methods address these issues by compressing explicit features of dynamic NeRFs, however, these methods often suffer from slow training and rendering speeds.

[0004] Recently, 3D Gaussian sputtering (3DGS) has shown superior performance in terms of rendering speed and quality for static scenes compared to NeRF-based methods. Several methods have attempted to extend 3DGS to dynamic environments by incorporating temporal correspondence or temporal dependencies, but these methods require all frames to be loaded into memory for training and rendering, which limits their practicality in streaming applications. 3DGStream models the inter-frame transformation and rotation of 3D Gaussians as a neural transform cache, which reduces the per-frame storage requirement of FVV. However, the overall data size is still large, hindering its ability to support efficient FVV transmission. Although a few studies have explored the compression of dynamic 3DGS, these methods encounter significant difficulties in handling real-world dynamic scenes containing backgrounds, which limits their practical utility. In addition, they independently optimize representation and compression, ignoring the rate-distortion (RD) trade-off during training, which ultimately limits the compression efficiency. Summary of the invention

[0005] In view of the defects in the prior art, an object of the present invention is to provide a 4D Gaussian compression method with a variable bit rate.

[0006] According to one aspect of the present invention, a variable bit rate 4D Gaussian compression method is provided, comprising:

[0007] The key frame, i.e., the first frame, is represented using the complete Gaussian compressed 3DGS. The implicit entropy model is used to estimate the bit rate after Gaussian compressed 3DGS, and analog quantization is introduced to make the compressed size variable bit rate.

[0008] In the first stage of non-key frame training, the motion of the inter-frame Gaussian primitives is estimated using a motion grid and two globally shared lightweight multi-layer perceptrons (MLPs) to obtain a Gaussian expression after motion estimation; the bit rate of the motion grid after compression is estimated using an implicit entropy model, and analog quantization is introduced to make the compressed size a variable bit rate; the motion grid, two globally shared lightweight MLPs and the corresponding implicit entropy model are jointly optimized;

[0009] In the second stage of non-key frame training, based on the Gaussian expression after motion estimation, sparse compensation Gaussian is used to perform motion compensation on the inter-frame change area, an implicit entropy model is used to estimate the bit rate after compression of the sparse compensation Gaussian, and analog quantization is introduced to make the compressed size a variable bit rate; the sparse compensation Gaussian and the corresponding implicit entropy model are jointly optimized;

[0010] Reconstructing the complete Gaussian expression of the current frame using the complete Gaussian expression of the first frame or the previous frame in the buffer, the motion grid of the current frame and the sparse compensation Gaussian, and storing it in the buffer for reconstruction of the next frame;

[0011] quantizing and entropy coding the trained motion grid and the sparse compensation Gaussian, compressing the motion grid and the sparse compensation size, obtaining a code stream with a variable bit rate, and realizing streaming transmission;

[0012] Based on the representation of the first frame, as well as the cyclic first-stage training of non-key frames, the second-stage training of non-key frames, the reconstruction of the current frame, the compression of the trained motion grid and the sparse compensation Gaussian, the compression and streaming of the entire volumetric video are completed.

[0013] Preferably, the use of complete Gaussian compression 3DGS to represent key frames includes:

[0014] Use a set of Gaussian primitives G to represent the entire three-dimensional space scene;

[0015] Each Gaussian unit It includes a set of optimizable parameters {μ; R; f; s; α}, where μ is the center position, R is the rotation matrix, f is the SH coefficient of the view-dependent color c, s is the scaling vector, and α is the transparency;

[0016] For the Gaoski The spatial distribution of a point x in Sure, Where ∑=Rss T R T ;

[0017] When rendering, the rendering color c of a pixel is calculated by alpha blending overlapping Gaussian primitives in depth order, specifically:

[0018]

[0019] Among them, α′ i is the projection of the transparency of the i-th Gaussian basis element on the image plane, c i is the color of the i-th Gaussian primitive in the viewing direction, and N represents the number of Gaussian primitives in the entire three-dimensional space.

[0020] Preferably, the method of using a motion grid and two globally shared lightweight multi-layer perceptrons MLP to perform motion estimation on inter-frame Gaussian primitives to obtain a Gaussian expression after motion estimation includes:

[0021] Position encoding is performed on the center position of the Gaussian primitive in the complete Gaussian expression of the previous frame loaded from the buffer to map it to multiple frequency bands;

[0022] The position is encoded in a multi-resolution motion grid M t Perform trilinear interpolation on the image to generate motion features of different scales;

[0023] The motion features of different scales are connected and input into two shared lightweight multi-layer perceptrons Φ μ and Φ R , obtain the motion estimation of the Gaussian basis element translation and rotation from the previous frame to the current frame;

[0024] Using the obtained motion estimation and the complete Gaussian expression of the previous frame in the buffer, the Gaussian expression G′ after the motion estimation transformation of the current frame is obtained. t .

[0025] Preferably, an implicit entropy model is used to estimate the bit rate after the moving grid or sparse compensation Gaussian compression, and analog quantization is introduced so that the compressed size is a variable bit rate, including:

[0026] By moving the mesh M t Or sparse compensation Gauss ΔG t Add uniform noise To simulate the quantization effect with a step size of q, the quantized value is obtained The size of the added noise is consistent with the quantization step size;

[0027] Use an implicit entropy model to approximate the quantized moving mesh or Sparse Compensated Gaussian The quantitative value of Probability mass function to estimate the compressed bit rate, specifically:

[0028]

[0029] where is the quantization value Probability mass function of is the quantization value Cumulative distribution function of

[0030] Preferably, the joint optimization of the motion grid, two globally shared lightweight MLPs, and the corresponding implicit entropy model includes:

[0031] Taking the weighted sum of the distortion loss and the estimated model bit rate loss as the total loss, and performing backpropagation on the gradient, specifically:

[0032]

[0033] where represents the total loss in the first stage, is the bit rate loss estimated from the quantized motion grid , N is the number of quantization values ; is the photometric loss, c g and are the true and reconstructed colors of the corresponding viewpoints respectively; is the D-SSIM evaluation index between the true image and the rendering result in training, λ 2 is the weight parameter; the parameter λ 1 balances the trade-off between the bit rate and the distortion, and controls the model size and the reconstruction quality.

[0034] Preferably, based on the Gaussian expression after motion estimation, using sparse compensation Gaussian to perform motion compensation on the inter-frame change region includes:

[0035] Based on the Gaussian expression G′ after the motion estimation transformation of the current frame t , identify the sub-optimal regions that need to be compensated, including regions where the gradient exceeds a predetermined threshold, and regions where the translation and rotation of the Gaussian basis elements exceed a predetermined threshold;

[0036] For the regions where the gradient exceeds the predetermined threshold, clone 1 Gaussian basis element located in that region;

[0037] For the regions where the translation and rotation of the Gaussian basis elements exceed the predetermined threshold, clone 2 Gaussian basis elements located in that region, and compress their scales to one percent.

[0038] Preferably, the joint optimization of the sparse compensation Gaussian and the corresponding implicit entropy model includes:

[0039] The weighted sum of the distortion loss and the estimated model bit rate loss is taken as the total loss, and the gradient is back-propagated as follows:

[0040]

[0041] in, is the SH coefficient of the sparse compensation Gaussian after quantization The estimated bit rate loss in the number of is the photometric loss; parameter λ 1 Balance the trade-off between bitrate and distortion, controlling model size and reconstruction quality.

[0042] Preferably, the cloned Gaussian basis ΔG t Normally distributed around the original Gaussian primitives and optimized in the second training phase.

[0043] Preferably, the reconstructing the complete Gaussian expression of the current frame by using the complete Gaussian expression of the first frame or the previous frame in the buffer, the motion grid of the current frame and the sparse compensation Gaussian, and storing it in the buffer for reconstruction of the next frame includes:

[0044] Once the training of the current frame is completed, that is, the first and second stage training is completed, the complete Gaussian representation of the current frame will be reconstructed Specifically:

[0045]

[0046] in and Represent the motion mesh and compensation Gaussian reconstructed for the current frame, respectively. Representative basis right Each Gaussian basis in The μ and R are updated; Stored in the reference buffer for reconstruction of the next frame.

[0047] Preferably, the step of quantizing and entropy coding the trained motion grid and the sparse compensation Gaussian, compressing the motion grid and the sparse compensation size, obtaining a code stream with a variable bit rate, and realizing streaming transmission includes:

[0048] After the above training of each frame is completed, the motion grid and sparse compensation Gaussian are quantized, and the formula is:

[0049]

[0050] Where q is the quantization step size, and x is all the data to be compressed, including the SH coefficient of the Gaussian sphere and the motion grid;

[0051] The quantized result is interval-encoded to obtain a bit stream, which contains the model information of the variable bit rate. The formula is:

[0052] B t =E(Q(q·x)-Q(q·min(x));ω t )

[0053] Among them B t is the code stream corresponding to time t, E represents the entropy encoder, ω t is the distribution of the corresponding data, which is obtained after the training.

[0054] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:

[0055] The variable bit rate 4D Gaussian compression method in the embodiment of the present invention realizes a wide range of variable bit rates through a single model while maintaining excellent rate-distortion performance, and can be used for streaming transmission to meet different requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0057] Figure 1 4D Gaussian compression method with variable bit rate in one embodiment of the present invention. DETAILED DESCRIPTION

[0058] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0059] Volumetric video will play a very important role in people's lives. However, there is currently a lack of a good volumetric video encoding and decoding framework, and there is no method for end-to-end joint optimization of volumetric video representation and compression, which will lead to the loss of dynamic details and reduced compression efficiency. At the same time, volumetric videos trained by a single model are difficult to be flexible, and it is rare to freely choose multiple rate-distortion performances at the decoding and rendering end according to device conditions and personal needs. To address this problem, in one embodiment of the present invention, a variable bit rate 4D Gaussian compression method is provided, specifically, as Figure 1As shown, the main steps are as follows:

[0060] S1: Use the complete Gaussian compressed 3DGS to represent the key frame, i.e. the first frame, use the implicit entropy model to estimate the bit rate after Gaussian compressed 3DGS, and introduce analog quantization to make the compressed size variable bit rate;

[0061] By executing the above step S1, an accurate three-dimensional geometric structure (3DGS) expression of the key frame or the first frame can be obtained, and the reconstruction quality is extremely high. In addition, the 3DGS model is small in size and supports variable bit rate.

[0062] S2: In the first stage of non-keyframe training, the motion of the inter-frame Gaussian primitives is estimated using a motion grid and two globally shared lightweight multi-layer perceptrons (MLPs) to obtain a Gaussian expression after motion estimation; the bit rate of the motion grid after compression is estimated using an implicit entropy model, and analog quantization is introduced to make the compressed size a variable bit rate; the motion grid, two globally shared lightweight MLPs, and the corresponding implicit entropy model are jointly optimized;

[0063] In the aforementioned step S2, based on the previous frame image, a Gaussian ball motion simulation is performed to obtain an initial representation of the non-key frame. The size of the motion grid is small, and variable bit rate encoding is used while ensuring the accuracy of motion estimation. For the parts that have appeared in the scene, the reconstruction quality is extremely high.

[0064] S3: The second stage of non-key frame training, based on the Gaussian expression after motion estimation, uses sparse compensation Gaussian to perform motion compensation for the inter-frame change area, uses the implicit entropy model to estimate the bit rate after sparse compensation Gaussian compression, and introduces analog quantization to make the compressed size variable bit rate; jointly optimizes the sparse compensation Gaussian and the corresponding implicit entropy model;

[0065] In the aforementioned step S3, based on the result of step S2, the motion compensation processing of the Gaussian sphere is performed, thereby obtaining the final representation of the non-key frame. The newly added Gaussian sphere is small in size, uses variable bit rate coding, and has high accuracy of motion compensation, which significantly improves the quality of the entire scene reconstruction.

[0066] S4: reconstructing the complete Gaussian expression of the current frame using the complete Gaussian expression of the first frame or the previous frame in the buffer, the motion grid of the current frame, and the sparse compensation Gaussian, and storing it in the buffer for reconstruction of the next frame;

[0067] In the aforementioned step S4, the final representation of the non-key frame is reconstructed. The representation is small in size and is encoded using a variable bit rate, while ensuring the high quality of scene reconstruction and preparing for the training of subsequent frames.

[0068] S5: quantize and entropy encode the trained motion grid and sparse compensation Gaussian, compress the size of the motion grid and sparse compensation Gaussian, obtain a variable bit rate code stream, and realize streaming transmission.

[0069] Based on the representation of the first frame, as well as the first stage training of non-key frames in loop S2, the second stage training of non-key frames in S3, the current frame reconstruction in S4, and the compressed trained motion grid and the sparse compensation Gaussian in S5, the compression and streaming of the entire volumetric video are completed.

[0070] The above-mentioned embodiments can improve the reconstruction quality and compression rate of volumetric video, improve rate-distortion performance, enhance its flexibility, and enable streaming transmission. The problems of excessive volumetric video storage, poor reconstruction quality, and inability to stream in the prior art are solved. In the above-mentioned embodiments, steps S1, S2, and S3 all use implicit entropy models to accurately calculate the bit rate of each model after compression, and incorporate analog quantization technology to achieve variable bit rate of compressed data size. With this method, the moving grid can have low entropy characteristics, the amount of data after compression encoding can be further reduced, and the robustness of the model to quantization loss is also enhanced.

[0071] In a preferred embodiment, a preferred solution of step S1 is provided, specifically:

[0072] A set of Gaussian primitives G is used as a point cloud to explicitly represent the entire 3D scene.

[0073] Each Gaussian unit Each of them consists of a set of optimizable parameters {μ; R; f; s; α}, where μ is the center position, R is the rotation matrix, f represents the SH coefficient of the view-dependent color c, s is the scaling vector, and α is the transparency. For a point x within a Gaussian basis, its spatial distribution is given by Sure, Where ∑=Rss T R T .

[0074] When rendering, the rendered color c of a pixel is calculated by alpha blending overlapping Gaussians in depth order, specifically:

[0075]

[0076] Among them, α′ i is the projection of the transparency of the i-th Gaussian basis element on the image plane, c i is the color of the i-th Gaussian primitive in the viewing direction. N represents the number of Gaussian primitives in the entire three-dimensional space.

[0077] In order to better represent and compress volumetric video, in a preferred embodiment of the present invention, a preferred solution for the first stage training of volumetric video is provided, that is, step S2, and the following steps can be adopted:

[0078] S2.1 uses a moving grid and two globally shared lightweight multi-layer perceptrons to estimate the inter-frame Gaussian primitives, including:

[0079] The complete Gaussian representation of the previous frame loaded from the buffer The center positions of the mid-gaussian primitives are positionally encoded to map to multiple frequency bands, specifically: Represents positional encoding.

[0080] Encode positions in a multi-resolution motion grid Perform trilinear interpolation on the image to generate motion features of different scales, where L is the resolution level;

[0081] Connect the motion features of different scales and input them into two shared lightweight multi-layer perceptrons Φ μ and Φ R , get the Gaussian basis element translation Δμ from the previous frame to the current frame t and rotation ΔR t Motion estimation achieves accurate motion prediction across multiple scales, captures necessary transformations and effectively reduces inter-frame redundancy, specifically:

[0082]

[0083] Here, interp(·) represents the interpolation operation of the grid.

[0084] Using the obtained motion estimation and the complete Gaussian expression of the previous frame in the buffer, the Gaussian expression G′ after the motion estimation transformation of the current frame can be obtained t , specifically:

[0085]

[0086] Among them, C means including f t-1 ,s t-1 and α t-1 Fixed parameters of Represents the Gaussian expression of the previous frame after reconstruction;

[0087] Represents all Gaussian spheres in the previous frame; represents the result obtained by applying all Gaussian balls of the previous frame to the moving mesh of this frame; μ t-1 Indicates the position of the Gaussian ball in the previous frame, Δμ tThe variable representing the position of the Gaussian ball in this frame relative to the previous frame, ΔR t The variable representing the rotation of the Gaussian ball in this frame relative to the previous frame, R t-1 Indicates the rotation of the Gaussian sphere in the previous frame.

[0088] S2.2, uses a compact implicit entropy model to accurately estimate the bit rate of the motion mesh after compression, and introduces analog quantization to make the compressed size variable bit rate, including:

[0089] Add uniform noise to the moving mesh and the newly added Gaussian primitives To simulate the quantization effect with a step size of q, making the training process robust while preserving the gradient flow;

[0090] Use a tiny and trainable implicit entropy model to approximate the quantized motion mesh or Sparse Compensated Gaussian The quantitative value of The probability mass function of To estimate the bit rate after compression, specifically:

[0091]

[0092] P CDF is the cumulative distribution function. is the quantized data to be compressed. The entropy model can approximate the quantized data to be compressed by calculating the cumulative distribution function (CDF) of hat{y} The probability mass function (PMF) of

[0093] S2.3, Jointly optimize the moving mesh M t , two globally shared lightweight multilayer perceptrons Φ μ and Φ R , and the corresponding implicit entropy model, in the process, the weighted sum of the distortion loss and the estimated model bit rate loss is taken as the total loss, the gradient is back-propagated, and the model is trained, specifically:

[0094]

[0095] in, is the quantized motion grid The estimated bit rate loss in the number of is the luminosity loss, c g and They are the true and reconstructed colors corresponding to the viewing angles; is the D-SSIM evaluation index between the real image and the rendering result in training, λ 2 is the weight parameter; parameter λ 1The trade-off between bitrate and distortion is balanced, thus controlling the model size and reconstruction quality.

[0096] Similarly, in order to better represent and compress the volumetric video, in another preferred embodiment of the present invention, a preferred solution for the second stage training of the volumetric video is provided, that is, step S3, and the following steps can be adopted:

[0097] S3.1, using sparse compensation Gaussian to perform motion compensation for areas with significant changes between frames, including:

[0098] Gaussian expression G′ after the motion estimation transformation of the current frame t Based on the above, suboptimal areas that need to be compensated are identified, mainly areas with significant gradient changes and large Gaussian primitives that undergo large changes in motion estimation;

[0099] For regions with significant gradient changes, when the gradient exceeds a predefined gradient threshold τ g When , clone a Gaussian primitive at that position, denoted as to ensure accurate representation of newly observed elements;

[0100] For larger Gaussian primitives that undergo large transformations in motion estimation, when the translation of the primitive |Δμ t | and rotation |ΔR t When the predefined threshold is exceeded, τ μ and τ R , clone the two Gaussian primitives at that position and compress their scale to one hundredth to To capture detailed motion dynamics more accurately;

[0101] The new compensated Gaussian basis ΔG t Press around the original Gaussian primitive distribution and is optimized in the second training phase.

[0102] S3.2, uses a compact implicit entropy model to accurately estimate the bit rate after sparse compensated Gaussian compression, and introduces analog quantization to make the compressed size variable bit rate, including:

[0103] By adding uniform noise To simulate the quantization effect with a step size of q, making the training process robust while preserving the gradient flow;

[0104] Use a tiny and trainable implicit entropy model to approximate the motion mesh M t Quantized value after simulation quantization The probability mass function of To estimate the bit rate after compression, specifically:

[0105]

[0106] S3.3, Jointly optimize sparse compensation Gaussian ΔG t And the corresponding implicit entropy model, in the process, the weighted sum of the distortion loss and the estimated model bit rate loss is taken as the total loss, the gradient is back-propagated, and the model is trained, specifically:

[0107]

[0108] in, is the SH coefficient of the sparse compensation Gaussian after quantization The estimated bit rate loss in the number of is the photometric loss, which is defined the same as in the first stage; parameter λ 1 The trade-off between bitrate and distortion is balanced, thereby controlling the model size and reconstruction quality. A similar strategy is also applied to the Gaussian representation of the first frame (key frame).

[0109] In a preferred embodiment, step S4: reconstruct the complete Gaussian expression of the current frame using the complete Gaussian expression of the previous frame in the buffer, the motion grid of the current frame and the sparse compensation Gaussian, and store it in the buffer for reconstruction of the next frame, specifically:

[0110] Once the training of the current frame is completed, the complete Gaussian representation of the current frame will be reconstructed Specifically:

[0111]

[0112] in and Represent the motion mesh and compensation Gaussian reconstructed for the current frame, respectively. Representative basis right Each Gaussian basis in The μ and R are updated.

[0113] Of course, you need to Stored in the reference buffer to facilitate reconstruction of the next frame.

[0114] In the above embodiment, the training of the motion mesh and the sparse compensation Gaussian has been completed. In a preferred embodiment, step S5 is implemented: the trained motion mesh and the sparse compensation Gaussian are quantized and entropy encoded, the model size is compressed, and a variable bit rate code stream is obtained to achieve streaming transmission. Specifically:

[0115] After the above training of each frame is completed, the motion grid and sparse compensation Gaussian are quantized, and the formula is:

[0116]

[0117] Where q is the quantization step size, and x represents all the data to be compressed, including the spherical coefficient of the Gaussian ball and the motion grid.

[0118] The quantized result is interval-encoded to obtain a bit stream, which contains the model information of the variable bit rate. The formula is:

[0119] B t =E(Q(q·x)-Q(q·min(x));ω t )

[0120] Among them B t is the code stream corresponding to time t, E represents the entropy encoder, ω t is the distribution of the corresponding data, which is obtained after the training.

[0121] Due to the low entropy and high reconstruction quality of the feature grid, the result obtained after the compression model is small and retains a high reconstruction quality after restoration. The final trained neural radiance field model (i.e., a 4D Gaussian compression framework with variable bit rate) can be evaluated for its reconstruction quality using peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), and its bit rate is evaluated using MB per frame. In order to fully analyze the rate-distortion (RD) performance, Bjontegaard Delta Bit-Rate (BDBR) and Bjontegaard Delta PSNR (BD-PSNR) are used, and the rendering efficiency is evaluated by calculating the number of rendered frames per second (FPS).

[0122] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various modifications or variations within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be used in any combination without conflicting with each other.

Claims

1. A variable bit rate 4D Gaussian compression method, characterized in that: include: The key frame, i.e., the first frame, is represented using the complete Gaussian compressed 3DGS. The implicit entropy model is used to estimate the bit rate after Gaussian compressed 3DGS, and analog quantization is introduced to make the compressed size variable bit rate. In the first stage of non-keyframe training, the motion grid and two globally shared lightweight multi-layer perceptrons (MLPs) are used to estimate the motion of the inter-frame Gaussian primitives and obtain the Gaussian expression after motion estimation. Using an implicit entropy model to estimate the bit rate of the motion mesh after compression, and introducing simulated quantization so that the compressed size is a variable bit rate; jointly optimizing the motion mesh, two globally shared lightweight MLPs and the corresponding implicit entropy model; In the second stage of non-key frame training, based on the Gaussian expression after motion estimation, sparse compensation Gaussian is used to perform motion compensation on the inter-frame change area, an implicit entropy model is used to estimate the bit rate after compression of the sparse compensation Gaussian, and analog quantization is introduced to make the compressed size a variable bit rate; the sparse compensation Gaussian and the corresponding implicit entropy model are jointly optimized; Reconstructing the complete Gaussian expression of the current frame using the complete Gaussian expression of the first frame or the previous frame in the buffer, the motion grid of the current frame and the sparse compensation Gaussian, and storing it in the buffer for reconstruction of the next frame; quantizing and entropy coding the trained motion grid and the sparse compensation Gaussian, compressing the motion grid and the sparse compensation size, obtaining a code stream with a variable bit rate, and realizing streaming transmission; Based on the representation of the first frame, as well as the cyclic first-stage training of non-key frames, the second-stage training of non-key frames, the reconstruction of the current frame, the compression of the trained motion grid and the sparse compensation Gaussian, the compression and streaming of the entire volumetric video are completed.

2. The variable bit rate 4D Gaussian compression method according to claim 1, characterized in that: The key frame is represented by using a complete Gaussian compression 3DGS, including: Use a set of Gaussian primitives G to represent the entire three-dimensional space scene; Each Gaussian It includes a set of optimizable parameters {μ; R; f; s; α}, where μ is the center position, R is the rotation matrix, f is the SH coefficient of the view-dependent color c, s is the scaling vector, and α is the transparency; For the Gaoski The spatial distribution of a point x in Sure, Where ∑=Rss T R T ; When rendering, the rendering color c of a pixel is calculated by alpha blending overlapping Gaussian primitives in depth order, specifically: Among them, α′ i is the projection of the transparency of the i-th Gaussian basis element on the image plane, c i is the color of the i-th Gaussian primitive in the viewing direction, and N represents the number of Gaussian primitives in the entire three-dimensional space.

3. The variable bit rate 4D Gaussian compression method according to claim 1, characterized in that: The motion estimation of the inter-frame Gaussian primitives using the motion grid and two globally shared lightweight multi-layer perceptrons MLPs to obtain the Gaussian expression after motion estimation includes: Position encoding is performed on the center position of the Gaussian primitive in the complete Gaussian expression of the previous frame loaded from the buffer to map it to multiple frequency bands; The position is encoded in a multi-resolution motion grid M t Perform trilinear interpolation on the image to generate motion features of different scales; The motion features of different scales are connected and input into two shared lightweight multi-layer perceptrons Φ μ and Φ R , obtain the motion estimation of the Gaussian basis element translation and rotation from the previous frame to the current frame; Using the obtained motion estimation and the complete Gaussian expression of the previous frame in the buffer, the Gaussian expression G′ after the motion estimation transformation of the current frame is obtained. t .

4. The variable bit rate 4D Gaussian compression method according to claim 3, characterized in that: The bit rate after motion grid or sparse compensation Gaussian compression is estimated using an implicit entropy model, and analog quantization is introduced to make the compressed size variable bit rate, including: By moving the mesh M t Or sparse compensation Gauss ΔG t Add uniform noise To simulate the quantization effect with a step size of q, the quantized value is obtained The size of the added noise is consistent with the quantization step size; Use an implicit entropy model to approximate the quantized moving mesh or Sparse Compensated Gaussian The quantitative value of The probability mass function of To estimate the bit rate after compression, specifically: in For quantitative value The probability mass function of For quantitative value The cumulative distribution function of .

5. The variable bit rate 4D Gaussian compression method according to claim 3, characterized in that: The joint optimization of the motion grid, two globally shared lightweight MLPs and a corresponding implicit entropy model includes: The weighted sum of the distortion loss and the estimated model bit rate loss is taken as the total loss, and the gradient is back-propagated as follows: in, represents the total loss in the first stage, is the quantized motion grid The estimated bit rate loss in the number of is the luminosity loss, c g and They are the true and reconstructed colors corresponding to the viewing angles; It is the D-SSIM evaluation index between the real picture and the rendering result in training, λ2 is the weight parameter; parameter λ1 balances the trade-off between bit rate and distortion, controlling the model size and reconstruction quality.

6. The variable bit rate 4D Gaussian compression method according to claim 1, characterized in that: The method of performing motion compensation on the inter-frame change region based on the Gaussian expression after the motion estimation using a sparse compensation Gaussian includes: Gaussian expression G′ after the motion estimation transformation of the current frame t Based on the above, suboptimal regions that need to be compensated are identified, including regions where the gradient exceeds a predetermined threshold, and regions where the translation and rotation of the Gaussian basis exceed a predetermined threshold; For a region where the gradient exceeds a predetermined threshold, cloning a Gaussian basis element located in the region; For a region where the translation and rotation of the Gaussian primitive exceeds a predetermined threshold, two Gaussian primitives located in the region are cloned and their scales are compressed to one hundredth.

7. The variable bit rate 4D Gaussian compression method according to claim 6, characterized in that: The jointly optimizing the sparse compensation Gaussian and the corresponding implicit entropy model includes: The weighted sum of the distortion loss and the estimated model bit rate loss is taken as the total loss, and the gradient is back-propagated as follows: in, is the SH coefficient of the sparse compensation Gaussian after quantization The estimated bit rate loss in the number of is the photometric loss; the parameter λ1 balances the trade-off between bitrate and distortion, controlling the model size and reconstruction quality.

8. The variable bit rate 4D Gaussian compression method according to claim 7, characterized in that: Cloned Gaussian basis ΔG t Normally distributed around the original Gaussian primitives and optimized in the second training phase.

9. The variable bit rate 4D Gaussian compression method according to claim 1, characterized in that: The method of reconstructing the complete Gaussian expression of the current frame by using the complete Gaussian expression of the first frame or the previous frame in the buffer, the motion grid of the current frame and the sparse compensation Gaussian, and storing it in the buffer for reconstruction of the next frame includes: Once the training of the current frame is completed, that is, the first and second stage training is completed, the complete Gaussian representation of the current frame will be reconstructed Specifically: in and Represent the motion mesh and compensation Gaussian reconstructed for the current frame, respectively. Representative basis right Each Gaussian basis in The μ and R are updated; Stored in the reference buffer for reconstruction of the next frame.

10. The variable bit rate 4D Gaussian compression method according to claim 1, characterized in that: The step of quantizing and entropy coding the trained motion grid and the sparse compensation Gaussian, compressing the motion grid and the sparse compensation size, obtaining a code stream with a variable bit rate, and realizing streaming transmission includes: After the above training of each frame is completed, the motion grid and sparse compensation Gaussian are quantized, and the formula is: Where q is the quantization step size, and x is all the data to be compressed, including the SH coefficient of the Gaussian sphere and the motion grid; The quantized result is interval-encoded to obtain a bit stream, which contains the model information of the variable bit rate. The formula is: B t =E(Q(q·x)-Q(q·min(x));ω t ) Among them B t is the code stream corresponding to time t, E represents the entropy encoder, ω t is the distribution of the corresponding data, which is obtained after the training.

Citation Information

Patent Citations

  • End-to-end video compression method and system based on deep learning and storage medium

    CN111405283A

  • Signal processing method based on deep neural network

    CN112203093A

  • Hierarchical progressive coding framework method and system for volume video

    CN118890487A

  • Motion-compensated compression of dynamic voxelized point clouds

    US20170347120A1

  • Gaussian Mixture Model Entropy Coding

    US20240340425A1

Cited By

  • 4D Gaussian sputtering compression method based on UV mapping

    CN120897069A

  • A 4D Gaussian sputtering compression method based on UV mapping

    CN120897069B

  • Layered representation compression and progressive transmission method of Gaussian splash volume video

    CN120915960A

  • Dynamic scene volume video streaming method, apparatus, and storage medium

    CN122640545A