Multi-view video compression method and application thereof
By constructing voxel space in multi-viewpoint video encoding, initializing anchor information, and optimizing anchor characteristics using neural Gaussian model and quantized offset, combining entropy coding and hash coding, the problem of insufficient compression rate in the existing technology is solved, and higher compression rate and rendering quality are achieved.
Patent Information
- Application Number
- CN202510284768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-11
AI Technical Summary
The existing multi-view video encoding method fails to effectively utilize the correlation between viewpoints, resulting in insufficient encoding compression rate and cannot meet the compression requirements of high-resolution three-dimensional video.
By constructing voxel space, initializing anchor point information, and optimizing anchor point features using neural Gaussian model and quantized offset, combining entropy coding and hash coding, reducing data redundancy and improving compression efficiency.
It achieves higher compression rate and rendering quality, reduces the bit rate required for video transmission, and improves the encoding efficiency of multi-view video.
Smart Images

Figure CN120302045A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to volumetric video encoding and compression, and more specifically, relates to a multi-viewpoint video compression method and its application. Background Art
[0002] In recent years, with the growth of people's demand for digital media vision, the presentation form of multimedia information has undergone great changes: from images to videos, from standard-definition videos to high-definition videos, and from single two-dimensional (2D) videos to three-dimensional (3D) videos with freely switchable viewpoints. Accompanied by this, the amount of multimedia data has shown a geometric growth, and there is an urgent need for a video encoding and compression technology that can efficiently compress high-resolution three-dimensional videos.
[0003] Among many 3D video encodings, the direct encoding method is based on the encoding strategy of traditional 2D video encoding methods, such as simulcast 3D video encoding, which independently compresses each video stream. Taking the unbalanced stereoscopic video encoding method as an example, different quantizers are used for the two viewpoints respectively, so that the qualities of the two viewpoints are different after decoding, and some bitrates can be saved. However, this encoding method does not utilize the correlation between the left and right viewpoints and does not remove the redundant information between viewpoints, so the encoding compression ratio is greatly limited.
[0004] If we want to improve the 3D volumetric video encoding efficiency, we need to consider the temporal correlation of 3D videos and the correlation between viewpoints. The 3D video encoding based on motion estimation and disparity estimation thus emerged. As early as more than a decade ago, MPEG-2 Multiview Profile proposed to utilize this feature to improve the encoding efficiency of 3D videos by combining the cross-correlation between the left and right viewpoints and the spatio-temporal correlation within the same viewpoint. Related to this, in order to eliminate the viewpoint correlation, methods based on the concept of disparity compensation to eliminate the viewpoint correlation, the disparity compensation method in the transform domain, the method of disparity estimation, and their multi-viewpoint video encoding schemes have been proposed. The current multi-viewpoint video encoding standard is based on motion estimation and disparity estimation. In order to improve the compression ratio as much as possible, a viewpoint-temporal pyramid prediction structure based on hierarchical B-frames is adopted during encoding, and this structure is adopted by the official test model JMVM of MVC.
[0005] In the latest MPEG, an Immersive Video Working Group (MIV) has been established. MIV is part of the ISO / IEC 23090 MPEG-I standard. This working group is dedicated to researching and developing encoding standards and technical solutions for immersive video. The standard aims to store and distribute this content through existing and future networks, enabling users to freely select viewing perspectives and directions in a limited viewing space with 6 degrees of freedom (6DoF) for playback. The input of the MIV codec is based on Multi-View Video + Depth (MVD), and each source view provides information represented by frames of geometric information (such as spatial information) and attribute samples (such as texture, transparency, surface normal, reflectivity), along with attributes such as view parameters and camera parameters for 3D reconstruction. A key function of the MIV encoder is to extract patches from the input views, combine them to form one or more attribute and geometry atlases, and eliminate inter-viewpoint redundancy by re-projecting pixels between different views to achieve compression encoding.
[0006] However, the compression ratio of existing methods still cannot meet the current requirements. Summary of the Invention
[0007] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a multi-viewpoint video compression method and its application, aiming to improve the compression ratio of 3D volume video to reduce the bit rate required for video transmission.
[0008] To achieve the above object, according to one aspect of the present invention, a multi-viewpoint video compression method is provided, including:
[0009] S1. Construct a voxel space and initialize the information of each anchor point therein, including static features, dynamic features, and offset and scaling information, based on the sparse point cloud obtained from the initial frame of the multi-viewpoint video to be compressed;
[0010] S2. Add the corresponding quantization offsets to the information of each anchor point at the current time of each anchor point to obtain the information of each anchor point with offsets; sample the anchor points, quantize and entropy-encode the information of each anchor point sample at the current time of each anchor point using the corresponding preset quantization offset parameters, and multiply the sum of the bitstream sizes of the three types of anchor point information of all encoded anchor points by a preset coefficient to obtain the loss function value Loss2;
[0011] Determine the three-dimensional positions of the neural Gaussians of each anchor point using the offset scaling information with offset of each anchor point; respectively input the static features and dynamic features with offset of each anchor point into the fully connected neural network model mlp of the static attributes and dynamic attributes of the neural Gaussians, and correspondingly obtain the static attribute information and dynamic attribute information of the neural Gaussians of each anchor point under the preset time perspective, and bring the dynamic attribute information into the time projection model to obtain the deformation attribute information of the neural Gaussians of each anchor point. The three-dimensional position of each neural Gaussian and each attribute information constitute a neural Gaussian point, which is used to render the image corresponding to the time perspective; compare the image with its real image to obtain the loss function value Loss1;
[0012] S3. Based on the sum of Loss2 and Loss1, optimize the parameters of all mlps and the anchor point information of all anchor points, and re-execute S2 until the termination condition is reached, and then execute S4;
[0013] S4. Perform hash encoding on the positions of each anchor point determined by the sparse point cloud respectively, correspondingly obtain the hash features and input them into the preset context mlp shared by each anchor point to obtain the quantization offset parameters of the anchor point information of each anchor point, which are used to quantize the corresponding anchor point information after iterative optimization and perform entropy encoding to obtain the three anchor point information bitstreams of all anchor points. Take the three anchor point information bitstreams of all anchor points and all mlps after iterative optimization as the coding compression results to complete the three-dimensional volume video coding compression.
[0014] Furthermore, the way to initialize the anchor point information therein is:
[0015] The static features, dynamic features and offset scaling information of each anchor point in the preset voxel space;
[0016] Determine the three-dimensional positions of the neural Gaussians of each anchor point using the offset scaling information of each anchor point; input the static features of each anchor point and the current perspective to be rendered into the fully connected neural network model mlp corresponding to each static attribute to obtain the static attribute information of the neural Gaussians of each anchor point of this kind; the three-dimensional position of each neural Gaussian and each static attribute information constitute a neural Gaussian point in the canonical space;
[0017] Obtain the image of the perspective to be rendered and the frame to be rendered based on all the neural Gaussian points in the canonical space; by comparing the image with its real image, optimize the parameters of all mlps and the static features and offset scaling information of all anchor points, repeat the above operations until the termination condition is reached, and take the static features and offset scaling information of each anchor point in the final voxel space and the preset dynamic features as the initialized anchor point information to complete the initialization of the anchor point information in the voxel space;
[0018] Then, the mlps used in the first iteration in S2 are the mlps corresponding to the neural Gaussian attributes optimized after the initialization of each anchor point information.
[0019] Further, the quantization offset corresponding to each type of anchor information of each anchor point is obtained by multiplying a corresponding preset quantization offset parameter by a random offset within the range of [0, 1], where the random offset is a parameter to be learned and optimized.
[0020] Further, before performing S1, the multi-view video to be compressed is divided into multiple GOPs, and S1-S5 are sequentially executed for each GOP in chronological order to implement encoding and compression. Among them, when compressing and encoding the Nth GOP, N≥2, all the initialized anchor points obtained by S1 corresponding to the first GOP are merged into the set of all initialized anchor points corresponding to the Nth GOP, so as to execute S2-S5 to implement the encoding and compression of the Nth GOP.
[0021] According to another aspect of the present invention, there is provided a multi-view video compression device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements a multi-view video compression method as described above.
[0022] According to another aspect of the present invention, there is provided a multi-view video decoding method, which is characterized by including:
[0023] Receiving the compression result sent by the encoding and compression end, including the bitstreams of three types of anchor information of all anchor points and all mlps, where the compression result is obtained by a multi-view video compression method as described above;
[0024] Decoding and inverse quantizing the bitstreams of three types of anchor information of all anchor points to obtain the static features, dynamic features, and offset and scaling information of each anchor point;
[0025] Using the offset and scaling information of each anchor point to determine the three-dimensional positions of the neural Gaussians of the anchor point; respectively inputting the static features and dynamic features of each anchor point into the fully connected neural network model mlp of the static attributes and dynamic attributes of the neural Gaussians, and correspondingly obtaining the static attribute information and dynamic attribute information of the neural Gaussians of the anchor point under the perspective of the time to be rendered, and bringing the dynamic attribute information into the time projection model to obtain the deformation attribute information of the neural Gaussians of the anchor point. The three-dimensional position of each neural Gaussian and each attribute information constitute a neural Gaussian point, which is used to render a picture of the perspective of the time to be rendered;
[0026] Combining the pictures of different time perspectives to obtain a multi-view video.
[0027] According to another aspect of the present invention, there is provided a multi-view video decoding device, including a memory and a processor, where the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements a multi-view video decoding method as described above.
[0028] According to another aspect of the present invention, there is provided a computer-readable storage medium having stored thereon a computer program, characterized in that when the computer program is executed by a processor, it implements a multi-view video compression method or a multi-view video decoding method as described above.
[0029] According to another aspect of the present invention, there is provided a computer program product comprising a computer program or instructions, characterized in that when the computer program or instructions are executed by a processor, they implement a multi-view video compression method or a multi-view video decoding method as described above.
[0030] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the technical solution provided by the present invention mainly has the following beneficial effects:
[0031] 1. The present invention proposes a multi-viewpoint video compression method. The three-dimensional positions of the neural Gaussians of each anchor point are determined using offset and scaling information. Then, the static features and dynamic features with offsets of each anchor point are respectively input into the fully connected neural network model MLP of the static attributes and dynamic attributes of the neural Gaussians, so as to obtain the static attribute information and dynamic attribute information of the neural Gaussians of each anchor point at the preset time perspective. Through this method, a relatively large number of neural Gaussians can be represented by a relatively small number of anchor point features and offset and scaling information, which can reduce the number of required points. To simulate the motion characteristics of the object, the dynamic attribute information is brought into the time projection model to obtain the deformation attribute information of the neural Gaussians of each anchor point. The three-dimensional position of each neural Gaussian and each attribute information constitute a neural Gaussian point. Using time projection to fit the neural Gaussians can integrate time information into the deformation modeling of the neural Gaussians, making the rendering result more realistic and efficient. To further compress the data, this method performs operations such as quantization and entropy coding on the anchor point information. By performing quantization and entropy coding on the anchor point information, the data storage overhead can be effectively reduced. To optimize the compression quality, this method considers introducing an adaptive quantization offset parameter. For each anchor point, this method introduces anchor point information with a quantization offset. Training with the anchor point information with offsets can dynamically adjust the quantization offset parameter, so as to dynamically adjust the quantization accuracy to adapt to the complexity of different data regions, while maximizing the compression efficiency while maintaining important information. To balance the rendering quality and compression size, this method sets in the training process to render the image corresponding to the time perspective; compare the image with its real image to obtain the loss function value Loss1; and add the bitstream sizes of the three types of anchor point information obtained by quantizing and entropy coding a part of the sampled anchor points and multiply by a preset coefficient to obtain the loss function value Loss2; use Loss2 to directly optimize quantization and entropy coding to minimize the bitstream size, while considering Loss1 to ensure the rendering quality. This method balances the amount of information and the compression ratio during the optimization process, so that the final bitstream can efficiently represent the anchor point information while reducing the impact of data compression on the rendering quality. Therefore, the present invention can improve the compression ratio.
[0032] 2. The present invention also proposes an optimization method for the anchor point information of each anchor point. To optimize the rendering quality, the first frame of the three-dimensional volume video is extracted and modeled through the high-quality 3DGS technology to form an initial anchor point model of the static scene obtained by reconstructing the single first frame as the initialization of the anchor point information for time-domain rendering (i.e., four-dimensional rendering at different times and different perspectives). The initial anchor point model obtained by this method has a great improvement in the accuracy of the initial point and can optimize the rendering quality.
[0033] 3. The present invention also proposes that before executing S1, the multi-view video to be compressed is divided into multiple GOPs, and S1-S5 are sequentially executed for each GOP in chronological order to achieve encoding and compression. Among them, when compressing and encoding the Nth GOP, N≥2, all the initialized anchor points obtained by S1 corresponding to the first GOP are merged into the set of all initialized anchor points corresponding to the Nth GOP, so as to execute S2-S5 to achieve the encoding and compression of the Nth GOP. By establishing the connection between the initial 3DGS point clouds of different GOPs, the GOP jitter during the encoding process is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of a multi-view video compression method provided by an embodiment of the present invention;
[0035] Figure 2 It is a schematic diagram of the co-iterative optimization process of the mlp and each anchor point information provided by an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the encoded diagram after iterative optimization provided by an embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the process of initializing each anchor point information provided by an embodiment of the present invention;
[0038] Figure 5 It is a schematic diagram of the decoding process provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present invention, and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0040] Embodiment 1
[0041] A multi-view video compression method, as Figure 1 shown, includes:
[0042] S1. According to the sparse point cloud obtained from the initial frame of the multi-view video to be compressed, construct a voxel space and initialize the information of each anchor point therein, including static features, dynamic features and offset scaling information;
[0043] S2. Add the corresponding quantization offsets to the respective anchor information of each anchor point at present to obtain the anchor information with offsets; sample the anchor points, quantize and entropy code the respective anchor information of each anchor point sample at present using the corresponding preset quantization offset parameters, add up the bitstream sizes of the three types of anchor information of all the anchor points obtained by coding, and multiply by a preset coefficient to obtain the loss function value Loss2;
[0044] Use the offset scaling information with offsets of each anchor point to determine the three-dimensional positions of the respective neural Gaussians of this anchor point; respectively input the static features and dynamic features with offsets of each anchor point into the fully connected neural network models mlp of the static attributes and dynamic attributes of the neural Gaussians, and correspondingly obtain the static attribute information and dynamic attribute information of the respective neural Gaussians of this anchor point under a preset time perspective, and bring the dynamic attribute information into the time projection model to obtain the deformation attribute information of the respective neural Gaussians of this anchor point. The three-dimensional position and each attribute information of each neural Gaussian form a neural Gaussian point, which is used to render the picture corresponding to the time perspective; compare this picture with its real picture to obtain the loss function value Loss1;
[0045] S3. Based on the sum of Loss2 and Loss1, optimize the parameters of all mlps and the anchor information of all anchor points, and re-execute S2 until the termination condition is reached, and then execute S4;
[0046] S4. Perform hash coding on the respective anchor point positions determined by the sparse point cloud, correspondingly obtain hash features and input them into the preset context mlp shared by each anchor point to obtain the quantization offset parameters of the respective anchor information of the corresponding anchor point, which are used to quantize and entropy code the corresponding anchor information after iterative optimization, obtain the bitstreams of the three types of anchor information of all anchor points, and use the bitstreams of the three types of anchor information of all anchor points and all mlps after iterative optimization as the coding compression result to complete the three-dimensional volume video coding compression.
[0047] In S2, the calculation of loss1 is as follows: The offset-scaled information with offsets of each anchor point is used to determine the three-dimensional positions of the neural Gaussians of each anchor point; the static features with offsets of each anchor point and the current view to be rendered are input into the fully-connected neural network model mlp corresponding to each static attribute, and the static attribute information of the neural Gaussians of each anchor point for this type of static attribute is obtained; the dynamic features with offsets of each anchor point, the current view to be rendered, and the time of the frame to be rendered are input into the mlp corresponding to each dynamic attribute, and the polynomial coefficients of the dynamic attributes of the neural Gaussians of each anchor point for this type of dynamic attribute are obtained; the three-dimensional positions and each attribute information of each neural Gaussian form a neural Gaussian point in the canonical space; based on the static attribute information to be projected and its corresponding dynamic attribute information of each neural Gaussian point in the canonical space, the time projection model of this type of static attribute to be projected is used to project the static attribute information to be projected into the corresponding deformation space of the frame to be rendered, and a neural Gaussian point in the deformation space is obtained; the static attribute to be projected is selected from one or more of the static attributes and the three-dimensional positions; the image of the view to be rendered and the frame to be rendered is obtained based on all the neural Gaussian points in the deformation space; the image is compared with its ground truth image to obtain the loss function value Loss1.
[0048] To reduce the storage and computational costs and improve the efficiency of neural Gaussian rendering, this method uses the offset-scaled information to determine the three-dimensional positions of the neural Gaussians of each anchor point, and then inputs the static features with offsets and the dynamic features of each anchor point into the fully-connected neural network models mlp of the static attributes and dynamic attributes of the neural Gaussians respectively, so as to obtain the static attribute information and dynamic attribute information of the neural Gaussians of each anchor point under the preset time view. By this method, a relatively large number of neural Gaussians can be represented by a relatively small number of anchor point features and offset-scaled information, and the number of required points can be reduced. To simulate the motion characteristics of the object, the dynamic attribute information is brought into the time projection model to obtain the deformation attribute information of the neural Gaussians of each anchor point. The three-dimensional positions and each attribute information of each neural Gaussian form a neural Gaussian point. Using time projection to fit the neural Gaussians can incorporate time information into the deformation modeling of the neural Gaussians, making the rendering result more realistic and efficient. It can be seen that the static features, dynamic features, and offset-scaled information of each anchor point are respectively used to decode the static attribute information, dynamic attribute information, and three-dimensional positions of the corresponding neural Gaussians in the preset canonical space.
[0049] Specifically, in this solution, in order to further compress the data, the method performs operations such as quantization and entropy coding on the anchor point information. By performing quantization and entropy coding on the anchor point information, the data storage overhead can be effectively reduced; in order to optimize the compression quality, the method considers introducing an adaptive quantization offset parameter. For each anchor point, the method introduces anchor point information with a quantization offset, and training with the obtained anchor point information with offsets can dynamically adjust the quantization offset parameter, thereby enabling dynamic adjustment of the quantization accuracy to adapt to the complexity of different data regions, while maximizing the compression efficiency while maintaining important information.
[0050] Regarding the entire S2, in order to balance the rendering quality and the compression size, as Figure 2 shown, during the training process of S2, the picture corresponding to the time perspective obtained by rendering is set; the picture is compared with its true picture to obtain the loss function value Loss1; and the sum of the bitstream sizes of the three types of anchor point information obtained by quantizing and entropy coding some of the sampled anchor points is multiplied by a preset coefficient to obtain the loss function value Loss2; Loss2 is used to directly optimize the quantization and entropy coding to minimize the bitstream size, while considering Loss1 to ensure the rendering quality. This method balances the amount of information and the compression ratio during the optimization process, so that the final bitstream can efficiently represent the anchor point information while reducing the impact of data compression on the rendering quality. Compared with the Loss2 branch that does not sample the anchor points and perform quantization and entropy coding on the anchor point information, if operations such as sampling, quantization, and entropy coding are not performed, more bits are required to store high-precision floating-point numbers, and the information entropy before entropy coding is also higher, resulting in relatively poor compression effect. The scheme of directly storing all anchor point information will lead to a relatively large bitstream size, affecting the storage and transmission efficiency. Through quantization and entropy coding, this method can reduce information redundancy while still maintaining good rendering quality, thereby improving the compression ratio of multi-viewpoint videos.
[0051] In addition, in S2, a preset proportion of anchor points are sampled from all anchor points as anchor point samples, and the current static features, dynamic features, and offset scaling information of each anchor point sample are respectively quantized using corresponding preset quantization offset parameters and the quantization values are entropy coded according to the types of anchor point information to obtain three types of anchor point information bitstreams. During training, in order to improve the training speed, the method samples anchor points and only performs quantization offset parameter quantization and entropy coding on some of the anchor points during training, rather than directly storing all data points, which can effectively reduce the time required for training.
[0052] Specifically for S4, as Figure 3As shown, hash encoding is performed on the positions of each anchor point determined by the sparse point cloud, and the corresponding hash features are obtained and input into a preset context MLP shared by each anchor point to obtain the quantization offset parameters of each anchor point's information, which are used to quantify the corresponding anchor point information after iterative optimization. Entropy encoding is performed on the quantization values of various anchor point information of each anchor point to obtain the bit rates of three types of anchor point information for all anchor points. The bit rates of three types of anchor point information for all anchor points and all MLPs after iterative optimization are used as the encoding compression results.
[0053] In order to save the quantization offset parameters of each anchor point, this method first performs hash encoding on the anchor point positions to obtain compact hash features; then inputs these hash features into the shared context MLP to output the corresponding quantization offset parameters; finally, quantizes the optimized anchor point information and performs entropy encoding to obtain a more compact bitstream. This method further reduces the storage overhead and makes the final encoding result more efficient.
[0054] As a preferred embodiment, as Figure 4 shown, the way to initialize the information of each anchor point is:
[0055] Preset the static features, dynamic features, and offset scaling information of each anchor point in the voxel space;
[0056] Use the offset scaling information of each anchor point to determine the three-dimensional positions of each neural Gaussian of the anchor point; input the static features of each anchor point and the current view to be rendered into the fully connected neural network model MLP corresponding to each static attribute to obtain the information of this static attribute of each neural Gaussian of the anchor point; the three-dimensional position of each neural Gaussian and the information of each static attribute form a neural Gaussian point in the canonical space;
[0057] Based on all the neural Gaussian points in the canonical space, obtain the image of the view to be rendered and the frame to be rendered; by comparing this image with its real image, optimize the parameters of all MLPs and the static features and offset scaling information of all anchor points, and repeat the above operations until the termination condition is reached. Take the static features and offset scaling information of each anchor point in the final voxel space and the preset dynamic features as the initialized information of each anchor point to complete the initialization of the information of each anchor point in the voxel space;
[0058] Then, the MLPs used in the first iteration in S2 are the MLPs corresponding to the neural Gaussian attributes optimized after the initialization of each anchor point's information.
[0059] To optimize the rendering quality, the first frame of the 3D volume video is extracted and modeled using high-quality 3DGS technology to form an initial anchor model of the static scene obtained by reconstructing the individual first frame, which is used as the initialization of the anchor information for temporal rendering (i.e., four-dimensional rendering at different times and different viewpoints). The initial anchor model obtained by this method has a great improvement in the accuracy of the initial point and can optimize the rendering quality.
[0060] As a preferred embodiment, the quantization offset corresponding to each anchor information of each anchor is obtained by multiplying the corresponding preset quantization offset parameter by a random offset within the range of [0, 1], where the random offset is a parameter to be learned and optimized.
[0061] Since the number of frames in a general 3D volume video is large, to ensure the encoding quality, when implementing the method of this embodiment, as a preferred embodiment, the 3D volume video is divided into multiple segments, and the number of frames in each segment meets a preset value, such as 15 frames, which is used as a GOP (i.e., Group Of Picture). When performing encoding operations each time, the volume video is encoded and compressed in units of 15 frames. In addition, when encoding and compressing a 3D volume video, the number of frames corresponding to each GOP can be adjusted according to the actual situation, and they can be the same or different between different GOPs, and the GOP length is not limited.
[0062] Before executing S1, the multi-viewpoint video to be compressed is divided into multiple GOPs, and S1 - S5 are sequentially executed for each GOP in chronological order to achieve encoding and compression. Among them, when compressing and encoding the Nth GOP, N≥2, all the initialized anchors obtained by S1 corresponding to the first GOP are merged into the set of all initialized anchors corresponding to the Nth GOP to execute S2 - S5 to achieve the encoding and compression of the Nth GOP, and by establishing the connection between the initial 3DGS point clouds of different GOPs, the GOP jitter during the encoding process is reduced.
[0063] In addition, it should be noted that the above static attributes may include opacity, color, rotation information, and size. The above attributes to be projected may include three-dimensional position, opacity, and rotation, and the dynamic attributes include position change coefficient, opacity change coefficient, and rotation information change coefficient.
[0064] The mathematical form of each time projection model is preferably a polynomial form.
[0065] Specifically, as a preference, the time projection model of the three-dimensional position is expressed as:
[0066]
[0067] In the formula, μ i(t) represents the three-dimensional position of the i-th neural Gaussian at time t, μ i (t0) represents the three-dimensional position of the i-th neural Gaussian in the canonical space, b i,k represents the polynomial coefficient, k represents the polynomial order, n p represents the upper limit of k.
[0068] Specifically, as an option, the temporal projection model of the opacity is expressed as:
[0069] σ i (t) = σ i (t0) + tanh(s i )(t - t0)
[0070] In the formula, σ i (t) represents the opacity of the i-th neural Gaussian at time t, σ i (t0) represents the opacity of the i-th neural Gaussian in the canonical space, s i represents the opacity change coefficient.
[0071] Specifically, as an option, the temporal projection model of the rotation information is expressed as:
[0072]
[0073] In the formula, q i (t) represents the rotation information of the i-th neural Gaussian at time t, q i (t0) represents the rotation information of the i-th neural Gaussian in the canonical space, k represents the polynomial order, n p represents the upper limit of k, c i,k represents the rotation information change coefficient.
[0074] This embodiment overcomes the defects in the existing 3D video coding technology that new view generation and temporal interpolation cannot be achieved, and provides a pre-trained guided form 3D volumetric video coding and compression method based on the 4DGS technology, thereby reducing the network bandwidth requirements for transmitting a large amount of volumetric video data, and being able to perform video decoding from any perspective in real time at the terminal, improving the user's viewing experience.
[0075] It should be noted that the method of this embodiment is a method that achieves compression by encoding a 3D volumetric video into a compact 3D point cloud. Limited by the performance of dynamic Gaussian 3D solid modeling, the number of input 3D video perspectives should be large enough and the perspectives should be dense enough. Otherwise, the reconstruction quality of the point cloud will also be damaged, resulting in poor video quality in the rendered video. That is, the better the video quality and compression performance will be with more initial perspectives.
[0076] Embodiment 2
[0077] A multi-view video compression device, including a memory and a processor, where the memory stores a computer program, and is characterized in that when the processor executes the computer program, it implements a multi-view video compression method as described in the first embodiment above.
[0078] The related technical solutions are the same as those in the first embodiment and will not be elaborated here.
[0079] Embodiment Three
[0080] A multi-view video decoding method, as Figure 5 shown, includes:
[0081] Receiving the compression result sent by the encoding and compression end, including the three anchor point information bitstreams of all anchor points and all mlps, where the compression result is obtained by a multi-view video compression method as described in the first embodiment above;
[0082] Decoding and inverse quantizing the three anchor point information bitstreams of all anchor points to obtain the static features, dynamic features, and offset and scaling information of each anchor point;
[0083] Using the offset and scaling information of each anchor point to determine the three-dimensional positions of the neural Gaussians of the anchor point; respectively inputting the static features and dynamic features of each anchor point into the fully connected neural network model mlp of the static attributes and dynamic attributes of the neural Gaussians, correspondingly obtaining the static attribute information and dynamic attribute information of the neural Gaussians of the anchor point in the perspective of the time to be rendered, and bringing the dynamic attribute information into the time projection model to obtain the deformation attribute information of the neural Gaussians of the anchor point. The three-dimensional position of each neural Gaussian and each attribute information form a neural Gaussian point, which is used to render a picture in the perspective of the time to be rendered;
[0084] Combining the pictures of different time perspectives to obtain a multi-view video.
[0085] An exemplary implementation manner, as Figure 4 shown, first uses the 3DGS technology to perform saturated reconstruction on the first frame of the three-dimensional video to obtain a high-quality and high-precision 3DGS point cloud model, and then as Figure 2 shown, inputs the model into the 4DGS technology to perform three-dimensional dynamic modeling on the objects in the dynamic space captured by it to obtain a voxel model that can express the dynamic scene, and then as Figure 3 shown, uses the HAC compression algorithm to compress the model to obtain a compressed model. After the compressed model is transmitted through the network, as Figure 5 shown, after decompressing at the terminal to obtain a decompressed voxel model, finally obtain two-dimensional videos and images from any perspective through rendering.
[0086] The related technical solutions are the same as those in the first embodiment and will not be elaborated here.
[0087] Embodiment 4
[0088] A multi-view video decoding device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, a multi-view video decoding method as described in the embodiment is implemented.
[0089] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0090] Embodiment 5
[0091] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a multi-view video compression method as described in the above embodiment or a multi-view video decoding method as described in the above embodiment is implemented.
[0092] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0093] The related technical solutions are the same as above and will not be elaborated here.
[0094] Embodiment 6
[0095] The embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method of the above embodiment of the present application.
[0096] The related technical solutions are the same as above and will not be elaborated here.
[0097] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-viewpoint video compression method, characterized in that, Including: S1. Construct a voxel space and initialize the information of each anchor point therein according to the sparse point cloud obtained from the initial frames of the multi-view video to be compressed, including static features, dynamic features, and offset and scaling information; S2. Add the corresponding quantization offsets to the current information of each anchor point respectively to obtain the information of each anchor point with offsets; sample the anchor points, quantize and entropy-encode the current information of each anchor point sample respectively using the corresponding preset quantization offset parameters, and multiply the sum of the bitstream sizes of the three types of anchor point information of all the encoded anchor points by a preset coefficient to obtain the loss function value Loss2; Determine the three-dimensional positions of the neural Gaussians of each anchor point using the offset and scaling information with offsets of each anchor point; input the static features and dynamic features with offsets of each anchor point into the fully connected neural network models mlp of the static attributes and dynamic attributes of the neural Gaussians respectively, and correspondingly obtain the static attribute information and dynamic attribute information of the neural Gaussians of each anchor point under the preset time perspective, and bring the dynamic attribute information into the time projection model to obtain the deformation attribute information of the neural Gaussians of each anchor point. The three-dimensional position of each neural Gaussian and the attribute information thereof form a neural Gaussian point, which is used to render the picture corresponding to the time perspective; compare this picture with its real picture to obtain the loss function value Loss1; S3. Based on the sum of Loss2 and Loss1, optimize the parameters of all mlps and the information of all anchor points, and re-execute S2 until the termination condition is reached, and then execute S4; S4. Perform hash encoding on the positions of each anchor point determined by the sparse point cloud respectively, correspondingly obtain hash features and input them into the preset context mlp shared by each anchor point to obtain the quantization offset parameters of the information of each anchor point, which are used to quantize the corresponding anchor point information after iterative optimization and perform entropy encoding to obtain the bitstreams of the three types of anchor point information of all anchor points. Take the bitstreams of the three types of anchor point information of all anchor points and all mlps after iterative optimization as the encoding and compression results to complete the 3D volume video encoding and compression.
2. The multi-viewpoint video compression method according to claim 1, wherein The way to initialize the information of each anchor point therein is: Preset the static features, dynamic features, and offset and scaling information of each anchor point in the voxel space; Determine the three-dimensional positions of the neural Gaussians of each anchor point using the offset and scaling information of each anchor point; input the static features of each anchor point and the current perspective to be rendered into the fully connected neural network model mlp corresponding to each static attribute to obtain the static attribute information of the neural Gaussians of each anchor point of this type; the three-dimensional position of each neural Gaussian and the static attribute information thereof form a neural Gaussian point in the canonical space; Obtain the image of the perspective to be rendered and the frame to be rendered based on all the neural Gaussian points in the canonical space; by comparing this image with its real picture, optimize the parameters of all mlps and the static features and offset and scaling information of all anchor points, repeat the above operations until the termination condition is reached, and take the static features and offset and scaling information of each anchor point in the final voxel space and the preset dynamic features as the initialized information of each anchor point to complete the initialization of the information of each anchor point in the voxel space; Then, each MLP adopted in the first iteration of S2 is the MLP corresponding to the neural Gaussian attribute optimized after the initialization of each anchor point information.
3. A multi-viewpoint video compression method according to claim 1, characterized in that The quantization offset corresponding to each type of anchor point information of each anchor point is obtained by multiplying a corresponding preset quantization offset parameter by a random offset within the interval [0, 1], where the random offset is a parameter to be optimized by learning.
4. A multi-view video compression method according to any one of claims 1 to 3, characterized in that, Before executing S1, the multi-view video to be compressed is divided into multiple GOPs, and S1-S5 are sequentially executed for each GOP in chronological order to achieve encoding compression. Among them, when compressing and encoding the Nth GOP, N≥2, all the initialized anchor points obtained by S1 corresponding to the first GOP are merged into the set of all initialized anchor points corresponding to the Nth GOP, so as to execute S2-S5 to achieve the encoding compression of the Nth GOP.
5. A multi-view video compression device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that When the processor executes the computer program, it implements a multi-view video compression method according to any one of claims 1 to 4.
6. A multi-viewpoint video decoding method, characterized in that, Including: Receiving the compression result sent by the encoding and compression end, including the bitstreams of the three types of anchor point information of all anchor points and all MLPs, where the compression result is obtained by a multi-view video compression method according to any one of claims 1 to 4; Decoding and inverse quantizing the bitstreams of the three types of anchor point information of all anchor points to obtain the static features, dynamic features, and offset scaling information of each anchor point; Using the offset scaling information of each anchor point to determine the three-dimensional positions of the neural Gaussians of each anchor point; respectively inputting the static features and dynamic features of each anchor point into the fully connected neural network model MLP of the static attribute and dynamic attribute of the neural Gaussian, and correspondingly obtaining the static attribute information and dynamic attribute information of the neural Gaussians of each anchor point at the time perspective to be rendered. Then, bringing the dynamic attribute information into the time projection model to obtain the deformation attribute information of the neural Gaussians of each anchor point. The three-dimensional position of each neural Gaussian and each attribute information constitute a neural Gaussian point, which is used to render the picture at the time perspective to be rendered; Combining the pictures at different time perspectives to obtain a multi-view video.
7. A multi-view video decoding device, including a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements a multi-view video decoding method according to claim 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a multi-view video compression method according to any one of claims 1 to 4 or a multi-view video decoding method according to claim 6.
9. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, it implements a multi-view video compression method according to any one of claims 1 to 4 or a multi-view video decoding method according to claim 6.
Citation Information
Cited By
Immersive video coding method and system based on 3DGS
CN121547592A