Video encoding method and apparatus for optimizing bitrate allocation

By directly optimizing the variation posterior parameters and quantization step size, the problem that the bit rate allocation in the prior art cannot be achieved is solved, and the rate distortion performance and encoding efficiency of video encoding are improved.

CN115695800BActive Publication Date: 2025-07-22TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211132263.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-07-22
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

The existing video encoding technology cannot achieve accurate pixel-level allocation in code rate allocation, resulting in insufficient encoding efficiency and quality.

Method used

By directly optimizing the variational posterior parameters and quantization step sizes without relying on the empirical model, using improved rate distortion cost expressions and gradient descent optimization methods, the pixel-level code rate allocation is achieved.

Benefits of technology

The rate distortion performance of video encoding is improved, the encoding quality is improved, and the code rate is reduced, and the code rate is realized at the pixel level is improved, which is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115695800B_ABST
    Figure CN115695800B_ABST
Patent Text Reader

Abstract

The present invention provides a video encoding method and apparatus for optimizing bitrate allocation, including: using the first variational posterior parameter obtained by evenly distributed variational inference for each video frame of a group of pictures as the initial value of the second variational posterior parameter corresponding to the video frame; then, by using an improved rate-distortion cost expression, taking the minimum rate-distortion cost of the group of pictures as the objective, performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame of the group of pictures to obtain the optimal second variational posterior parameter and quantization interval for each video frame of the group of pictures; and encoding each video frame of the group of pictures by using the optimal quantization interval and the optimal second variational posterior parameter of each video frame of the group of pictures. The present invention directly optimizes the variational posterior parameter and quantization step without relying on an empirical model, realizes bitrate allocation at the pixel level, and thus achieves the effect of saving bitrate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video compression, and particularly to a video encoding method and apparatus for optimizing bitrate allocation. Background Art

[0002] Bitrate allocation is an important topic in video encoding and decoding, which is the task of allocating bitrates to different frames / regions according to minimizing the bitrate cost. Among them, in the rate-distortion cost R + λD, D represents the total distortion, that is, the error between the video reconstruction and the video source; R represents the total bitrate, which determines the bandwidth / storage cost of the video encoding bitstream in the channel / storage medium; λ represents the Lagrange multiplier that weighs the total distortion D and the total bitrate R.

[0003] Currently, there are roughly three categories of bitrate allocation schemes. The first category constructs a rate-Lagrange multiplier model (R-λ model) and a distortion-Lagrange multiplier model (D-λ model) based on the rate-distortion model (R-D model). At the same time, based on the characteristics of video frames, a rate dependency model and a quality dependency model are empirically established. Then, based on these empirical models, the optimal bitrate allocation is solved. However, this type of scheme is applicable to traditional non-differentiable encoders and performs poorly on deep learning encoders. Moreover, deep learning encoders are differentiable everywhere, and the use of empirical models is neither accurate nor necessary. The second category follows the rate-Lagrange multiplier model (R-λ model), the distortion-Lagrange multiplier model (D-λ model) and the rate dependency model, and further proposes a quality dependency model applicable to deep learning encoders. Then, based on these empirical models, the optimal bitrate allocation is solved. However, this type of scheme also has the problem that the use of empirical models is neither accurate nor necessary. The third category uses an encoder parameter optimization method to optimize the overall rate distortion for each GoP. However, optimizing the encoder parameters rather than the direct parameters of the variational posterior makes the optimization process unable to perceive the spatial distribution of pixels, resulting in the bitrate allocation being limited to the frame level and unable to achieve pixel-level bitrate allocation.

[0004] In summary, the bitrate allocation technology for video encoding and decoding needs to be further improved. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a video encoding method and apparatus for optimizing bitrate allocation, which directly optimizes the variational posterior parameters and quantization step sizes without relying on empirical models, thereby achieving pixel-level bitrate allocation.

[0006] In a first aspect, the present invention provides a video coding method for optimizing bitrate allocation, the method comprising:

[0007] When encoding each group of pictures in the target video, using an encoder to perform amortized variational inference on each video frame within the group of pictures to obtain first variational posterior parameters of each video frame within the group of pictures;

[0008] Initializing second variational posterior parameters of each video frame within the group of pictures to its own first variational posterior parameters, and performing gradient descent optimization on the second variational posterior parameters and quantization intervals of each video frame within the group of pictures with the minimum rate-distortion cost of the group of pictures as the target to obtain optimal second variational posterior parameters and optimal quantization intervals of each video frame within the group of pictures;

[0009] Using the optimal quantization intervals and optimal second variational posterior parameters of each video frame within the group of pictures to implement the encoding of each video frame within the group of pictures;

[0010] Wherein, the rate-distortion cost of the group of pictures is calculated using an improved rate-distortion cost expression;

[0011] The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameters of each video frame of the group of pictures to a second bitstream obtained by quantizing the second variational posterior parameters of each video frame of the group of pictures based on the quantization intervals of each video frame of the group of pictures.

[0012] According to the video coding method for optimizing bitrate allocation provided by the present invention, changing the first bitstream obtained by quantizing the second variational posterior parameters of each video frame of the group of pictures to a second bitstream obtained by quantizing the second variational posterior parameters of each video frame of the group of pictures based on the quantization intervals of each video frame of the group of pictures specifically means:

[0013] Changing the first bitstream obtained by quantizing the second variational posterior parameter y of the i-th video frame of the group of pictures i to the second bitstream obtained by quantizing the second variational posterior parameter y of the i-th video frame of the group of pictures based on the quantization interval Δ i of the i-th video frame of the group of pictures i wherein, i ∈ (1 to k), k represents the total number of video frames within the group of pictures,

[0014] and represents the rounding operation symbol. represents the rounding operation symbol.

[0015] According to the video coding method for optimizing bitrate allocation provided by the present invention, the minimum rate-distortion cost is equivalent to the maximum variational evidence lower bound;

[0016] The maximum variational evidence lower bound is specifically calculated according to the following formula:

[0017]

[0018] wherein, represents the entropy model parameterized by and represents the decoder corresponding to the encoder, represents the set of the second bitstreams corresponding to the first to the (i - 1)-th video frames of the picture group, represents the set of the second bitstreams corresponding to the first to the i-th video frames of the picture group.

[0019] According to the video coding method for optimized bitrate allocation provided by the present invention, the is specifically calculated according to the following formula:

[0020]

[0021] wherein, represents the cumulative distribution function of

[0022] According to the video coding method for optimized bitrate allocation provided by the present invention, the encoding of each video frame within the picture group is implemented by using the optimal quantization interval and the optimal second variational posterior parameter of each video frame within the picture group, and includes:

[0023] Performing entropy coding on to obtain the encoding result of the i-th video frame within the picture group;

[0024] Decoding the picture group includes:

[0025] Performing entropy decoding on the encoding result of the i-th video frame within the picture group to obtain the corresponding entropy decoding result;

[0026] Based on using the encoder to decode the entropy decoding result to obtain the reconstructed frame of the i-th video frame within the picture group;

[0027] wherein, represents the optimal second variational posterior parameter of the i-th video frame within the picture group, represents the optimal quantization interval of the i-th video frame within the picture group, represents the rounding operation symbol.

[0028] According to the video encoding method for optimizing bitrate allocation provided by the present invention, during the process of performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame within the group of pictures with the minimum rate-distortion cost of the group of pictures as the target, Gumbel Softmax is used to perform differentiable relaxation on the rounding operation.

[0029] According to the video encoding method for optimizing bitrate allocation provided by the present invention, before performing entropy decoding on the encoding result of the i-th video frame within the group of pictures, it further includes:

[0030] Storing the as a 32-bit floating-point number, and using a channel to transmit the encoding result of the i-th video frame within the group of pictures and the to the decoding end.

[0031] In a second aspect, the present invention provides a video encoding device for optimizing bitrate allocation, and the device includes:

[0032] An amortized encoding module, when encoding each group of pictures in the target video, uses an encoder to perform amortized variational inference on each video frame within the group of pictures to obtain the first variational posterior parameter of each video frame within the group of pictures;

[0033] An optimal solution module, initializes the second variational posterior parameter of each video frame within the group of pictures to its own first variational posterior parameter, and performs gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame within the group of pictures with the minimum rate-distortion cost of the group of pictures as the target to obtain the optimal second variational posterior parameter and optimal quantization interval of each video frame within the group of pictures;

[0034] An optimized encoding module, uses the optimal quantization interval and optimal second variational posterior parameter of each video frame within the group of pictures to implement the encoding of each video frame within the group of pictures;

[0035] wherein, the rate-distortion cost of the group of pictures is calculated using an improved rate-distortion cost expression;

[0036] The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures to a second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures based on the quantization interval of each video frame in the group of pictures.

[0037] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the video encoding method for optimizing bitrate allocation as described in the first aspect.

[0038] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the video encoding method for optimizing bitrate allocation as described in the first aspect.

[0039] The video encoding method and apparatus for optimizing bitrate allocation provided by the present invention improve SAVI, enabling it to directly optimize the variational posterior parameters and quantization intervals during bitrate allocation, and the effect is equivalent to optimizing the Lagrangian multipliers λ={λ1, λ2…λ k} corresponding to k video frames within a group of pictures. In practical applications, AVI is used to calculate the first variational posterior parameter of each video frame within a group of pictures under the Lagrangian multiplier λ0. On this basis, the improved SAVI is used to encode each video frame within the group of pictures, achieving pixel-level bitrate allocation without relying on an empirical model, thereby improving the rate-distortion performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is a flowchart of the video encoding method for optimizing bitrate allocation provided by the present invention;

[0042] Figure 2 is a comparison diagram between the video encoding method provided by the present invention and the traditional deep learning video encoding method;

[0043] Figure 3 is a schematic diagram of the cooperation between the encoder and the semi-averaging provided by the present invention;

[0044] Figure 4 is a structural diagram of the video encoding apparatus for optimizing bitrate allocation provided by the present invention;

[0045] Figure 5 is a structural diagram of an electronic device for implementing the video encoding method for optimizing bitrate allocation provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0047] The following elaborates on the abbreviations specific to this field.

[0048] I-frame: Intra-frame, an intra-coded frame that does not require reference to other frames during encoding and decoding.

[0049] P-frame: Predictive-frame, a forward-predicted frame that requires reference to the previous I-frame or P-frame during encoding and decoding.

[0050] GoP: Group of Pictures, a group of pictures formed by the pictures between two adjacent I-frames.

[0051] bpp: bits per pixel, the average number of bits required to encode the color information of each pixel.

[0052] PSNR: Peak Signal-to-Noise Ratio, a peak signal-to-noise ratio, which is an objective metric used to measure the quality of image reconstruction and is defined as where MAX I is the maximum value representing the image color, and MSE is the mean square error between the original image and the reconstructed image. The unit of PSNR is decibel (dB).

[0053] Variational Posterior: Let the original image distribution be p(x), and the latent variable distribution be p(y|x). However, p(y|x) itself is unsolvable due to its high dimensionality. The variational posterior is q*(y|x) = argmin DKL(q(y|x)||p(y|x)). Here, q(y|x) is a set of solvable distributions, and q*(y|x) is the distribution among them that minimizes the KL divergence (DKL) between q(y|x) and p(y|x). This distribution is called the variational posterior and can often be solved by optimization methods.

[0054] VI: Variational Inference, is the process of converting the original image distribution into the variational posterior q(y|x).

[0055] AVI: Amortized Variational Inference. It refers to a type of VI where the same set of parameters is used for all original images. Most of the work belongs to AVI.

[0056] SAVI: Semi - Amortized Variational Inference. It refers to a VI that first initializes the latent variable distribution parameters using AVI and then optimizes the distribution for each image using optimization.

[0057] Gumbel - Softmax (Gumbel - Normalized Exponent): Aims to solve the problem of gradient propagation in discontinuous posterior distributions and can be used in SAVI for discrete variational posteriors.

[0058] The following combines Figures 1 - 5 to describe the video encoding method and device for optimizing bitrate allocation of the present invention.

[0059] In a first aspect, the present invention provides a video encoding method for optimizing bitrate allocation, as Figure 1 shown, the method includes:

[0060] S11, when encoding each group of pictures in the target video, use an encoder to perform amortized variational inference on each video frame within the group of pictures to obtain the first variational posterior parameter of each video frame within the group of pictures;

[0061] The present invention provides a pair of encoder - decoder, and are the encoder parameters and decoder parameters respectively. The encoder and decoder make full use of a specific Lagrange multiplier λ0. When encoding the i - th frame x i ∈R H×W with H×W pixels in a group of pictures with k frames, first use the encoder to obtain the variational posterior parameter (i.e., latent variable) y i of the i - th frame of video, and the formula is Here, the encoder can be understood as an inference model, represents to In the encoder, by quantifying y i to obtain After that, an parameterized entropy model is used as the probability mass function. After having the probability mass function, the bitrate is used for encoding

[0062] Correspondingly, during the decoding process, x i is reconstructed to obtain the reconstructed frame The formula is When d(·,·) is the mean square error, the distortion can be written as

[0063] Since amortized variational inference uses the same Lagrange multiplier λ0, the calculation formula for the rate-distortion cost is as follows:

[0064]

[0065] When encoding each frame of video in the picture group, the same λ0 is used, which means that the encoding between each frame of video does not affect each other. For each frame of video, the rate-distortion cost (R-D cost) is optimal. However, because there is inter-frame dependence during the encoding of the picture group, when using the same λ0 to encode each frame or region of the picture group, it is sub-optimal for the entire picture group. Based on this, the subsequent work of the present invention is carried out.

[0066] S12, initialize the second variational posterior parameter of each video frame in the picture group as its own first variational posterior parameter, and perform gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame in the picture group with the minimum rate-distortion cost of the picture group as the target, so as to obtain the optimal second variational posterior parameter and optimal quantization interval of each video frame in the picture group;

[0067] Among them, the rate-distortion cost of the picture group is calculated using an improved rate-distortion cost expression;

[0068] The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame of the picture group to the second bitstream obtained by quantizing the second variational posterior parameter of each video frame of the picture group based on the quantization interval of each video frame of the picture group;

[0069] The goal of video coding is to minimize the error D between the video reconstruction and the video source to improve the viewing experience, and to minimize the bit rate R in the channel / storage medium to reduce the bandwidth / storage cost. Therefore, bit rate allocation is performed with the minimum rate-distortion cost of the picture group as the target.

[0070] The present invention expects to use different R-D trade-off parameters λ i i ∈ (1~k) to encode the video frames {x1~x k} in the picture group, achieve pixel-level bit allocation, and the video frames are limited to use decoder parameters Decodable.

[0071] For this reason, the present invention proposes an equivalent bit rate allocation scheme, which uses a quantization interval Δ i i ∈ (1~k) to enhance the bitstream i ∈ (1 ~ k), and directly optimize y i and Δ i . This is equivalent to constructing an improved SAVI encoder.

[0072] The working principle of the SAVI encoder is as follows: By default, the quantization interval used in fully amortized variational inference is 1, and the group of pictures x containing k frames of video 1:k has a corresponding quantization interval of Δ 1:k ;

[0073] Then, the bitstream 1:k corresponding to the group of pictures x containing k frames of video is subjected to quantization interval enhancement to obtain an enhanced bitstream After that, replace in with Replace with Replace with to obtain the rate - distortion cost calculation formula of the SAVI encoder;

[0074] Among them,

[0075] When applied, assume that the group of pictures x 1:k , and the variational posterior parameters obtained using fully amortized variational inference are

[0076] Take as the initial variational posterior parameters of the first video frame to the k - th video frame within the group of pictures. On this basis, with the goal of minimizing the rate - distortion cost of the group of pictures, continuously perform gradient descent to solve for the quantization interval Δ i and the variational posterior parameter y i to seek the global optimal solution. In practice, the number of gradient descent steps is about 2000 to converge.

[0077] S13. Use the optimal quantization interval and the optimal second variational posterior parameter of each video frame within the group of pictures to implement the encoding of each video frame within the group of pictures.

[0078] The introduction of the quantization interval Δ actually expands the entropy model (enhances the probability mass function of the entropy model), which is reflected in the decoder as applying an enhancement parameter to the decoder parameters (fixed value), making the decoder decode more flexibly.

[0079] Figure 2 This invention's video coding method is compared with traditional deep - learning video coding methods, such as Figure 2As shown in the figure, an improved semi-amortized encoder is added after the encoder at the sending end of the present invention for semi-amortized variational inference (SAVI). The improved semi-amortized encoder uses a cyclic iteration method for bitrate allocation, directly controlling the variational posterior parameter y and the quantization interval Δ, thereby improving the video quality. The changed time complexity is mainly reflected at the sending end. In fact, no change is required at the receiving end, nor will it increase the decoder complexity. However, this scheme can significantly reduce the loss D and the bitrate R. On average, this scheme can improve the PSNR by 1 dB or reduce the bitrate by 25%.

[0080] The video coding method for optimizing bitrate allocation provided by the present invention improves SAVI and directly optimizes the variational posterior parameter and the quantization interval during bitrate allocation, and its effect is equivalent to optimizing the Lagrangian multipliers λ = {λ1, λ2... λ k} corresponding to k video frames within the group of pictures. In practical applications, the first variational posterior parameter of each video frame within the group of pictures is calculated using AVI under the Lagrangian multiplier λ0. On this basis, each video frame within the group of pictures is encoded using the improved SAVI, and pixel-level bitrate allocation is achieved without relying on an empirical model, improving the rate-distortion performance.

[0081] Based on the above embodiments, as an alternative embodiment, changing the first bitstream obtained by quantifying the second variational posterior parameter of each video frame of the group of pictures to the second bitstream obtained by quantifying the second variational posterior parameter of each video frame of the group of pictures based on the quantization interval of each video frame of the group of pictures specifically includes:

[0082] Changing the first bitstream obtained by quantifying the second variational posterior parameter y i of the i-th video frame of the group of pictures to the second bitstream obtained by quantifying the second variational posterior parameter y i of the i-th video frame of the group of pictures based on the quantization interval Δ i of the i-th video frame of the group of pictures

[0083] where i ∈ (1~k), k represents the total number of video frames within the group of pictures, represents the rounding operation symbol.

[0084] Specifically, the unquantified variational posterior parameter y i , when quantifying at an interval of 1, can obtain Performing inverse quantization on can obtain

[0085] The unquantified variational posterior parameter y i , when quantifying at an interval of Δ i , can obtain Performing inverse quantization on Inverse quantization can obtain

[0086] Therefore, the pre-enhancement bitstream is The post-enhancement bitstream is

[0087] The present invention enhances the quantization interval of the latent variable bitstream to enhance the probability mass function of the entropy model, and further improves the rate-distortion performance.

[0088] Based on the above embodiments, as an alternative embodiment, the minimum rate-distortion cost is equivalent to the maximum variational evidence lower bound;

[0089] The maximum variational evidence lower bound is specifically calculated according to the following formula:

[0090]

[0091] where represents the entropy model parameterized by , represents the decoder corresponding to the encoder, represents the set of the second bitstreams corresponding to the first video frame to the (i - 1)-th video frame of the picture group, represents the set of the second bitstreams corresponding to the first video frame to the i-th video frame of the picture group.

[0092] Specifically, is equivalent to -R i , is equivalent to D i , so maximizing is equivalent to minimizing the R - D cost.

[0093] The present invention gives the optimization objective of the improved semi-amortized encoder, which is convenient for the iterative optimization of the variational posterior parameters and the quantization interval.

[0094] Based on the above embodiments, as an alternative embodiment, the is specifically calculated according to the following formula:

[0095]

[0096] where represents cumulative distribution function of.

[0097] In the encoding and decoding process of the encoder ( parameterized) - decoder ( parameterized), The improved semi-average sharing encoder replaces with (i.e., ) and replaces to obtain expansion formula.

[0098] The provision of the above formula facilitates the calculation of the code rate of the improved semi-average sharing encoder.

[0099] Based on the above embodiment, as an alternative embodiment, the encoding of each video frame in the image group by using the optimal quantization interval and the optimal second variational posterior parameter of each video frame in the image group includes:

[0100] Perform entropy encoding on to obtain the encoding result of the i-th video frame in the image group;

[0101] That is, use as the quantization interval to quantize to obtain Then perform entropy encoding on to obtain the corresponding binary code stream.

[0102] Decoding the image group includes:

[0103] Perform entropy decoding on the encoding result of the i-th video frame in the image group to obtain the corresponding entropy decoding result;

[0104] Based on Use the encoder to decode the entropy decoding result to obtain the reconstructed frame of the i-th video frame in the image group;

[0105] where represents the optimal second variational posterior parameter of the i-th video frame in the image group, represents the optimal quantization interval of the i-th video frame in the image group, represents the rounding operation symbol.

[0106] That is, perform entropy decoding on the binary code stream to obtain

[0107] Multiply by to obtain

[0108] Use the decoder to decode to obtain the reconstructed frame of x i ​

[0109] The improvement of the present invention is mainly reflected in the encoding process, and the decoding process is adaptively changed without increasing the complexity of the decoder.

[0110] Based on the above embodiments, as an alternative embodiment, in the process of performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame in the image group with the minimum rate-distortion cost of the image group as the target, Gumbel Softmax is used to perform differentiable relaxation on the rounding operation.

[0111] Figure 3 An example of the schematic diagram of the cooperation between the encoder and the semi-amortized is shown. Since the rounding operation is non-differentiable, and the improved semi-amortized encoder searches for the global optimal solution through the gradient descent method, which requires differentiability everywhere. Therefore, the present invention introduces Gumbel Softmax to perform differentiable relaxation on the rounding operation to ensure the normal progress of gradient descent optimization.

[0112] Based on the above embodiments, as an alternative embodiment, before performing entropy decoding on the encoding result of the i-th video frame in the image group, it further includes:

[0113] Storing the as a 32-bit floating-point number, and using a channel to transmit the encoding result of the i-th video frame in the image group and the to the decoding end.

[0114] The present invention stores the as a 32-bit floating-point number because the bit rate it occupies can be ignored, reducing the bandwidth / storage cost.

[0115] In summary, the present invention has the following advantages:

[0116] First: Directly optimize the variational posterior and quantization interval, which is a very fine-grained parameterization of the bitstream. This scheme can fully cover all possibilities of the bitstream and is a pixel-level bit allocation.

[0117] Second: It does not rely on the empirical R-D model and completely starts from the gradient properties of the differentiable codec itself to achieve the optimal bit rate allocation.

[0118] Next, the video encoding device for optimizing the bit rate allocation provided by the present invention will be described. The video encoding device for optimizing the bit rate allocation described below can be mutually corresponding and referenced with the video encoding method for optimizing the bit rate allocation described above. Figure 4 An example of the structural schematic diagram of the video encoding device for optimizing the bit rate allocation is shown, as Figure 4 shown, the device includes:

[0119] The equalized coding module 21, when coding each group of pictures in the target video, uses an encoder to perform equalized variational inference on each video frame in the group of pictures, and obtains the first variational posterior parameter of each video frame in the group of pictures;

[0120] The optimal solution module 22 initializes the second variational posterior parameter of each video frame in the group of pictures as its own first variational posterior parameter, and performs gradient descent optimization on the second variational posterior parameter and the quantization interval of each video frame in the group of pictures with the minimum rate-distortion cost of the group of pictures as the target, and obtains the optimal second variational posterior parameter and the optimal quantization interval of each video frame in the group of pictures;

[0121] The optimized coding module 23 uses the optimal quantization interval and the optimal second variational posterior parameter of each video frame in the group of pictures to implement the coding of each video frame in the group of pictures;

[0122] Among them, the rate-distortion cost of the group of pictures is calculated using an improved rate-distortion cost expression;

[0123] The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures to the second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures based on the quantization interval of each video frame in the group of pictures.

[0124] On the basis of the above embodiments, as an alternative embodiment, the changing the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures to the second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the group of pictures based on the quantization interval of each video frame in the group of pictures is specifically:

[0125] Changing the first bitstream obtained by quantizing the second variational posterior parameter y of the i-th video frame in the group of pictures i to the second bitstream obtained by quantizing the second variational posterior parameter y of the i-th video frame in the group of pictures based on the quantization interval Δ of the i-th video frame in the group of pictures i where i ∈ (1~k), k represents the total number of video frames in the group of pictures, i and represents the rounding operation symbol.

[0126] where i ∈ (1~k), k represents the total number of video frames in the group of pictures, and represents the rounding operation symbol.

[0127] On the basis of the above embodiments, as an alternative embodiment, the minimum rate-distortion cost is equivalent to the maximum variational evidence lower bound;

[0128] The maximum variational evidence lower bound is specifically calculated according to the following formula:

[0129]

[0130] wherein, represents an entropy model parameterized by a parametric entropy model, represents the decoder corresponding to the encoder, represents the set of second bitstreams corresponding to the first to the (i - 1)-th video frames of the picture group, represents the set of second bitstreams corresponding to the first to the i-th video frames of the picture group.

[0131] Based on the above embodiments, as an alternative embodiment, the is specifically calculated according to the following formula:

[0132]

[0133] wherein, represents the cumulative distribution function of

[0134] Based on the above embodiments, as an alternative embodiment, the optimization encoding module is specifically configured to:

[0135] perform entropy encoding on to obtain the encoding result of the i-th video frame within the picture group;

[0136] The apparatus further includes a decoding module, and the decoding module is configured to decode the picture group, including:

[0137] an entropy decoding unit configured to perform entropy decoding on the encoding result of the i-th video frame within the picture group to obtain a corresponding entropy decoding result;

[0138] a reconstruction unit configured to, based on use the encoder to decode the entropy decoding result to obtain the reconstructed frame of the i-th video frame within the picture group;

[0139] wherein, represents the optimal second variational posterior parameter of the i-th video frame within the picture group, represents the optimal quantization interval of the i-th video frame within the picture group, represents the rounding operation symbol.

[0140] Based on the above embodiments, as an alternative embodiment, during the process of performing gradient descent optimization on the second variational posterior parameters and quantization intervals of each video frame in the image group with the minimum rate-distortion cost of the image group as the target, Gumbel Softmax is used to perform differentiable relaxation on the rounding operation.

[0141] Based on the above embodiments, as an alternative embodiment, the apparatus further includes a transmission module, configured to, before performing entropy decoding on the encoding result of the i-th video frame in the image group, store the as a 32-bit floating-point number, and transmit the encoding result of the i-th video frame in the image group and the to the decoding end through a channel.

[0142] Figure 5 Illustrates a schematic physical structure diagram of an electronic device, as Figure 5 shown. The electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a video encoding method for optimizing bitrate allocation. The method includes: when encoding each image group in the target video, using an encoder to perform amortized variational inference on each video frame in the image group to obtain the first variational posterior parameter of each video frame in the image group; initializing the second variational posterior parameter of each video frame in the image group as its own first variational posterior parameter, and performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame in the image group with the minimum rate-distortion cost of the image group as the target to obtain the optimal second variational posterior parameter and optimal quantization interval of each video frame in the image group; using the optimal quantization interval and optimal second variational posterior parameter of each video frame in the image group to implement the encoding of each video frame in the image group; where the rate-distortion cost of the image group is calculated using an improved rate-distortion cost expression; the improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the quantized image group to a second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the image group based on the quantization interval of each video frame in the image group.

[0143] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0144] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video coding method for optimizing bitrate allocation provided by each of the above methods. The method includes: when encoding each group of pictures (GOP) in the target video, using an encoder to perform amortized variational inference on each video frame in the GOP to obtain the first variational posterior parameter of each video frame in the GOP; initializing the second variational posterior parameter of each video frame in the GOP to its own first variational posterior parameter, and performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame in the GOP with the minimum rate-distortion cost of the GOP as the target to obtain the optimal second variational posterior parameter and optimal quantization interval of each video frame in the GOP; using the optimal quantization interval and optimal second variational posterior parameter of each video frame in the GOP to implement the encoding of each video frame in the GOP; wherein, the rate-distortion cost of the GOP is calculated using an improved rate-distortion cost expression; the improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the GOP to the second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the GOP based on the quantization interval of each video frame in the GOP. On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the video coding method for optimizing bitrate allocation provided by each of the above methods. The method includes: when encoding each group of pictures (GOP) in the target video, using an encoder to perform amortized variational inference on each video frame in the GOP to obtain the first variational posterior parameter of each video frame in the GOP; initializing the second variational posterior parameter of each video frame in the GOP to its own first variational posterior parameter, and performing gradient descent optimization on the second variational posterior parameter and quantization interval of each video frame in the GOP with the minimum rate-distortion cost of the GOP as the target to obtain the optimal second variational posterior parameter and optimal quantization interval of each video frame in the GOP; using the optimal quantization interval and optimal second variational posterior parameter of each video frame in the GOP to implement the encoding of each video frame in the GOP; wherein, the rate-distortion cost of the GOP is calculated using an improved rate-distortion cost expression; the improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameter of each video frame in the GOP to the second bitstream obtained by quantizing the second variational posterior parameter of each video frame in the GOP based on the quantization interval of each video frame in the GOP.The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video coding method for optimizing bitrate allocation, characterized in that, The method includes: When encoding each group of pictures (GOP) in the target video, using an encoder to perform amortized variational inference on each video frame within the GOP to obtain the first variational posterior parameters of each video frame within the GOP; Initializing the second variational posterior parameters of each video frame within the GOP to their respective first variational posterior parameters, and performing gradient descent optimization on the second variational posterior parameters and quantization intervals of each video frame within the GOP with the minimum rate-distortion cost of the GOP as the objective to obtain the optimal second variational posterior parameters and optimal quantization intervals of each video frame within the GOP; Using the optimal quantization intervals and optimal second variational posterior parameters of each video frame within the GOP to achieve the encoding of each video frame within the GOP; Wherein, the rate-distortion cost of the GOP is calculated using an improved rate-distortion cost expression; The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameters of each video frame of the GOP to a second bitstream obtained by quantizing the second variational posterior parameters of each video frame of the GOP based on the quantization intervals of each video frame of the GOP.

2. The video encoding method for optimizing bitrate allocation according to claim 1, wherein The changing the first bitstream obtained by quantizing the second variational posterior parameters of each video frame of the GOP to a second bitstream obtained by quantizing the second variational posterior parameters of each video frame of the GOP based on the quantization intervals of each video frame of the GOP is specifically: The second variational posterior parameter y of the i-th video frame of the quantized picture group i The obtained first bitstream is changed to be based on the quantization interval Δ of the i-th video frame of the picture group i The second variational posterior parameter y of the i-th video frame of the quantized picture group i The obtained second bitstream where \(i\in(1\sim k)\), and \(k\) represents the total number of video frames in the image group. represents the rounding operation symbol.

3. The video coding method for optimizing bitrate allocation according to any one of claims 1 or 2, characterized in that The minimum rate-distortion cost is equivalent to the maximum variational evidence lower bound; The maximum variational evidence lower bound Specifically, it is calculated according to the following formula: Among them, represents an entropy model parameterized by a parameter, represents the decoder corresponding to the encoder, represents the set of second bitstreams corresponding to the first to the (i - 1)-th video frames of the group of images, represents the set of second bitstreams corresponding to the first to the i-th video frames of the group of images, x i represents the i-th video frame of the group of images, i ∈ (1 to k), where k represents the total number of video frames in the group of images, represents the rounding operation symbol, λ0 is the Lagrange multiplier, d(·, ·) is the mean square error, and log represents the natural logarithm.

4. The video encoding method for optimizing bitrate allocation according to claim 3, characterized in that The said It is specifically calculated according to the following formula: Among them, denotes the cumulative distribution function.

5. The video coding method for optimizing bitrate allocation according to claim 1, wherein The using the optimal quantization intervals and optimal second variational posterior parameters of each video frame within the GOP to achieve the encoding of each video frame within the GOP includes: Perform entropy coding to obtain the coding result of the i-th video frame within the image group; Decoding the GOP, including: Performing entropy decoding on the encoding result of the i-th video frame within the GOP to obtain the corresponding entropy decoding result; Based on Use the encoder to decode the entropy decoding result to obtain a reconstructed frame of the i-th video frame within the picture group; Among them, represents the optimal second variational posterior parameter of the i-th video frame within the image group, represents the optimal quantization interval of the i-th video frame within the image group, represents the rounding operation symbol.

6. The video coding method for optimizing bitrate allocation according to claim 4, characterized in that, During the process of performing gradient descent optimization on the second variational posterior parameters and quantization intervals of each video frame within the GOP with the minimum rate-distortion cost of the GOP as the objective, using Gumbel Softmax to perform differentiable relaxation on the rounding operation.

7. The video encoding method for optimizing code rate allocation according to claim 5, wherein Before performing entropy decoding on the encoding result of the i-th video frame within the GOP, it further includes: Store the as a 32-bit floating-point number, and use a channel to transmit the encoding result of the i-th video frame in the image group and the to the decoding end.

8. A video encoding device for optimizing bitrate allocation, characterized in that, The apparatus includes: An amortized encoding module, which when encoding each GOP in the target video, uses an encoder to perform amortized variational inference on each video frame within the GOP to obtain the first variational posterior parameters of each video frame within the GOP; An optimal solution module, which initializes the second variational posterior parameters of each video frame within the GOP to their respective first variational posterior parameters, and performs gradient descent optimization on the second variational posterior parameters and quantization intervals of each video frame within the GOP with the minimum rate-distortion cost of the GOP as the objective to obtain the optimal second variational posterior parameters and optimal quantization intervals of each video frame within the GOP; An optimized encoding module, which uses the optimal quantization intervals and optimal second variational posterior parameters of each video frame within the GOP to achieve the encoding of each video frame within the GOP; Among them, the rate-distortion cost of the image group is calculated using an improved rate-distortion cost expression; The improved rate-distortion cost expression changes the first bitstream obtained by quantizing the second variational posterior parameters of each video frame in the image group to a second bitstream obtained by quantizing the second variational posterior parameters of each video frame in the image group based on the quantization interval of each video frame in the image group.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the video coding method for optimizing bitrate allocation according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the video coding method for optimizing bitrate allocation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Code rate control method for stabilizing video quality

    CN101754003A

  • Image compression method and image compression device

    CN113014927A