Three-dimensional wavelet scalable video rate control method based on video content perceptual characteristics and system thereof
By using an R-λ model based on video content-aware characteristics and an adaptive Lagrange multiplier selection algorithm, the problem of unavailable quantization parameter QP in a 3D wavelet scalable video coding system is solved, achieving precise bitrate control and improved rate-distortion performance while maintaining the continuity of video quality.
Patent Information
- Application Number
- CN202411201807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2044-08-29
AI Technical Summary
In existing 3D wavelet scalable video coding systems, the quantization parameter QP is unavailable, making it difficult to accurately establish the RQ model. The coding performance of the R-ρ model degrades under multiple transform sizes. The selection algorithm for the Lagrange multiplier λ affects the bitrate distortion performance and coding efficiency of the video encoding and decoding system, making it difficult to achieve precise bitrate control.
An R-λ model based on video content awareness characteristics is adopted. An adaptive Lagrange multiplier λi_SSIM is constructed by combining the wavelet filter type, subband coupling phenomenon and temporal subband content awareness characteristics through the adaptive Lagrange multiplier selection algorithm. The R-λ model is established, and the bitrate control of the three-dimensional wavelet scalable video is realized through adaptive bitrate control technology.
Without increasing complexity, precise bitrate control was achieved, improving the rate-distortion performance of the codec, maintaining the continuity of video quality, and obtaining smoother reconstructed video quality.
Smart Images

Figure CN119788862B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video encoding and decoding technology, and relates to bitrate control of three-dimensional wavelet scalable video, specifically to a three-dimensional wavelet scalable video bitrate control method and system based on video content perception characteristics. Background Technology
[0002] The primary objective of rate control is to achieve the highest possible reconstructed video quality within a given target bitrate, thereby meeting the requirements of channel transmission bandwidth and storage device space. Rate control is crucial for stable and reliable video transmission, and researchers can design appropriate rate control strategies for different application scenarios. Reasonable and effective rate control techniques can, on the one hand, improve channel utilization while ensuring video quality; on the other hand, they are of great significance for overcoming bandwidth limitations and reducing storage costs. One of the core issues in rate control is estimating the bitrate-distortion function of the encoded video sequence.
[0003] Rate control technology has evolved alongside video coding technology. Each generation of video coding standards recommends its corresponding rate control algorithm integrated into the test model, such as TM5 for MPEG-2, TMN8 for H.263, VM8 for MPEG-4, JM for H.264 / AVC, and HM for HEVC. In earlier video coding standards like H.261 and MPEG-1, rate control was achieved solely by adjusting the coding quantization step size using buffer state feedback. To precisely control the output bitrate, the rate control algorithms recommended by H.263, MPEG-4, H.264 / AVC, HEVC, and the AVS video coding standard all employ rate-distortion models. Currently, rate control models are mainly classified into three categories: RQ model, R-ρ model, and R-λ model.
[0004] (1) For the RQ model, the rate control method using the RQ model requires obtaining the quantization parameter QP. However, in the 3D wavelet scalable video coding system, the quantization parameter QP is unavailable. It exists implicitly in the bit plane truncation stage rather than being explicitly expressed, making it very difficult to establish an RQ rate control model. Furthermore, in the 3D wavelet-based scalable video coding system, all motion vector information is retained without compression. Only the residual information is encoded and compressed, and bit allocation is performed according to the sub-bit plane generated by encoding, thereby achieving compression of low-frequency and high-frequency sub-bands. However, quantization is only effective for residual information, so it is difficult to accurately establish a model describing the relationship between non-residual information (such as motion vector information, coding mode information, etc.) and QP. Therefore, adjusting the quantization parameter QP has little significance for the bit overhead of non-residual information. In addition, the quantization parameter QP can only be selected as an integer, and the quantization step size Qstep doubles every 6 increments of QP. Moreover, QP can only be selected as some discrete values, which also restricts the development of the accuracy of achieving the target bit rate through the quantization parameter QP.
[0005] (2) For the R-ρ model, the R-ρ model bit rate control method has good coding performance under the condition of using fixed size transformation. However, applying the R-ρ model to the three-dimensional wavelet video coding framework that supports multiple transformation sizes leads to a decrease in coding performance. Therefore, the bit rate control algorithm based on the R-ρ model is difficult to apply to the three-dimensional wavelet video coding framework.
[0006] (3) For the R-λ model, the Lagrange multiplier λ is a key factor in controlling the bit rate. It can calculate the parameter λ through adaptive bit rate allocation and parameter update. Since the algorithm for selecting λ directly affects the bit rate distortion performance and coding efficiency of the video encoding and decoding system, it is essential to select the optimal λ for encoding in order to achieve a more accurate bit rate control effect. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a three-dimensional wavelet scalable video bitrate control method and system based on video content perception characteristics. Without significantly increasing complexity, it achieves precise bitrate control, significantly improves the rate-distortion performance of the codec, better maintains the continuity of video quality, and obtains more stable reconstructed video quality.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] On the one hand, the present invention provides a three-dimensional wavelet scalable video bitrate control method based on video content perception characteristics, specifically including the following steps:
[0010] Step 1: In the motion compensation time-domain filtering stage, based on the influence of wavelet filter type, subband coupling phenomenon, and time-domain subband content awareness characteristics on the propagation of reconstruction errors, an adaptive Lagrange multiplier selection algorithm for the wavelet domain is constructed, and the optimal adaptive Lagrange multiplier λ for the wavelet domain is selected. i_SSIM The adaptive Lagrange multiplier selection algorithm is used to adaptively select Lagrange multipliers for each time-domain sub-band.
[0011] Step 2: Establish an R-λ model based on the perceptual characteristics of video content;
[0012] Step 3: Based on the R-λ model, the bitrate of the three-dimensional wavelet scalable video is controlled by adaptive bitrate control technology.
[0013] Furthermore, the specific implementation process of step 1 is as follows:
[0014] Step 1.1: Based on the wavelet basis coefficients selected for time-domain wavelet decomposition and the time-domain subband coupling phenomenon, construct the distortion relationship model between the time-domain subband frame and the reconstructed video frame, i.e., the following formula (7);
[0015] Step 1.2: Calculate the temporal sub-band weight factor based on the mutual information value of the luminance components of adjacent sub-band frames, the gradient value of each sub-band frame and its texture consistency. The calculation process is shown in the following formula (8).
[0016] Step 1.3: Adaptively select the current optimal wavelet domain adaptive Lagrange multiplier based on the structural similarity index of video sub-band frames.
[0017] Furthermore, the adaptive selection method based on the structural similarity index of video sub-band frames in step 1.3 yields the rate-distortion optimization objective function J. SSIM (F i )for:
[0018]
[0019] In formula (17), Weighting distortion for high-frequency subbands; Weighting distortion for low-frequency subbands; λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; For high-frequency subband SSIM distortion; This is for low-frequency subband SSIM distortion.
[0020] Furthermore, λ i_SSIM The calculation formula is as follows:
[0021]
[0022]
[0023] In formula (9), is an empirical parameter for different video sequences, with a value of 6.2; MAD is the mean absolute error between the original subband frame and the reconstructed subband frame; η is a parameter related to the temporal subband activity factor; R T ω is the target bitrate. i Let MSE be the weight factor for the i-th temporal sub-band frame; in formula (16), MSE = (MAD) 2 ; c2 is the variance of the residual pixel values before wavelet transform; c2 is a positive constant; λ i Let be the Lagrange multiplier of the i-th frame.
[0024] Furthermore, λ i The acquisition process is as follows:
[0025]
[0026] In formula (6), d i For the distortion of the i-th temporal subband frame, r i β is the bitrate of the i-th temporal subband frame; q These are parameters related to wavelet transform; is the variance of the residual pixel values before wavelet transform; Q is the total number of wavelet coefficients; k′ and η are parameters related to the temporal subband activity factor, which can be adaptively adjusted according to different video content and different temporal subband activity.
[0027] Where, β q These are parameters for different wavelet decomposition levels after wavelet transform, tailored to different frame modes. Their specific values are as follows: β q (q=1,2,3)
[0028]
[0029] in, Then formula (6) simplifies to:
[0030]
[0031] Then the Lagrange multiplier λ of the i-th frame i The calculation formula is:
[0032]
[0033] In formula (8), is an empirical parameter for different video sequences, with a value of 6.2; MAD is the mean absolute error between the original subband frame and the reconstructed subband frame; η is a parameter related to the temporal subband activity factor; R T ω is the target bitrate. i is the weighting factor for the i-th time-domain sub-band frame.
[0034] Furthermore, in step 2, the R-λ model is:
[0035]
[0036] In formula (18), λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; c1 is the variance of the residual pixel values before wavelet transform; c2 is a positive constant; MSE is the mean square error of the original sub-band frame and the reconstructed sub-band frame; R T Target bitrate; η is an empirical parameter for different video sequences, with a value of 6.2; η is a parameter related to the temporal subband activity factor; ω i is the weighting factor for the i-th temporal sub-band frame.
[0037] Further, step 3, based on the R-λ model, implements bitrate control for the three-dimensional wavelet scalable video using adaptive bitrate control technology, specifically including:
[0038] Step 3.1: GOP-level bit allocation, the target bit for each GOP level. GOP The calculation formula is:
[0039]
[0040]
[0041] In formula (19), Bit avg_frame The average number of bits per video frame is calculated using formula (20); N coded N is the number of frames that have been encoded. GOP N is the number of video frames contained within a GOP; SW is the sliding window size, SW≥N GOP Bit coded R is the total number of bits consumed across all encoded frames; in formula (20), R tar The target bitrate is the target bitrate; framerate is the frame rate.
[0042] Step 3.2, Frame-level bit allocation, frame-level target bit frame The calculation is obtained according to formula (21), that is, the remaining bits in the GOP are allocated to the remaining frames according to the weight of different video frames;
[0043] Bit frame = (Bit GOP -Coded GOP )×ω frame (twenty one)
[0044] In formula (21), Bit GOP The target number of bits at the GOP level; Coded GOP ω represents the number of bits consumed to encode the current GOP. frame The weight coefficients for the current video frame are calculated using formula (22);
[0045]
[0046] In formula (22), E frame Energy of the current frame; E GOP The current GOP energy is represented by W; the video frame width is represented by H; the video frame height is represented by I(x,y); and the pixel value at (x,y) is represented by I(x,y).
[0047] Step 3.3: Basic Unit-Level Bit Allocation. Define the basic unit as different time-domain sub-bands after MCTF time-domain decomposition, and the target bit at the basic unit level. sub The result is obtained according to formula (23):
[0048] Bit sub = (Bit frame -Bit header -Coded frame )×ω sub (twenty three)
[0049] In formula (23), Bit frame The target number of bits for the current frame; Bit header The number of bits consumed for encoding header information; Coded frame ω represents the number of bits already consumed in the current frame. sub These are the weighting coefficients for the time-domain sub-band;
[0050] Step 3.4: Implement bitrate control for 3D wavelet scalable video using an adaptive bitrate control model.
[0051] Furthermore, in step 3.4, the adaptive rate control model is calculated using formula (24):
[0052] λ=α·2 βR (twenty four)
[0053] In formula (24), α and β are model parameters related to the characteristics of the video sequence, which can be obtained through adaptive updating, i.e.:
[0054] α new =α old +2δ·(lnλ real -lnλ theory )×α old (25)
[0055] β new =β old +2δ·(lnλ real -lnλ theory )·Rln2 (26)
[0056] In formula (25), α old The value of α before the update; α new The updated value of α; λ real The actual λ value used during encoding; λ theory λ is the theoretical value of λ obtained from the R-λ model; δ is the perturbation parameter, and δ∈(0,1); in formula (26), R is the actual number of encoded bits; β old The value of β before the update; β new The updated β value; λ real The actual λ value used during encoding; λ theory λ is the theoretical value of λ obtained from the R-λ model; δ is the perturbation parameter, and δ∈(0,1).
[0057] On the other hand, the present invention also provides a system based on the three-dimensional wavelet scalable video bitrate control method described above in part or in whole, including an encoder, a stream extractor, and a decoder connected in sequence; at the encoder end, the original video is first analyzed by the MCTF analysis module, and the temporal sub-band weights are obtained according to the obtained temporal sub-bands, and then an adaptive λ selection module is used; the temporal sub-bands are processed by the spatial domain analysis module to form a three-dimensional wavelet spatiotemporal sub-band, and then the bitstream information is jointly organized into a fully embeddable coded bitstream by the quantization module, the 3D-EZBC encoding module and the motion vector lossless encoding module; the stream extractor processes the received fully embeddable coded bitstream according to the process of "constructing an R-λ model → bit allocation → bitrate control" to obtain a scalable coded bitstream; the decoder decodes the received scalable coded bitstream to generate a reconstructed video.
[0058] Furthermore, in the system, the proportion r of disconnected pixels is used. u The number of MCTF decomposition levels is adjusted adaptively, specifically as follows:
[0059] Define r m r represents the proportion of multi-connected pixels. u r represents the proportion of unconnected pixels. c r represents the proportion of connected pixels; where r m≥0, r u ≥0, r c ≥0, and r m +r c +r u =1; where r u The definition is as follows:
[0060] r u =Uncon_num / Pixel_num (3)
[0061] In formula (3), Uncon_num is the number of unconnected pixels; Pixel_num is the total number of pixels in each frame; when r u The MCTF process terminates when the given threshold T = 0.5 is exceeded.
[0062] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0063] 1) In the motion compensation temporal filtering stage, this invention fully considers the impact of wavelet filter type, subband coupling phenomenon, and temporal subband content-aware characteristics on the propagation of reconstruction errors. First, based on the wavelet basis coefficients selected for temporal wavelet decomposition and the temporal subband coupling phenomenon, a distortion relationship model between temporal subband frames and reconstructed video frames is constructed. Second, temporal subband weighting factors are calculated based on the mutual information values of the luminance components of adjacent subband frames, the gradient values of each subband frame, and their texture consistency. Finally, Lagrange multipliers are adaptively selected based on the Structural Similarity Index (SSIM) of the video subband frames. This algorithm can adaptively select Lagrange multipliers for each temporal subband content-aware model, which not only accurately obtains the MCTF error propagation model but also achieves better video reconstruction quality.
[0064] 2) The inventors of this application have discovered that there is a more robust relationship between the bit rate R and the Lagrange multiplier λ in a three-dimensional wavelet video encoding and decoding system. When establishing the R-λ model, the activity factor of the video sub-band frame is first obtained by calculating the color histogram and directional gradient histogram of the time-domain sub-band. Then, an R-λ model based on the video content perception characteristics is established. Finally, the coding parameters are adaptively adjusted based on the proposed R-λ model to achieve the goal of controlling the bit rate.
[0065] In summary, compared with the bitrate control methods used in current classic 3D wavelet video encoding and decoding schemes, the bitrate control method provided by this invention, without significantly increasing complexity, can achieve precise bitrate control, significantly improve the rate-distortion performance of the encoder and decoder, better maintain the continuity of video quality, and obtain more stable reconstructed video quality. Attached Figure Description
[0066] The accompanying drawings are incorporated in and form part of this specification, and together with the description serve to explain the principles of the invention.
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0068] Figure 1 A flowchart of a three-dimensional wavelet scalable video bitrate control method based on video content awareness characteristics provided by the present invention;
[0069] Figure 2 A system composition connection diagram of the three-dimensional wavelet scalable video bitrate control method provided by the present invention;
[0070] Figure 3 The processing flowchart of the stream extractor provided by the present invention;
[0071] Figure 4 A comparison chart of rate-distortion performance of different coding schemes provided in Example 2 at different target bit rates;
[0072] Figure 5 The figure shows the comparison results of the average SSIM (MSSIM) of the test video sequences with different resolutions provided in Example 2 on different three-dimensional wavelet video codec systems;
[0073] Figure 6 A comparison chart of the accuracy of bitrate control in different three-dimensional wavelet video codecs provided in Example 2;
[0074] Figure 7 The average PSNR standard deviation of different three-dimensional wavelet test video sequences provided in Example 2. Detailed Implementation
[0075] Exemplary embodiments will now be described in detail. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples consistent with some aspects of the invention as detailed in the appended claims.
[0076] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0077] Explanation of relevant technical terms:
[0078] MCTF: Motion Compensated Temporal Filtering. Rate control technology consists of two steps: First, allocating an appropriate number of bits to each level of coding unit, typically including Group of Pictures (GOP), frame, and Basic Unit (BU) levels; second, striving to achieve the pre-allocated number of bits for each level. Rate control technology allows the encoder to select appropriate coding parameters from a set of discrete, legal coding parameters to achieve the target bit rate. 3D-EZBC: Three-dimensional Embedded Zero-Block Coding. SSIM: Structural Similarity Index. SAD: Sum of Absolute Differences. SAD distortion is a common image and video quality assessment metric, primarily used to measure the degree of difference between two image or video frames. MAD: Mean Absolute Error. MSE: Mean Square Error. LMS: Least Mean Square Error. NRMSE: Normalized Root Mean Square Error. PSNR: Peak Signal to Noise Ratio.
[0079] The rate-distortion (RD) curve is a convex function, and λ is the slope of the RD curve. Assuming that D and R in the rate-distortion function are differentiable, λ can be calculated by minimizing the Lagrange cost function, i.e., equations (1) and (2) below:
[0080]
[0081]
[0082] In rate-distortion optimization of video coding, the Lagrange multiplier λ is a key coding parameter balancing bitrate and distortion. It determines the choice of coding mode during encoding and is responsible for coordinating the video coding bitrate and distortion, playing a crucial role in the encoder. λ, bitrate, and distortion are interconnected. The bitrate control method based on the R-λ model uses the Lagrange multiplier λ as the main factor controlling the bitrate and calculates the parameter λ through adaptive bitrate allocation and parameter updates.
[0083] Example 1
[0084] See Figure 1 As shown, this embodiment provides a three-dimensional wavelet scalable video bitrate control method based on video content awareness characteristics, specifically including the following steps:
[0085] Step 1: In the motion compensation time-domain filtering stage, based on the influence of wavelet filter type, subband coupling phenomenon, and time-domain subband content awareness characteristics on the propagation of reconstruction errors, an adaptive Lagrange multiplier λ selection algorithm in the wavelet domain is constructed. The specific process is as follows:
[0086] (1) Adaptive adjustment of the time-domain wavelet decomposition series:
[0087] After motion estimation, pixels are categorized into four types: connected pixels, unconnected pixels, multi-connected pixels, and intra pixels. For most video sequences, only a few pixels in the current frame have no matching pixels in the reference frame; these are called intra pixels. This invention omits the analysis of this type of pixel (intra pixels).
[0088] Let r m r represents the proportion of multi-connected pixels. u r represents the proportion of unconnected pixels. c R represents the proportion of connected pixels, where r m ≥0, r u ≥0, r c ≥0, and r m +r c +r u =1. Different types of video frames exhibit varying pixel connectivity after motion estimation, and the proportion of different pixel types differs in low-frequency frames. Furthermore, within the same GOP, the proportion of each pixel type in low-frequency frames varies across different MCTF temporal decomposition levels. As the decomposition level increases, the proportion of connected pixels in low-frequency frames decreases, while the proportion of multi-connected and non-connected pixels increases.
[0089] During the MCTF temporal decomposition process, when unconnected pixels appear in low-frequency frames, motion estimation cannot find a matching pixel, leading to periodic fluctuations in the reconstructed video quality. In particular, when bidirectional prediction matching results in too many unconnected pixels, temporal analysis cannot continue. Therefore, we use the proportion r of unconnected pixels... u The number of MCTF decomposition levels is adjusted adaptively. When r u When the threshold T = 0.5 is exceeded, we will terminate the MCTF process. The given threshold T is related to the number of time-domain decomposition levels. Where r... u The definition is shown in formula (3):
[0090] r u =Uncon_num / Pixel_num (3)
[0091] In formula (3), Uncon_num is the number of unconnected pixels; Pixel_num is the total number of pixels in each frame.
[0092] (2) Adaptive Lagrange multiplier selection
[0093] λ can be seen as a trade-off between distortion and bitrate. A larger λ means a larger proportion of R in the cost function J, implying a higher compression rate requirement during mode selection, while neglecting image quality to some extent. A smaller λ means a smaller proportion of R in the cost function J, implying that the encoder needs to sacrifice some bits to improve image quality. Therefore, we can allocate more bits to the temporal subbands containing more detailed information by reducing the value of λ. Since the distortion of a video frame is a linear combination of the distortions of all temporal subbands, and the wavelet coefficients of each temporal subband follow a Gaussian distribution, the subband distortion after wavelet transform is:
[0094]
[0095] In formula (4), d i For the distortion of the i-th temporal subband frame, r i Let be the bitrate of the i-th temporal subband frame; Let be the variance of the q-th wavelet coefficient; Q be the total number of wavelet coefficients; and k′ and η be parameters related to the time-domain subband activity factor. However, in practical applications, calculating the variance of wavelet coefficients is not easy; therefore, we can make the following approximation:
[0096]
[0097] In formula (5), β q For parameters related to wavelet transform: Let be the variance of the residual pixel values before wavelet transform; substituting formula (5) into formula (4), we get formula (6):
[0098]
[0099]
[0100] In formula (6), It can be approximately estimated by MAD, and formula (6) can be expressed as:
[0101]
[0102] In formula (7), Therefore, the Lagrange multiplier λ of the i-th frame i Obtained through the following formula:
[0103]
[0104] Based on formula (8), the Lagrange multiplier λ of the i-th frame i It can be approximated by the following formula:
[0105]
[0106] In formula (9), To suit different video sequences, an empirical parameter of 6.2 is used here.
[0107] To achieve optimal perceived quality within bitrate constraints, rate-distortion optimization coding mode selection guided by subjective distortion is necessary. This invention considers incorporating the structural similarity index SSIM distortion into the adaptive MCTF framework in 3D wavelet video coding to select the optimal coding mode, thereby maximizing subjective rate-distortion performance.
[0108] Define the SSIM-based Lagrange cost function J SSIM for:
[0109] J SSIM =D SSIM +λ SSIM ·R (10)
[0110] In formula (10), D SSIM The SSIM distortion is calculated using formula (11); R is the coding rate; λ SSIM It is a Lagrange multiplier used in rate-distortion optimization based on SSIM to balance the coding rate R and distortion D.
[0111]
[0112] In the rate-distortion optimization framework based on SSIM, SSIM distortion is used instead of SAD distortion to achieve better subjective quality, but the challenge lies in λ. i_SSIM How to determine this? Since different frames and macroblocks have different rate-distortion characteristics, if λ could be adaptively adjusted at the macroblock level based on the saliency of image content... i_SSIM This allows for a better balance between bit rate and SSIM distortion, further improving subjective rate distortion performance.
[0113] To minimize the Lagrange cost function of the SSIM-based distortion coding mode, we differentiate equation (10) and set the derivative to zero, then we have:
[0114]
[0115] Assuming that the distortion of each temporal sub-band frame is independent and identically distributed, and follows an additive white Gaussian noise distribution with a mean of 0, then SSIM can be approximated as:
[0116]
[0117] The relationship between MAD and MSE can be approximated as: MSE = (MAD) 2 Therefore, the SSIM distortion d of the i-th frame i_SSIM It can be approximated as:
[0118]
[0119]
[0120] Combining formulas (15) and (9), we can see that
[0121]
[0122] In formula (16), λ i_SSIM For the SSIM-based Lagrange multiplier of the i-th frame, c1 represents the variance of the original subband frames; c2 represents a positive constant.
[0123] In summary, the rate-distortion optimization objective function based on the adaptive Lagrange multiplier λ selection algorithm is shown in Equation (17):
[0124]
[0125] In formula (17), Weighting distortion for high-frequency subbands; Weighting distortion for low-frequency subbands; λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; For high-frequency subband SSIM distortion; This is for low-frequency subband SSIM distortion.
[0126] Step 2: Establish an R-λ model based on video content-aware characteristics. The specific process is as follows:
[0127] Define the adaptive R-λ model for the i-th sub-band frame as follows:
[0128]
[0129] In formula (18), λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; c2 is the variance of the original subband frame; c2 is the positive constant; MSE is the mean square error of the original subband frame and the reconstructed subband frame; R TThat is the target bitrate.
[0130] Step 3: Based on the R-λ model, bitrate control of the 3D wavelet scalable video is achieved through adaptive bitrate control technology. The specific process is as follows:
[0131] Step 3.1, GOP-level bit allocation:
[0132] To enable better video transmission over the network, the GOP level should implement variable bitrate control, i.e., VBR control. The number of bits allocated to each GOP level can be adaptively adjusted according to the network status. In the bitrate control algorithm proposed in this invention, the target bits of the GOP level are calculated using the following formula (19). GOP .
[0133]
[0134] In formula (19), Bit avg_frame The average number of bits per video frame is calculated using the following formula (20); N coded N is the number of frames that have been encoded. GOP N is the number of video frames contained within a GOP; SW is the sliding window size, SW≥N GOP Bit coded The total number of bits consumed for all encoded frames.
[0135]
[0136] In formula (20), R tar The target bitrate is denoted as 'target bitrate', and the frame rate is denoted as 'frame rate'.
[0137] Step 3.2, Frame-level bit allocation: Frame-level target bit frame The calculation is performed according to formula (21), that is, the remaining bits in the GOP are allocated to the remaining frames according to the weights of different video frames:
[0138] Bit frame = (Bit GOP -Coded GOP )×ω frame (twenty one)
[0139] In formula (21), Bit GOP The target number of bits at the GOP level; Coded GOP ω represents the number of bits consumed to encode the current GOP. frame The weight coefficients for the current video frame are obtained from formula (22):
[0140]
[0141] In formula (22), E frame Energy of the current frame; E GOP The current GOP energy is represented by W; the video frame width is represented by H; the video frame height is represented by I(x,y); and the pixel value of the pixel at (x,y) is represented by I(x,y).
[0142] Step 3.3, Basic Unit-Level Bit Allocation:
[0143] Target bit allocation at the basic unit level involves proportionally distributing the remaining bits of the current video frame into the remaining basic units of the current frame. Here, a basic unit refers to a different temporal sub-band after MCTF temporal decomposition. Basic unit level target bit sub It is calculated according to formula (23).
[0144] Bit sub = (Bit frame -Bit header -Coded frame )×ω sub (twenty three)
[0145] In formula (23), Bit frame The target number of bits for the current frame; Bit header The number of bits consumed for encoding header information; Coded frame ω represents the number of bits already consumed in the current frame. sub These are the weighting coefficients for the time-domain subband.
[0146] Step 3.4, Rate Control: Determine appropriate coding parameters based on the established R-λ model, simplifying the R-λ model shown in formula (18) as follows:
[0147] λ=α·2 βR (twenty four)
[0148] In formula (24), α and β are model parameters related to the characteristics of the video sequence, which can be obtained through adaptive updating, i.e.:
[0149] α new =α old +2δ·(lnλ real -lnλ theory )×α old (25)
[0150] β new =β old +2δ·(lnλ real -lnλ theory )·Rln2 (26)
[0151] In formulas (25) and (26), R is the actual number of encoded bits; αold The value of α before the update; α new The updated α value; β old The value of β before the update; β new The updated β value; λ real The actual λ value used during encoding; λ theory λ is the theoretical value of λ obtained from the R-λ model. δ is the perturbation parameter, and δ∈(0,1). It is worth noting that δ in formula (25) and formula (26) can take different values.
[0152] To verify the rationality and accuracy of the adaptive update parameters provided by this invention, the following verifications were conducted:
[0153] Taking the logarithm of formula (24) yields the following:
[0154] lnλ=lnα+ln2 βR (27)
[0155] Therefore, we have:
[0156]
[0157] The actual λ used real The value and the theoretically obtained λ theory The squared error between the values is:
[0158] e 2 =(lnλ) real -lnλ theory ) 2 (29)
[0159]
[0160] Using the LMS method for one iteration, we can obtain the following from formula (30):
[0161]
[0162] From formula (25), we can obtain:
[0163] lnα new =lnα old +2δ·(lnλ real -lnλ theory (33)
[0164] therefore,
[0165]
[0166] After Taylor expansion of formula (34) and ignoring higher-order terms, we get:
[0167]
[0168] Similarly, we can conclude that:
[0169]
[0170] As can be seen from formulas (35) and (36), they are consistent with formulas (25) and (26) respectively, so the adaptive update model parameters obtained from them are reasonable.
[0171] Example 2
[0172] Based on Example 1, in order to evaluate the performance of the proposed three-dimensional wavelet scalable video encoding and decoding framework and the bitrate control method based on video content awareness characteristics, 14 standard test videos with YUV format, 8-bit depth and 4:2:0 sampling rate were selected for the experiment, including five types of resolution: CIF (352×288), 4CIF (704×576), 720p (1280×720), 1080p (1920×1080) and 2K (2560×1600). The target bitrates extracted for test videos of different resolutions are also different (CIF and 4CIF: 256,384,512,640,768,896,1024kbps; 720p: 384,512,640,768,896,1024,1536kbps; 1080p and 2K: 2048,3072,4096,5120,6144,7168,10240kbps).
[0173] It should be noted that the algorithm implementation and model verification mentioned in this embodiment were performed on a desktop computer with an Intel quad-core (i5-2400@3.10GHz CPU) and 8GB of RAM, using the Matlab R2012b and VC++ 6.0 programming environment. During the experiment, 16 frames were selected for the Group of Pictures (GOP) (GOP=16), MCTF temporal decomposition was performed at level 4, motion estimation used a full search with 1 / 4 pixel precision, the horizontal and vertical search ranges were [-16, 15], the block size ranged from 4×4 to 128×128, and other encoding parameters remained consistent with the default configuration file. Classic 3D wavelet scalable video codec systems MC-EZBC, RPI-MC-EZBC, RWTH-MC-EZBC, and ENH-MC-EZBC can achieve good coding performance in the field of 3D wavelet video codecs. This embodiment will use the bitrate control scheme of the above codec systems as a comparative experiment, and the performance evaluation indicators are as follows:
[0174] (1) PSNR index
[0175] Since the human eye is most sensitive to the luminance component, we only considered the luminance component of the video sequence in the experiment. This paper uses the Dolby number (decibel, dB) of the peak signal-to-noise ratio (PSNR) as an objective indicator for evaluating video quality. As a commonly used objective video quality evaluation indicator, the peak signal-to-noise ratio (PSNR) is defined as follows: given an original image f(x,y) of size W×H and a distorted image g(x,y), the PSNR of image f(x,y) is calculated by formula (37):
[0176]
[0177] In formula (37), MSE is the mean squared error; f max Let f(x,y) be the maximum grayscale value of the original input image. For example, in a commonly used 8-bit grayscale image, f... max The value is 255.
[0178] (2) SSIM index
[0179] PSNR has long been favored due to its clear physical meaning and high computational efficiency. For the same video content, the higher the PSNR value, the better the video quality. However, the evaluation results of PSNR are often inconsistent with human subjective evaluation results. Studies have shown that calculating only pixel-level differences cannot accurately measure the visual quality of an image. The structural similarity index SSIM estimates the quality of a distorted image by comparing the similarity of brightness, contrast, and structural information between the original image and the distorted image. The calculation formula is shown in formula (38) below.
[0180]
[0181] In formula (38), c1 and c2 are positive constants, used to ensure that the entire fraction can obtain a stable calculation result when the denominator approaches zero; μ f The mean of the original image; μ g The mean of the distorted image; The variance of the original image; σ represents the variance of the distorted image. f,g SSIM represents the covariance between the original image and the distorted image. The SSIM value ranges from [0,1]. A larger SSIM value indicates less image distortion and better image quality. In most practical applications, using PSNR and SSIM to measure objective image quality is feasible and convenient. Therefore, this embodiment uses PSNR and SSIM as objective evaluation indicators for the quality of the decoded video.
[0182] In the experiment, we used MSSIM (Mean SSIM) as an evaluation index to measure the subjective quality of the reconstructed video, which can be calculated by the following formula (39).
[0183]
[0184] In formula (39), K is the total number of video frames; SSIM i The SSIM value of the i-th frame is obtained by formula (38).
[0185] (3)PSNR STD index
[0186] In encoders containing MCTF structures, there are periodic fluctuations in reconstructed video quality (PSNR). Using average PSNR alone as a video quality evaluation metric is incomplete; video quality fluctuation should also be considered as another indicator of the performance of a 3D wavelet video encoder. Therefore, we use PSNR to measure reconstructed video quality, and the standard deviation of PSNR (PSNR Standard Deviation) is also important. STD The variability in reconstructed video quality is measured using the following definition:
[0187]
[0188] In formula (40), PSNR i Let be the PSNR value of the i-th reconstructed video frame; K is the total number of video frames. A higher PSNR value indicates better rate-distortion performance; simultaneously, PSNR... STD The smaller the value, the smaller the fluctuation in the quality of the reconstructed video.
[0189] (4) NRMSE index
[0190] We use the standard root mean square error (NRMSE) to measure the accuracy of rate control, and the calculation formula is shown in formula (41) below.
[0191]
[0192] In formula (41), The actual coding bitrate of the i-th frame; Let i be the target bitrate of the i-th frame; NRMSE represents the average actual bitrate; K represents the total number of video frames. The smaller the NRMSE value, the more precise the bitrate control, and vice versa.
[0193] (5)R 2 index
[0194] To evaluate the accuracy of the algorithm model, we use R. 2R is a quantitative measure of the deviation between the estimated and theoretical values of λ. 2 The calculation is as follows:
[0195]
[0196] In formula (42), X i This represents the theoretical value of the i-th data point. This is the estimated value of the i-th data point; This is the average of all data. When... At that time, R 2 The maximum value of 1 is reached. R 2 The closer the value is to 1, the closer the estimated value is to the theoretical value, and the more accurate the model is.
[0197] The comparative analysis yielded the following experimental results and their analysis:
[0198] ①Analysis and discussion of the accuracy of the R-λ model
[0199] In our experiment, we tested the λ values at different MCTF decomposition levels of the video. i_SSIM R 2 The values are shown in Table 1. As can be seen from Table 1, R... 2 The value is almost 1. Therefore, the λ-adaptive selection model can fit different coding rates well, and the proposed rate control method and its mathematical model are reasonable.
[0200] Table 1. R values for different MCTF decomposition levels of video sequences. 2 Values Table 1.TheR 2 values of each MCTF level for test video sequences.
[0201]
[0202]
[0203] ② Rate-distortion performance analysis and discussion
[0204] To verify the rate-distortion performance of the 3D wavelet scalable video codec under the rate-distortion optimization criterion and the adoption of a video content-aware rate control method, our coding scheme (Ours) will be compared with the most representative 3D wavelet motion-compensated subband coding schemes (ENH-MC-EZBC, RWTH-MC-EZBC, RPI-MC-EZBC, and MC-EZBC). Let "Scheme 1", "Scheme 2", "Scheme 3", and "Scheme 4" represent the comparisons between our coding scheme and the ENH-MC-EZBC, RWTH-MC-EZBC, RPI-MC-EZBC, and MC-EZBC schemes, respectively. Figure 4 A comparison of rate-distortion performance of different coding schemes at different target bit rates is presented. Figure 4 As can be seen, our scheme can improve the average PSNR gain by 0.79-2.58dB compared with the other four schemes. Figure 5 The average SSIM (MSSIM) of test video sequences at different resolutions on different 3D wavelet video codec systems is compared. Figure 5 Thus, our proposed bitrate control algorithm can better reflect the perceptual quality of the reconstructed video under different target bitrates.
[0205] ③Analysis and discussion of bitrate control precision
[0206] Figure 6 A comparison of the accuracy of bitrate control in different three-dimensional wavelet video codecs is presented. Figure 6 Experimental results show that our proposed λ-domain rate control scheme achieves more accurate rate control. The effect is particularly significant for video sequences with rich texture information and scene transitions. For example, compared to ENH-MC-EZBC, our rate control scheme performs better, achieving more precise rate control. Our proposed rate control method has an average NRMSE of 0.89%–2.18%. The average NRMSE for rate control in the ENH-MC-EZBC codec is 2.93%–4.89%. Compared to the rate control schemes in RWTH-MC-EZBC, RPI-MC-EZBC, and MC-EZBC, the rate control method proposed in this embodiment also achieves superior rate control accuracy.
[0207] ④ Discussion on video quality fluctuation analysis
[0208] Figure 7 The results of comparing the quality fluctuations of reconstructed videos using different three-dimensional wavelet video codecs are presented. Figure 7Experimental results show that the proposed λ-domain bitrate control method based on video content awareness has a smaller PSNR standard deviation than other 3D wavelet video codecs, resulting in more stable reconstructed video quality. Our proposed bitrate control scheme achieves an average PSNR of... STD The value is 1.5. Compared with ENH-MC-EZBC, the standard deviation of PSNR of our proposed algorithm can be reduced by 2.08. At the same time, the reconstructed video quality stability obtained by using the bitrate control algorithm proposed in this paper is also better than that obtained by using the bitrate control methods in RWTH-MC-EZBC, RPI-MC-EZBC and MC-EZBC.
[0209] In summary, this embodiment provides a rate control method for a 3D wavelet scalable video codec system. This method first adaptively selects Lagrange multipliers for each temporal sub-band frame; then, it constructs an R-λ model based on the proposed adaptive Lagrange multiplier selection algorithm; finally, it proposes a λ-domain rate control method based on video content-aware characteristics, effectively improving rate control accuracy by adaptively adjusting coding parameters.
[0210] Furthermore, based on this invention, deep learning frameworks can be tailored to different application scenarios, autonomously learning low-to-high-level features from massive amounts of images. This demonstrates superior performance in artificial intelligence applications and has broad application value. In future research, we can fully leverage the advantages of wavelet transform and deep neural networks, employing an adaptive wavelet neural network as the encoding core. We can use deep learning frameworks to design a video encoding bitrate control algorithm with the highest encoding efficiency to improve video encoding performance and effectively reduce video data volume while improving video frame quality. In addition, we can introduce more biologically and neurologically inspired technologies, such as attention, concept formation, metacognition, and memory, which have promising application prospects.
[0211] Example 3
[0212] Based on the previous embodiments, this embodiment also provides a system for a three-dimensional wavelet scalable video bitrate control method, combined with... Figure 2As shown, this system is based on the ENH-MC-EZBC encoding and decoding system platform and consists of three parts: an encoder, a bitstream extractor, and a decoder. Its working principle is as follows: At the encoder end, the original video sequence is first analyzed by the MCTF analysis module. Based on the obtained temporal sub-bands, temporal sub-band weights are obtained, and adaptive λ selection is performed. Then, the temporal sub-bands are analyzed in the spatial domain to form a three-dimensional wavelet spatiotemporal sub-band. This sub-band is then quantized, and the 3D-EZBC encoded bitstream and motion vector lossless encoded bitstream information are jointly organized into a fully embeddable encoded bitstream. At the bitstream extractor end, an R-λ model is constructed, and bit allocation and rate control schemes are implemented to obtain an efficient and flexible scalable encoded bitstream. At the decoder end, the inverse operation of the encoder end is performed to generate the reconstructed decoded video sequence.
[0213] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention.
[0214] It should be understood that the present invention is not limited to the content already described above, and various modifications and changes can be made without departing from its scope. The scope of the present invention is limited only by the appended claims.
Claims
1. A three-dimensional wavelet scalable video bitrate control method based on video content-aware characteristics, characterized in that, Specifically, the steps include the following: Step 1: In the motion compensation time-domain filtering stage, based on the influence of wavelet filter type, subband coupling phenomenon, and time-domain subband content awareness characteristics on the propagation of reconstruction errors, an adaptive Lagrange multiplier selection algorithm for the wavelet domain is constructed, and the optimal adaptive Lagrange multiplier λ for the wavelet domain is selected. i_SSIM The adaptive Lagrange multiplier selection algorithm is used to adaptively select Lagrange multipliers for each time-domain sub-band. The specific implementation process of step 1 is as follows: Step 1.1: Based on the wavelet basis coefficients selected for time-domain wavelet decomposition and the time-domain subband coupling phenomenon, construct a distortion relationship model between the time-domain subband frame and the reconstructed video frame; Step 1.2: Calculate the temporal sub-band weighting factor based on the mutual information values of the luminance components of adjacent sub-band frames, the gradient values of each sub-band frame, and its texture consistency. Step 1.3: Adaptively select the current optimal wavelet domain using adaptive Lagrange multipliers based on the structural similarity index of video sub-frames; the method of adaptive selection based on the structural similarity index of video sub-frames in Step 1.3 yields the rate-distortion optimization objective function J. SSIM (F i )for: In formula (17), Weighting distortion for high-frequency subbands; Weighting distortion for low-frequency subbands; λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; For high-frequency subband SSIM distortion; For low-frequency subband SSIM distortion; Step 2: Establish an R-λ model based on video content perception characteristics; In Step 2, the R-λ model is: In formula (18), λ i_SSIM Let be the SSIM-based Lagrange multiplier for the i-th frame; c1 is the variance of the residual pixel values before wavelet transform; c2 is a positive constant; MSE is the mean square error of the original sub-band frame and the reconstructed sub-band frame; R T Target bitrate; η is an empirical parameter for different video sequences, with a value of 6.2; η is a parameter related to the temporal subband activity factor; ω i λ is the weighting factor for the i-th temporal sub-band frame; i_SSIM The calculation formula is as follows: In formula (9), is an empirical parameter for different video sequences, with a value of 6.2; MAD is the mean absolute error between the original subband frame and the reconstructed subband frame; η is a parameter related to the temporal subband activity factor; R T ω is the target bitrate. i is the weighting factor for the i-th temporal sub-band frame; In formula (16), MSE=(MAD) 2 ; c2 is the variance of the residual pixel values before wavelet transform; c2 is a positive constant; λ i Let Lagrange multiplier be the i-th frame; Step 3: Based on the R-λ model, the bitrate of the three-dimensional wavelet scalable video is controlled by adaptive bitrate control technology.
2. The three-dimensional wavelet scalable video bitrate control method based on video content awareness characteristics according to claim 1, characterized in that, λ i The acquisition process is as follows: In formula (6), d i For the distortion of the i-th temporal subband frame, r i β is the bitrate of the i-th temporal subband frame; q Here, represents the parameters related to the wavelet transform; Q is the total number of wavelet coefficients; k′ and η are parameters related to the time-domain subband activity factor; where, Then formula (6) simplifies to: Then the Lagrange multiplier λ of the i-th frame i The calculation formula is:
3. The three-dimensional wavelet scalable video bitrate control method based on video content awareness characteristics according to claim 1, characterized in that, Step 3: Based on the R-λ model, implement bitrate control for the 3D wavelet scalable video using adaptive bitrate control technology, specifically including: Step 3.1: GOP-level bit allocation, the target bit for each GOP level. GOP The calculation formula is: In formula (19), Bit avg_frame The average number of bits per video frame is calculated using formula (20); N coded N is the number of frames that have been encoded. GOP N is the number of video frames contained within a GOP; SW is the sliding window size, SW≥N GOP Bit coded R is the total number of bits consumed across all encoded frames; in formula (20), R tar The target bitrate is the target bitrate; framerate is the frame rate. Step 3.2, Frame-level bit allocation, frame-level target bit frame The calculation is obtained according to formula (21), that is, the remaining bits in the GOP are allocated to the remaining frames according to the weight of different video frames; Bit frame =(Bit GOP -Coded GOP )×ω frame (21) In formula (21), Bit GOP The target number of bits at the GOP level; Coded GOP ω represents the number of bits consumed to encode the current GOP. frame The weight coefficients for the current video frame are calculated using formula (22); In formula (22), E frame Energy of the current frame; E GOP The current GOP energy is represented by W; the video frame width is represented by H; the video frame height is represented by I(x,y); and the pixel value at (x,y) is represented by I(x,y). Step 3.3: Basic Unit-Level Bit Allocation. Define the basic unit as different time-domain sub-bands after MCTF time-domain decomposition, and the target bit at the basic unit level. sub The result is obtained according to formula (23): Bit sub =(Bit frame -Bit header -Coded frame )×ω sub (23) In formula (23), Bit frame The target number of bits for the current frame; Bit header The number of bits consumed for encoding header information; Coded frame ω represents the number of bits already consumed in the current frame. sub These are the weighting coefficients for the time-domain sub-band; Step 3.4: Implement bitrate control for 3D wavelet scalable video using an adaptive bitrate control model.
4. The three-dimensional wavelet scalable video bitrate control method based on video content awareness characteristics according to claim 3, characterized in that, In step 3.4, the adaptive bit rate control model is calculated using formula (24): λ=a·2 βR (24) In formula (24), R is the actual number of encoded bits, and α and β are model parameters related to the characteristics of the video sequence, which can be obtained through adaptive updating, i.e.: a new =a old +2δ·(lnλ real -lnλ theory )×a old (25) b new =b old +2δ·(lnλ real -lnλ theory )·Rln2 (26) In formula (25), α old The value of α before the update; α new The updated α value; λ real The actual λ value used during encoding; λ theory λ is the theoretical value of λ obtained from the R-λ model; δ is the perturbation parameter, and δ∈(0,1); In formula (26), R is the actual number of encoded bits; β old The value of β before the update; β new The updated β value; λ real The actual λ value used during encoding; λ theory λ is the theoretical value of λ obtained from the R-λ model; δ is the perturbation parameter, and δ∈(0,1).
5. A system based on the three-dimensional wavelet scalable video bitrate control method according to any one of claims 1 to 4, comprising an encoder, a bitstream extractor, and a decoder connected in sequence, characterized in that, At the encoder end, the original video is first analyzed by the MCTF analysis module, and then the temporal sub-band weights are obtained based on the obtained temporal sub-bands, and an adaptive λ selection module is used. The temporal sub-bands are then processed by the spatial analysis module to form a three-dimensional wavelet temporal-space sub-band, and then the bitstream information is jointly organized into a fully embeddable coded bitstream by the quantization module, the 3D-EZBC encoding module and the motion vector lossless encoding module. The bitstream extractor processes the received fully embeddable coded bitstream according to the process of "constructing an R-λ model → bit allocation → bit rate control" to obtain a scalable coded bitstream. The decoder decodes the received scalable encoded bitstream to generate a reconstructed video.
6. The system of the three-dimensional wavelet scalable video bitrate control method according to claim 5, characterized in that, Using the proportion of unconnected pixels r u The number of MCTF decomposition levels is adjusted adaptively, specifically as follows: Define r m r represents the proportion of multi-connected pixels. u r represents the proportion of unconnected pixels. c r represents the proportion of connected pixels; where r m ≥0, r u ≥0, r c ≥0, and r m +r c +r u =1; where r u The definition is as follows: r u =Uncon_num / Pixel_num (3) In formula (3), Uncon_num is the number of unconnected pixels; Pixel_num is the total number of pixels in each frame. When r u The MCTF process terminates when the given threshold T = 0.5 is exceeded.
Citation Information
Patent Citations
Code rate control method based on three-dimensional wavelet video coding
CN113259662A
Efficient rate allocation for multi-resolution coding of data
US20040228537A1