Adaptive context initialization for entropy coding in video coding
Patent Information
- Application Number
- JP2024552506
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-12
- Filing Date
- 2023-03-09
- Publication Date
- 2026-02-25
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 269,090, filed March 9, 2022, entitled "Entropy Coding Method," U.S. Provisional Patent Application No. 63 / 363,703, filed April 27, 2022, entitled "Entropy Coding Method," U.S. Provisional Patent Application No. 63 / 366,218, filed June 10, 2022, entitled "Entropy Coding Method," U.S. Provisional Patent Application No. 63 / 367,710, filed July 5, 2022, entitled "Entropy Coding Method," and U.S. Provisional Patent Application No. 63 / 368,240, filed July 12, 2022, entitled "Entropy Coding Method," the entire contents of which are incorporated herein by reference.
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to video processing and, more particularly, to context initialization for entropy coding in video coding. [Background technology]
[0003] The ubiquity of camera-equipped devices such as smartphones, tablets, and computers has made it easier than ever to capture videos and images. However, even a short video can contain a significant amount of data. Video coding techniques (including video encoding and decoding) compress video data into smaller sizes, allowing a variety of videos to be stored and transmitted. Video coding is used in a wide range of applications, including digital television broadcasting, video transmission over the Internet and mobile networks, real-time applications (e.g., video chat, video conferencing, etc.), digital versatile discs (DVDs) and Blu-ray discs, etc. It is desirable to improve the efficiency of video coding schemes in order to reduce the consumption of storage space for storing videos and / or network bandwidth for transmitting videos. Summary of the Invention
[0004] Some embodiments relate to context initialization for entropy coding in video coding. In one example, a method for decoding video from a video bitstream representing the video includes: accessing a binary string from the video bitstream; the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice preceding the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on the entropy coding model having the initial context value; reconstructing a frame of the video based at least in part on the decoded slice; displaying the reconstructed frame together with other frames of the video.
[0005] In another example, a non-transitory computer-readable medium stores program code, the program code executable by one or more processing devices to perform operations including: accessing a binary string from a video bitstream representing a video; the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice previous to the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on an entropy coding model having the initial context value; reconstructing a frame of the video based at least in part on the decoded slice; and displaying the reconstructed frame together with other frames of the video.
[0006] In yet another example, a system includes a processing device and a non-transitory computer-readable medium, the non-transitory computer-readable medium communicatively coupled to the processing device. The processing device is configured to execute program code stored on the non-transitory computer-readable medium, thereby performing operations including: accessing a binary string from a video bitstream representing a video; the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice preceding the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on an entropy coding model having the initial context value; reconstructing a frame of the video based at least in part on the decoded slice; and displaying the reconstructed frame together with other frames of the video.
[0007] In one example, a method for decoding video from a video bitstream representing the video includes: accessing a binary string from the video bitstream; the binary string representing a partition of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored for CTUs in a previous partition based on an initial context value associated with a previous partition of the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; reconstructing a frame of the video based at least in part on the decoded partition; and displaying the reconstructed frame.
[0008] In another example, a non-transitory computer readable medium stores program code, the program code executable by one or more processing devices to perform operations including: accessing a binary string from a video bitstream of a video; the binary string representing a partition of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored for CTUs in a previous partition based on an initial context value associated with a previous partition of the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on an entropy coding model with the initial context value; reconstructing a frame of the video based at least in part on the decoded partition; and displaying the reconstructed frame.
[0009] In yet another example, a system includes a processing device and a non-transitory computer readable medium, the non-transitory computer readable medium communicatively coupled to the processing device. The processing device is configured to execute program code stored on the non-transitory computer readable medium, thereby performing operations including: accessing a binary string from a video bitstream of a video; the binary string representing a partition of the video; determining an initial context value of an entropy coding model for the partition by transforming a context value stored for a CTU in a previous partition based on an initial context value associated with a previous partition of the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; reconstructing a frame of the video based at least in part on the decoded partition; and displaying the reconstructed frame.
[0010] In one example, a method for decoding video from a video bitstream representing the video includes: accessing a binary string from the video bitstream; the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored in a buffer for a CTU in the previous frame based on an initial context value associated with a previous frame of the frame, a quantization parameter for the previous frame, and a slice quantization parameter for the frame; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context values; replacing the context values stored in the buffer with context values for the CTU in the frame determined upon decoding the partition; reconstructing a frame of the video based at least in part on the decoded partition; displaying the reconstructed frame.
[0011] In another example, a non-transitory computer readable medium stores program code, the program code executable by one or more processing devices to perform operations including: accessing a binary string from a video bitstream of a video; the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model of the partition by transforming context values stored in a buffer for a CTU in the previous frame based on an initial context value associated with a previous frame of the frame, a slice quantization parameter of the previous frame, and a slice quantization parameter of the frame; decoding the partition by decoding at least a portion of the binary string based on an entropy coding model having the initial context values; replacing the context values stored in the buffer with context values of the CTUs in the frame determined upon decoding the partition; reconstructing a frame of the video based at least in part on the decoded partition; and displaying the reconstructed frame.
[0012] In yet another example, a system includes a processing device and a non-transitory computer readable medium, the non-transitory computer readable medium communicatively coupled to the processing device. The processing device is configured to execute program code stored on the non-transitory computer readable medium, thereby performing operations including: accessing a binary string from a video bitstream of a video; the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model of the partition by transforming context values stored in a buffer for CTUs in the previous frame based on an initial context value associated with a previous frame of the frame, a slice quantization parameter of the previous frame, and a quantization parameter of the frame; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model having the initial context values; replacing the context values stored in the buffer with context values of CTUs in the frame determined upon decoding the partition; reconstructing a frame of the video based at least in part on the decoded partition; and displaying the reconstructed frame.
[0013] These exemplary embodiments are not described to limit or define the present disclosure, but to provide examples to aid in understanding the present disclosure. Additional embodiments are described in detail in the specification and further explanation is provided in the specification. [Brief description of the drawings]
[0014] The features, embodiments and advantages of the present disclosure will be better understood from the following detailed description when read in conjunction with the accompanying drawings, in which: [Figure 1] FIG. 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments described herein. [Diagram 2] FIG. 2 is a block diagram illustrating an example of a video decoder configured to implement embodiments described herein. [Diagram 3] FIG. 3 is a diagram illustrating an example of a division of coding tree units of an image in a video, according to some embodiments of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of context initialization from the previous frame (CIPF), according to some embodiments of the present disclosure. [Diagram 5] FIG. 5 is a diagram illustrating another example of context initialization from previous frame (CIPF) according to some embodiments of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an example of a random access picture group structure with associated temporal layer indices according to some embodiments of the present disclosure. [Figure 7] FIG. 7 illustrates an example process for decoding video encoded via entropy coding with adaptive context initialization according to some embodiments of the disclosure. [Figure 8] FIG. 8 is a diagram illustrating an example of motion compensation and entropy coding context initialization dependency of an image coding structure under a random access common test condition in which context initialization using a previous frame is applied. [Figure 9] FIG. 9 illustrates an example of inheriting context initialization from a previous frame in coding order without considering temporal layers and quantization parameters in the example image coding structure shown in FIG. 8 in accordance with some embodiments of the present disclosure. [Figure 10] FIG. 10 is a diagram illustrating an example of inheriting context initialization from a previous frame in a lower temporal layer in the example image coding structure shown in FIG. 8 according to some embodiments of the present disclosure. [Figure 11] FIG. 11 is a diagram illustrating example values for context initialization table translations according to some embodiments of the present disclosure. [Figure 12]FIG. 12 illustrates an example process for decoding video encoded with a random access image coding structure via entropy coding with adaptive context initialization, according to some embodiments of the present disclosure. [Figure 13] FIG. 13 is a diagram illustrating an example of applying context initialization using previous frame (CIPF) to low-latency common test conditions according to some embodiments of the present disclosure. [Figure 14] FIG. 14 is a diagram showing the behavior of the CIPF buffer in the example shown in FIG. [Figure 15] FIG. 15 is a diagram showing another example of the random access (RA) test conditions. [Figure 16] FIG. 16 is a diagram showing the behavior of the CIPF buffer in the example shown in FIG. [Figure 17] FIG. 17 illustrates an example of the behavior of the proposed CIPF buffer configuration under the RA test conditions shown in FIG. 15 according to some embodiments of the present disclosure. [Figure 18] FIG. 18 illustrates an example process for decoding a CIPF encoded video with adaptive context initialization and the presented buffer management according to some embodiments of the present disclosure. [Figure 19] FIG. 19 illustrates an example of a computing system that can be used to implement some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] Various embodiments provide a context initialization for entropy coding in video coding. As mentioned above, more and more video data is being generated, stored, and transmitted. It is beneficial to increase the efficiency of video coding techniques. One way to increase the efficiency is through entropy coding, where an entropy coding algorithm is applied to quantized samples of a video to reduce the size of data representing the video samples. In context-based binary arithmetic entropy coding, a coding engine estimates a context probability indicating the likelihood that the next binary symbol has a value of 1. Such an estimation requires an initial context probability estimate. One way to determine the initial context probability estimate is to use a context value of a CTU located at the center of a previous slice. However, such initialization may not be accurate because the previous slice likely does not have enough bits encoded in a context-based coding mode, and the context value of the CTU located at the center of the previous slice does not accurately reflect the context of the slice.
[0016] Various embodiments described herein solve these problems by enabling adaptive context initialization. Adaptive context initialization allows for selecting an initial context value of an entropy coding model of a current slice from multiple options based on a setting or configuration of the frame or slice. For example, the initial context value can be set to the context value of the last CTU in the previous slice or frame, the context value of a CTU located at the center of the previous slice or frame, or a default initial context value independent of the previous slice or frame.
[0017] In one embodiment, a syntax element can be used to indicate a CTU position for obtaining an initial context value from a previous slice or frame. If the syntax element has a first value (e.g., 1), the initial context value can be set to the context value stored for the CTU at the center of the previous slice or frame. If the syntax element has a second value (e.g., 0), the initial context value can be set to the context value stored for the last CTU of the previous slice or frame. Another syntax element can be used to indicate whether to use a context value from a previous slice or frame for initialization or to use a default initial context value. In some examples, both syntax elements can be transmitted in the picture header of the frame containing the slice or the slice header of the slice.
[0018] In a further embodiment, a syntax element may be used to indicate a threshold value for determining a CTU position for obtaining an initial context value from a previous slice or frame. A quantization parameter (QP) value of the previous slice or frame may be compared with the threshold value. If the QP value is not greater than the threshold value, the initial context value may be set to the context value of a central CTU of the previous slice or frame. Otherwise, the initial context value may be set to the context value of a last CTU of the previous slice or frame.
[0019] In another embodiment, the initialization can be based on a time layer index associated with a frame in a random access (RA) group of pictures (GOP) structure. For example, two syntax elements can be used: a first syntax element indicating a first threshold for determining whether to use an initial context value from a previous slice or frame, and a second syntax element indicating a second threshold for determining a CTU position for obtaining an initial context value from a previous slice or frame. The second threshold is set to be not greater than the first threshold. If the time layer index of the current slice is greater than the first threshold, the initial context value of the slice is set to a default initial context value. If the time layer index is not greater than the first threshold, the time layer index of the slice is compared with the second threshold. If the time layer index is not greater than the second threshold, the initial context value is determined to be the context value of the central CTU of the previous slice or frame. Otherwise, the initial context value is set to be the context value of the last CTU of the previous slice or frame.
[0020] When applying CIPF to a random access image coding structure, initialization inheritance from a previous slice or frame with the same slice quantization parameter may introduce additional dependencies between frames, limiting the parallel processing capability of both encoding and decoding. To solve this problem, the context initialization inheritance can be modified to remove these additional dependencies. For example, the context initialization of the current frame can be determined to the context value of the previous frame in the coding order without considering the temporal layer and slice quantization parameter. In another example, the initial context value can be determined to the context value of the previous frame in the lower temporal layer. In a further example, the initial context value can be determined to the context value of the reference frame of the current frame based on the motion compensation and prediction structure.
[0021] Furthermore, since the slice quantization parameters of the current image and the image to be inherited may be different, the inherited initial context values may be transformed based on the previous slice quantization parameters and the current slice quantization parameters. In one example, the transformation is performed based on a default initial context value determined using the quantization parameters of the previous slice or frame and a default initial context value determined using the quantization parameters of the current slice or frame. In another example, the transformation is performed based on an initial context value of the previous slice or frame, which may be determined based on the slice or frame before the previous slice or frame using the same method described herein.
[0022] In order to use context values from a previous frame or slice in the CIPF, a buffer is used to store the context values. The current CIPF uses five buffers to store context values, and each buffer is used to store context data of a frame with a corresponding time layer index. However, frames with the same time layer index may have different slice quantization parameters. Therefore, every time a new combination of time layer index and slice quantization parameter is observed, the context values of the frame with the new combination are pushed into the buffer, and the old data in the buffer is discarded. As a result, the context values of previous frames, especially frames with lower time layer indexes, may be discarded, and the CIPF cannot be applied to frames where the maximum coding gain can be obtained by applying the CIPF. This results in a decrease in coding efficiency.
[0023] To solve this problem, the CIPF buffer can be managed to store context values from each time layer in the buffer. As a result, the CIPF process can be applied to each eligible frame by using the context values stored in the buffer with the same time layer index. After coding the current frame, the existing context values in the buffer that have been used as the initial context values and have the same time layer index as the current frame are replaced with new context values. If the slice quantization parameters of the current frame and the previous frame are different, the stored context values can be transformed based on the two slice quantization parameters before being used in the entropy coding model. Optionally, the number of buffers can be increased to store context values of different slice quantization parameters in different time layers and to be used for frames with corresponding time layer indexes and slice quantization parameters.
[0024] As described herein, some embodiments improve video coding efficiency by enabling adaptive selection of context value initialization of an entropy coding model. Based on the configuration (e.g., slice QP and temporal layer index) of a current slice or frame and / or a previous slice or frame, an initial context value can be selected from a central CTU or a last CTU of a previous slice, so that the initial context value can be selected more accurately. As a result, the accuracy of the entropy coding model is improved, and coding efficiency is improved.
[0025] Furthermore, by allowing the inheritance of the initial context from a slice or frame having a slice quantization parameter different from the current slice quantization parameter, the additional inter-image dependency introduced by the context initialization inheritance in the random access image coding structure can be eliminated, thereby improving the parallel processing capability of the encoder and the decoder. The inherited initial context value can be transformed based on the quantization parameter of the previous slice or frame and the quantization parameter of the current slice. This transformation reduces or eliminates the inaccuracy of the initial context value estimation due to the difference between the slice quantization parameter of the current slice or frame and the slice quantization parameter of the previous slice or frame. As a result, the overall coding efficiency is improved.
[0026] By improving the buffer management to store the context values of each temporal layer in a buffer, the video coding efficiency is further improved. Furthermore, by converting the context values based on the slice quantization parameters of the previous frame and the current frame, the same buffer can be used to score the context values of frames in temporal layers with different slice quantization parameters. As a result, the total number of buffers does not change, and CIPF can be performed for each eligible frame. Compared to the existing buffer management, in which data in the buffer may be lost and CIPF cannot be utilized for some frames, the proposed buffer management allows CIPF to be applied to more frames, achieving higher coding efficiency.
[0027]
[0023] Referring now to the drawings, Figure 1 is a block diagram illustrating an example of a video encoder 100 configured to implement embodiments described herein. In the example shown in Figure 1, the video encoder 100 includes a partition module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, an in-loop filter module 120, an intra prediction module 126, an inter prediction module 124, a motion estimation module 122, a decoded image buffer 130, and an entropy coding module 116.
[0028] The input to the video encoder 100 is an input video 102 that includes a sequence of pictures (also called frames or images). In a block-based video encoder, for each picture, the video encoder 100 utilizes a partition module 112 to divide the picture into blocks 104, each block including a number of pixels. The blocks may be macroblocks, coding tree units, coding units, prediction units, and / or prediction blocks. One picture may include blocks having different sizes, and the block divisions for different pictures of the video may also be different. Each block may be encoded using different predictions, such as intra prediction, inter prediction, or intra and inter hybrid prediction.
[0029] Typically, the first image of a video signal is an intra-coded image that is encoded using only intra prediction. In intra prediction mode, a block of the image is predicted using only encoded data from the same image. An intra-coded image can be decoded without information from other images. To perform intra prediction, the video encoder 100 shown in FIG. 1 can use an intra prediction module 126. The intra prediction module 126 is configured to generate an intra prediction block (prediction block 134) using reconstructed samples in a reconstructed block 136 of an adjacent block of the same image. The intra prediction is performed based on the intra prediction mode selected for the block. Then, the video encoder 100 calculates the difference between the block 104 and the intra prediction block 134. This difference is called a residual block 106.
[0030] To further remove redundancy from the block, the transform module 114 transforms the residual block 106 into a transform domain by applying a transform to the samples in the block. Examples of transforms include, but are not limited to, a discrete cosine transform (DCT) or a discrete sine transform (DST). The transformed values may be referred to as transform coefficients that represent the residual block in the transform domain. In some examples, the residual block may be directly quantized without being transformed by the transform module 114. This is called transform skip mode.
[0031] The video encoder 100 may further quantize the transform coefficients using a quantization module 115, thereby obtaining quantized coefficients. Quantization involves dividing a sample by a quantization step size followed by rounding, while inverse quantization involves multiplying the quantized value by the quantization step size. Such a quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of a video sample (transformed or untransformed), thereby using fewer bits to represent the video sample.
[0032] Quantization of coefficients / samples within a block can be done independently, and such quantization methods are used in some existing video compression standards such as H.264, HEVC, and VVC. For an N×M block, the two-dimensional (2D) coefficients of the block can be converted into a one-dimensional (1D) array in some scan order for use in coefficient quantization and coding. Quantization of coefficients within a block can utilize scan order information. For example, quantization of a given coefficient in a block can depend on the state of the previous quantized value along the scan order. To further improve coding efficiency, multiple quantizers can be utilized. Which quantizer is used to quantize the current coefficient depends on the previous information of the current coefficient in the encoding / decoding scan order. Such a quantization scheme is called dependent quantization.
[0033] The degree of quantization can be adjusted using a quantization step size. For example, for scalar quantization, different quantization step sizes can be applied to obtain finer or coarser quantization. A smaller quantization step size corresponds to a finer quantization, and a larger quantization step size corresponds to a coarser quantization. The quantization step size can be indicated by a quantization parameter (QP). The quantization parameter is provided in the encoded bitstream of the video so that a video decoder can access and decode using the quantization parameter.
[0034] The quantized samples are then coded by an entropy coding module 116 to further reduce the size of the video signal. The entropy encoding module 116 is configured to apply an entropy encoding algorithm to the quantized samples. In some examples, the quantized samples are binarized into binary bins, which are further compressed into bits by a coding algorithm. Examples of binarization methods include, but are not limited to, a combination of truncated Rice (TR) and restricted k-th order Exp-Golomb (EGk) binarization, and k-th order Exp-Golomb binarization. To improve coding efficiency, a method of history-based Rice parameter derivation is utilized, where the Rice parameters derived for a transform unit (TU) are based on variables obtained or updated from a previous TU. Examples of entropy encoding algorithms include, but are not limited to, variable length coding (VLC) schemes, context adaptive VLC (CAVLC) schemes, arithmetic coding schemes, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding techniques. The entropy coded data is added to the output encoded video 132 bitstream.
[0035] As mentioned above, the reconstructed block 136 from the neighboring blocks is used for intra prediction of the block of the image. Generating the reconstructed block 136 of a block involves calculating a reconstructed residual of the block. The reconstructed residual can be determined by applying an inverse quantization and an inverse transform to the quantized residual of the block. The inverse quantization module 118 is configured to apply an inverse quantization to the quantized samples to obtain inverse quantized coefficients. The inverse quantization module 118 applies the inverse of the quantization scheme applied by the quantization module 115 by utilizing the same quantization step size as the quantization module 115. The inverse transform module 119 is configured to apply an inverse transform (e.g., an inverse DCT or an inverse DST) of the transform applied by the transform module 114 to the inverse quantized samples. The output of the inverse transform module 119 is the reconstructed residual of the block in the pixel domain. The reconstructed residual is added to the prediction block 134 of the block to obtain the reconstructed block 136 in the pixel domain. For blocks whose transform is skipped, the inverse transform module 119 is not applied to those blocks, and the dequantized samples are the reconstructed residuals of those blocks.
[0036] Blocks in subsequent images following the first intra-predicted image can be coded using either inter prediction or intra prediction. In inter prediction, a prediction of a block in an image is a prediction from one or more previously encoded video images. To perform inter prediction, video encoder 100 utilizes inter prediction module 124. Inter prediction module 124 is configured to perform motion compensation of the block based on a motion estimate provided by motion estimation module 122.
[0037] The motion estimation module 122 compares the current block 104 of the current image with the decoded reference image 108 used for motion estimation. The decoded reference image 108 is stored in the decoded image buffer 130. The motion estimation module 122 selects a reference block from the decoded reference image 108 that best matches the current block. The motion estimation module 122 further determines an offset between the position (e.g., x, y coordinates) of the reference block and the position of the current block. The offset is called a motion vector (MV) and is provided to the inter prediction module 124 together with the selected reference block. In some cases, multiple reference blocks are identified for the current block in multiple decoded reference images 108. Thereby, multiple motion vectors are generated and provided to the inter prediction module 124 together with the corresponding reference blocks.
[0038] The inter prediction module 124 performs motion compensation using the motion vector and other inter prediction parameters to generate a prediction of the current block (i.e., the inter prediction block 134). For example, based on the motion vector, the inter prediction module 124 can locate the prediction block pointed to by the motion vector in the corresponding reference image. When there are multiple prediction blocks, these prediction blocks are combined with some weights to generate the prediction block 134 of the current block.
[0039] For an inter-predicted block, the video encoder 100 may subtract the inter-predicted block 134 from the block 104 to generate a residual block 106. The residual block 106 may be transformed, quantized, and entropy coded in a manner similar to the residual of an intra-predicted block described above. Similarly, a reconstructed block 136 of an inter-predicted block may be obtained by inverse quantizing and inverse transforming the residual, followed by combining it with the corresponding prediction block 134.
[0040] To obtain a decoded image 108 used for motion estimation, the reconstructed block 136 is processed by an in-loop filter module 120. The in-loop filter module 120 is configured to smooth pixel transitions, thereby improving video quality. The in-loop filter module 120 can be configured to implement one or more in-loop filters, such as a de-blocking filter, a sample adaptive offset (SAO) filter, an adaptive loop filter (ALF), etc.
[0041] 2 is a block diagram illustrating an example of a video decoder 200 configured to implement embodiments described herein. The video decoder 200 processes encoded video 202 in a bitstream to generate a decoded image 208. In the example illustrated in FIG. 2, the video decoder 200 includes an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, an in-loop filter module 220, an intra prediction module 226, an inter prediction module 224, and a decoded image buffer 230.
[0042] The entropy decoding module 216 is configured to perform entropy decoding of the encoded video 202. The entropy decoding module 216 decodes the quantized coefficients, coding parameters including intra-prediction parameters and inter-prediction parameters, and other information. In some examples, the entropy decoding module 216 decodes the bitstream of the encoded video 202 into a binary representation and converts the binary representation into quantization levels of the coefficients. The entropy decoded coefficient levels are then inverse quantized by the inverse quantization module 218 and subsequently inverse transformed to the pixel domain by the inverse transform module 219. The inverse quantization module 218 and the inverse transform module 219 function similarly to the inverse quantization module 118 and the inverse transform module 119, respectively, as described above with respect to FIG. 1. The inverse transformed residual blocks may be added to the corresponding prediction blocks 234 to generate the reconstructed blocks 236. For blocks for which the transform has been skipped, the inverse transform module 219 is not applied to these blocks. The dequantized samples produced by the dequantization module 118 are used to generate a reconstructed block 236 .
[0043] A prediction block 234 for a particular block is generated based on the prediction mode of the block. If the coding parameters of the block indicate that the block is to be intra predicted, a reconstructed block 236 of a reference block in the same image is fed to the intra prediction module 226, which can generate the prediction block 234 for the block. If the coding parameters of the block indicate that the block is to be inter predicted, the prediction block 234 is generated by the inter prediction module 224. The intra prediction module 226 and the inter prediction module 224 function similarly to the intra prediction module 126 and the inter prediction module 124 of FIG. 1, respectively.
[0044] 1, inter prediction involves one or more reference pictures. The video decoder 200 generates a decoded image 208 of the reference picture by applying an in-loop filter module 220 to a reconstructed block of the reference picture. The decoded image 208 is stored in a decoded picture buffer 230 for use by the inter prediction module 224 and for output.
[0045]
[0033] Referring now to Figure 3, Figure 3 is a diagram illustrating an example of a division of coding tree units of an image in a video, according to some embodiments of the present disclosure. As described above with respect to Figures 1 and 2, to encode an image of a video, the image is divided into blocks, such as CTU 302 in VVC as shown in Figure 3. For example, CTU 302 can be a block of 128x128 pixels. The CTUs are processed according to an order, such as the order as shown in Figure 3.
[0046] Entropy Coding
[0047] JPEG2025508539000002.jpg49170
number
number
[0048] Note that some blocks in a slice can be coded in skip-mode without CABAC, e.g., to reduce the number of bits used for the slice. Blocks coded using skip-mode do not contribute to the construction of the context.
[0049] For each context variable, two variables pStateIdx0 and pStateIdx1 are initialized as follows: From the 6-bit table entry initValue, two 3-bit variables slopeIdx and offsetIdx are derived as follows:
number
number
number
number
number
[0050] pps_cabac_init_present_flag equal to 1 specifies that sh_cabac_init_flag is present in the slice header referencing the PPS. pps_cabac_init_present_flag equal to 0 specifies that sh_cabac_init_flag is not present in the slice header referencing the PPS.
[0051] [Table 1]
[0052] As shown in Table 2, the syntax element ph_inter_slice_allowed_flag is transmitted in picture_header_structure().
[0053] ph_inter_slice_allowed_flag equal to 0 specifies that all coded slices of the picture have sh_slice_type equal to 2. ph_inter_slice_allowed_flag equal to 1 specifies that there may be one or more coded slices in the picture with sh_slice_type equal to 0 or 1, or there may be none or more coded slices.
[0054] [Table 2]
[0055] As shown in Table 4, the syntax elements sh_slice_type and sh_cabac_init_flag are transmitted in slice_header().
[0056] sh_slice_type specifies the coding type of the slice based on 3.
[0057] [Table 3]
[0058] sh_cabac_init_flag specifies the method for determining which initialization tables are used in the context variable initialization process. If sh_cabac_init_flag is not present, it is inferred to be equal to 0.
[0059] [Table 4]
[0060] In some examples, a previously coded slice or frame can be utilized for CABAC initialization. FIG. 4 is a diagram illustrating an example of CABAC context initialization from the previous frame (CIPF). As shown in FIG. 4, if the current slice type is B or P, after coding the CTU to a specified position, first obtain and store the probability state (i.e., context value) of each context model. Then, the stored probability state is used as the initial probability state of the corresponding context model in the next B slice or P slice coded with the same quantization parameter (QP) and the same temporal ID (Tid). The CTU position for storing the probability state is calculated using the following formula:
number
[0061] As shown in Table 5, the syntax element sps_cipf_enabled_flag in a sequence parameter set (SPS) can be used to indicate whether context initialization from the previous frame is enabled. If the value of sps_cipf_enabled_flag is equal to 1, the context initialization from the previous frame is used for each slice associated with the SPS. If sps_cipf_enabled_flag is equal to 0, the same CABAC context initialization process as specified in VVC is applied to each slice associated with the SPS.
[0062] [Table 5]
[0063] In the VVC standard, the quantization parameter QP of a slice is derived as follows: As shown in Table 6, the syntax elements pps_no_pic_partition_flag, pps_init_qp_minus26, and pps_qp_delta_info_in_ph_flag are transmitted in a picture parameter set (PPS).
[0064] [Table 6]
[0065] The syntax element ph_qp_delta is transmitted in the picture_header_structure as shown in Table 7. ph_qp_delta specifies the initial value of QpY that will be used for coding blocks in the picture until changed by the value of CuQpDeltaVal in the coding unit layer. If pps_qp_delta_info_in_ph_flag is equal to 1, the initial value of the QpY quantization parameter for all slices of the picture, SliceQpY, is derived as follows:
number
[0066] [Table 7]
[0067] The syntax sh_qp_delta is transmitted in the slice_header_structure as shown in Table 8. sh_qp_delta specifies the initial value of QpY to be used for coding blocks in the slice until changed by the value of CuQpDeltaVal in the coding unit layer. If pps_qp_delta_info_in_ph_flag is equal to 0, the initial value of the QpY quantization parameter for the slice, SliceQpY, is derived as follows:
number
[0068] [Table 8] In the VVC standard, as shown in Tables 9 and 10, the number of temporal layers (sub-layers) is defined in a video parameter set (VPS) and a sequence parameter set (SPS). The value of vps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may exist in a layer specified by the VPS. The value of vps_max_sublayers_minus1 should be in the range of 0 to 6, inclusive.
[0069] [Table 9]
[0070] The value of sps_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers that may be present in each coded layer video sequence (CLVS) that references the SPS. If sps_video_parameter_set_id is greater than 0, the value of sps_max_sublayers_minus1 should be in the range of 0 to vps_max_sublayers_minus1, inclusive. Otherwise (sps_video_parameter_set_id is equal to 0) the following applies: The value of -sps_max_sublayers_minus1 should be in the range of 0 to 6, inclusive. The value of -vps_max_sublayers_minus1 is inferred to be equal to sps_max_sublayers_minus1. The value of -NumSubLayersInLayerInOLS
[0000]
[0000] is inferred to be equal to sps_max_sublayers_minus1 + 1. The value of -vps_ols_ptl_idx
[0000] is inferred to be equal to 0, and the value of vps_ptl_max_tid[ vps_ols_ptl_idx
[0000] ], i.e., vps_ptl_max_tid
[0000] , is inferred to be equal to sps_max_sublayers_minus1.
[0071] [Table 10]
[0072] Generally, transform coefficients consume most of the bits in a video bitstream. If more bits are spent on a slice, the context table or context values will be more compatible during coding from one CTU to another. Meanwhile, the texture may be different from the first CTU to the last CTU. In this case, a good trade-off can be achieved by initializing the context of the next slice with the context value of the CTU at the center of the slice, as shown in FIG. 4. However, if fewer bits are spent on a slice, many blocks are coded as skip mode. In this case, the context table cannot be customized to the texture because there are not enough context-coded blocks in the slice. In this case, the context of the next slice can be initialized using the context value of the CTU toward the end of the slice according to the encoding order (e.g., the order shown in FIG. 3). For example, the last CTU of the slice can be used to initialize the context of the next slice, as shown in FIG. 5. That is, instead of Equation 8, the following equation is used to calculate the CTU position for storing the probability state:
number
[0073] The CTU position used for the initialization of the context of the next slice can be adaptively switched between number 8 and number 11. In one embodiment, when sps_cipf_flag=1, an additional syntax sps_cipf_center_flag can be transmitted as shown in Table 11 below.
[0074] [Table 11]
[0075] sps_cipf_enabled_flag equal to 1 specifies that for each slice, the CTU position for storing the probability state is specified by sps_cipf_center_flag. sps_cipf_enabled_flag equal to 0 specifies that the probability state of each slice is reset to the default initial value. sps_cipf_center_flag specifies the CTU position for storing the probability state. sps_cipf_center_flag equal to 1 specifies that for each slice, the CTU position for storing the probability state is calculated by the following formula: CTU location = min ((W + C) / 2 + 1, C) sps_cipf_center_flag equal to 0 specifies that for each slice, the CTU position for storing the probability state is calculated by the following formula: CTU location = C W represents the number of CTUs in a CTU row, and C represents the total number of CTUs in a slice. If sps_cipf_center_flag is not present, the value of sps_cipf_center_flag is inferred to be equal to 0.
[0076] When the bitrate of the bitstream is high or the QP of each slice is low, more blocks are coded in the context-based mode (i.e., non-skip mode), and better coding efficiency is provided by setting sps_cipf_center_flag = 1. Conversely, when the bitrate of the bitstream is low or the QP of each slice is high, fewer blocks are coded in the context-based mode or non-skip mode, and better coding efficiency is provided by setting sps_cipf_center_flag = 0. Thus, in the second embodiment, a pre-determined threshold is transmitted in the SPS and compared with the slice QP value to determine whether to use the CTU at the center of the slice or the last CTU for context initialization of the next slice.
[0077] For example, as shown in Table 12, a pre-determined threshold cipf_QP_threshold can be transmitted in the SPS, and the location CTU_location of the CTU used to initialize the slice context can be determined by comparing it with the QP of the previous slice, sliceQP, and the value of cipf_QP_threshold, as follows: if sliceQP <= cipf_QP_threshold CTU location = min ((W + C) / 2 + 1, C) else CTU location = C
[0078] [Table 12]
[0079] sps_cipf_enabled_flag equal to 1 specifies that for each slice, the CTU location for storing the probability state is specified by sps_cipf_QP_threshold. sps_cipf_enabled_flag equal to 0 specifies that the probability state of each slice is reset to the default initial value. When sps_cipf_enabled_flag is equal to 1, sps_cipf_QP_threshold specifies the QP threshold used to control how to determine the CTU position for entropy initialization. If the slice QP specified in the slice header is not greater than this threshold, CTU location = min ((W + C) / 2 + 1, C). If not, CTU location = C. W represents the number of CTUs in a CTU row, and C represents the total number of CTUs in a slice.
[0080] In another embodiment, the context initialization of random access (RA) is considered. As part of RA CTC, a group of pictures (GOP) structure of RA is defined as shown in FIG. 6. In this GOP structure, pictures are divided into different temporal layers, for example, layer 0 to layer 5 in FIG. 6. In this example, I-frames and B-frames are in temporal layer 0, one B-frame is in temporal layer 1, and two B-frames are in temporal layer 2, as follows: A relatively low QP is applied to the pictures of the lower temporal layer, and a relatively high QP is applied to the pictures of the higher temporal layer. More bits are spent on the pictures of the lower temporal layer, and improved coding efficiency can be achieved. Thus, in the pictures of the higher temporal layer, many blocks are coded in skip mode, and in this case, the image quality of the reference frame is more important for coding efficiency.
[0081] In this case, coding efficiency can be improved by using the predefined sps_cipf_temporal_layer_threshold. For example, as shown in Table 13, syntax elements cipf_enabled_temporal_layer_threshold and cipf_center_temporal_layer_threshold can be transmitted in the SPS. cipf_center_temporal_layer_threshold is not greater than cipf_enabled_temporal_layer_threshold.
[0082] [Table 13]
[0083] sps_cipf_enabled_flag equal to 1 specifies that the CABAC context initialization process for each slice associated with the SPS is specified by the following syntax elements sps_cipf_enabled_temporal_layer_threshold and sps_cipf_center_temporal_layer_threshold. sps_cipf_enabled_flag equal to 0 specifies that the CABAC context initialization process for all slices is the same and is reset to the default initial value. sps_cipf_enabled_temporal_layer_threshold specifies the maximum value of Tid for which CABAC context initialization from the previous frame is applied. If the value of Tid for the current slice is greater than sps_cipf_enabled_temporal_layer_threshold, the CABAC context initialization process specified by VVC is applied. The value of sps_cipf_enabled_temporal_layer_threhsold should be in the range of 0 to sps_max_sublayers_minus1+1 inclusive. sps_cipf_center_temporal_layer_threshold specifies the maximum Tid value for which the CABAC context initialization specified by Figure 4 is applied. If the value of Tid of the current slice is greater than sps_cipf_center_temporal_layer_threshold, the CABAC context initialization specified by Figure 4 is applied, that is, as follows. if Tid <= sps_cipf_center_temporal_layer_threshold CTU location = min ((W + C) / 2 + 1, C) else CTU location = C Tid represents the time layer index, W represents the number of CTUs in a CTU row, and C represents the total number of CTUs in a slice. The value of sps_cipf_center_temporal_layer_threhsold should be in the range of 0 to sps_cipf_enabled_temporal_layerthreshold, inclusive.
[0084] Another advantage of using the syntax element sps_cipf_enabled_temporal_layer_threshold is that it reduces the context values that need to be stored. For example, in Figure 6, if the value of sps_cipf_enabled_temporal_layer_threshold is 5, then the CABAC context initialization values need to be stored for Tid2, Tid3, Tid4, and Tid5. However, if the value of sps_cipf_enabled_temporal_layer_threshold is 3, then the CABAC context initialization tables need to be stored only for Tid2 and Tid3. This is useful when the encoder or decoder has limited storage.
[0085] In another embodiment, cipf_enabled_flag is transmitted in picture_header or slice_header. If cipf_enabled_flag is transmitted in picture_header or slice_header, cipf_center_flag is also transmitted in picture_header or slice_header. The proposed syntax of SPS, PPS, picture_header and slice_header are shown in Table 14, Table 15, Table 16, and Table 17, respectively.
[0086] [Table 14]
[0087] [Table 15]
[0088] [Table 16]
[0089] [Table 17]
[0090] sps_cipf_enabled_flag equal to 1 specifies that the CABAC context initialization process for each slice associated with the SPS is specified by the syntax elements ph_cipf_enabled_flag and ph_cipf_center_flag in the picture_header_structure() or the syntax elements sh_cipf_enabled_flag and sh_cipf_center_flag in the slice_header(). sps_cipf_enabled_flag equal to 0 specifies that the CABAC context initialization process for each slice associated with the SPS is the same and is reset to the default initial value. pps_cipf_info_in_ph_flag equal to 1 specifies that ph_cipf_enabled_flag and ph_cipf_center_flag are transmitted in the picture_header_structure() syntax. pps_cipf_info_in_ph_flag equal to 0 specifies that ph_cipf_enabled_flag and ph_cipf_center_flag are not transmitted in the picture_header_structure() syntax and sh_cipf_enabled_flag and sh_cipf_center_flag are transmitted in the slice_header() syntax. ph_cipf_enabled_flag equal to 1 specifies that the CABAC context initialization from the previous frame applies to all slices in the associated image. ph_cipf_enabled_flag equal to 0 specifies that the CABAC context initialization from the previous frame does not apply to any slices in the associated image and the CABAC context initialization specified by VVC applies to all slices in the associated image. ph_cipf_center_flag equal to 1 specifies that for all slices in the associated image, the CTU positions for CABAC context initialization are obtained from the previous frame as follows: CTU location = min ((W + C) / 2 + 1, C) ph_cipf_enabled_flag equal to 0 specifies that for all slices in the associated image, the CTU positions for CABAC context initialization are obtained from the previous frame as follows: CTU location = C W represents the number of CTUs in a CTU row, and C represents the total number of CTUs in a slice. sh_cipf_enabled_flag equal to 1 specifies that the CABAC context initialization from the previous frame is applied to the associated slice. sh_cipf_enabled_flag equal to 0 specifies that the CABAC context initialization from the previous frame is not applied to the associated slice and the CABAC context initialization is reset to the default initial value. sh_cipf_center_flag equal to 1 specifies that for the associated slice, the CTU position for CABAC context initialization is obtained from the previous frame as follows: CTU location = min ((W + C) / 2 + 1, C) sh_cipf_enabled_flag equal to 0 specifies that for the associated slice, the CTU position for CABAC context initialization is obtained from the previous frame as follows: CTU location = C W represents the number of CTUs in a CTU row, and C represents the total number of CTUs in a slice.
[0091] As mentioned above, by adaptively switching the CTU position for CABAC context initialization from the previous slice between the center and bottom right, the context initialization can be more accurate for the slice, reducing entropy code bits and improving coding efficiency.
[0092] 7 illustrates an example process 700 for decoding video encoded via entropy coding with adaptive context initialization, according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video decoder 200) perform the operations illustrated in FIG. 7 by executing appropriate program code (e.g., program code implementing the entropy decoding module 216). For illustrative purposes, the process 700 is described with reference to several examples depicted in the figures. However, other implementations are possible.
[0093] At block 702, the process 700 includes accessing a video bitstream representing a video signal. The video bitstream is encoded by a video encoder using entropy coding with adaptive context initialization as presented herein. At block 704, which includes blocks 706-712, the process 700 includes reconstructing each frame of a video from the video bitstream. At block 706, the process 700 includes accessing a binary bit string from the video bitstream representing a slice of the frame. At block 708, the process 700 includes determining an initial context value (e.g., p(1) in equation (1)) for the entropy coding model of the slice. For a slice, the initial context value may be adaptively determined to one of the following three options: a context value stored for a CTU located near the center of the previous slice (e.g., a CTU shown in equation (8)), a context value stored for a CTU toward the end of the previous slice (e.g., the last CTU shown in equation (11), and a default initial context value specified in the VVC standard as shown in equation (5). In one example, the order of the CTUs in the previous slice is determined by the scan order, as described above with respect to FIG.
[0094] In one embodiment, a syntax element, such as the syntax element sps_cipf_center_flag described above, may be used to indicate a CTU position for obtaining an initial context value from a previous slice. If the value of the syntax element sps_cipf_center_flag is 1, the initial context value may be set to a context value stored for a center CTU of the previous slice. If the value of the syntax element sps_cipf_center_flag is 0, the initial context value may be set to a context value stored for a last CTU of the previous slice. Another syntax element, such as sps_cipf_enabled_flag, may be used to indicate whether to use a context value from a previous slice for initialization or to use a default initial context value. In some examples, the syntax elements sps_cipf_center_flag and sps_cipf_enabled_flag may be transmitted in a picture header (PH) of a frame including the slice or in a slice header (SH) of the slice. Thus, determination of the initial context value can be performed by extracting the syntax elements sps_cipf_center_flag and sps_cipf_enabled_flag from the bitstream and selecting an appropriate initial context value based on the values of the syntax elements.
[0095] In a further embodiment, a syntax element (e.g., sps_cipf_QP_threshold as mentioned above) may be used to indicate a threshold for determining a CTU position for obtaining an initial context value from a previous slice. A quantization parameter (QP) value of the previous slice may be compared with the threshold sps_cipf_QP_threshold. If the QP value is less than or equal to the threshold, the initial context value may be set to the context value of the central CTU of the previous slice. Otherwise, the initial context value may be set to the context value of the last CTU of the previous slice.
[0096] In another embodiment, the initialization can be based on a temporal layer index associated with a frame in a random access (RA) group of pictures (GOP) structure. For example, two syntax elements can be used, including a syntax element indicating a threshold for determining whether to use an initial context value from a previous slice (e.g., sps_cipf_enabled_temporal_layer_threshold, as described above), and a syntax element indicating a second threshold for determining a CTU position for obtaining an initial context value from a previous slice (e.g., sps_cipf_center_temporal_layer_threshold, as described above). sps_cipf_center_temporal_layer_threshold is set to be not greater than sps_cipf_enabled_temporal_layer_threshold. If the temporal layer index Tid of the current slice is greater than sps_cipf_enabled_temporal_layer_threshold, the initial context value of the slice is set to the default initial context value. If the temporal layer index Tid is not greater than sps_cipf_enabled_temporal_layer_threshold, the temporal layer index Tid of the slice is compared with sps_cipf_center_temporal_layer_threshold. If the temporal layer index Tid is not greater than sps_cipf_center_temporal_layer_threshold, the initial context value is determined to be the context value of the center CTU of the previous slice. Otherwise, the initial context value is set to the context value of the last CTU of the previous slice.
[0097] At block 710, the process 700 includes decoding the slice by decoding the entropy coded portion of the binary string using the entropy coding model with the determined initial context values. The entropy decoded values may represent the quantized and transformed residuals of the slice. At block 712, the process 700 includes reconstructing a frame based on the decoded slice. The reconstruction includes dequantizing and inverse transforming the entropy decoded values to reconstruct pixel samples of the slice, as described above with respect to FIG. 2. The operations of blocks 706-712 may be performed on other slices of the frame to reconstruct the frame. At block 714, the reconstructed frame may be output for display.
[0098] It should be understood that the above examples are illustrative and should not be construed as limiting. Different implementations can be adopted for adaptive context initialization. For example, instead of using the central CTU shown in equation (8), any CTU located in the central CTU row of the slice (e.g., central 1-5 CTU rows) can be used as the first of three options. Similarly, instead of using the last CTU as shown in equation (11), any CTU located in the last few CTU rows (e.g., last 1-3 CTU rows) can be used as the second option, as long as the CTU position in the first option precedes the CTU position in the second option. Furthermore, although some examples focus on applying CIPF to slices, the same method can be applied to frames using context values stored for a CTU in a previous frame or a CTU in the last slice in a previous frame (e.g., a central CTU or an end CTU).
[0099] FIG. 8 is a diagram showing an example of motion compensation and entropy coding context initialization dependency of the image coding structure of the random access (RA) common test condition (CTC) with CIPF. In FIG. 8, each box represents one frame. The letter in the box indicates the picture type of the frame, and the number indicates the picture order count (POC) of the frame in the display order. The number under the box indicates the position of the frame in the coding order. The right side of the figure shows the temporal layer index Tid of each temporal layer similar to that shown in FIG. 6. The left side of the figure shows the delta QP of each temporal layer, where the delta QP is the difference between the QP of the layer and the base QP. The dotted lines between the boxes indicate prediction dependencies, and the solid lines indicate CIPF dependencies. As can be seen from FIG. 8, the context initialization inheritance introduces additional dependencies between images, which will limit the parallel processing capabilities of both encoding and decoding. Several embodiments are presented herein to solve this problem.
[0100] In one example, as shown in the example of FIG. 9, the context initialization value is inherited from a previous picture in the coding order without considering the temporal layer and QP. In another example, as shown in the example of FIG. 10, the context initialization value is inherited from a previous picture of a lower temporal layer. In a further example, the context initialization table inheritance follows a motion compensation and prediction structure, using a reference frame for motion compensation as a "previous" frame to inherit the state of the context variables and initialize the context variables of the current frame. The context initialization value inheritance in this example can be illustrated by the motion compensation and prediction path shown as "prediction dependency" in the dotted lines in FIG. 9 and FIG. 10. Also, when the motion prediction and motion compensation involve multiple reference frames, the context values can be inherited from multiple frames. In this case, the initialization of the context values can be a combination of these inherited values, such as an average or a weighted average. In some examples, coding standards such as VVC, ECM, AVC, and HEVC support multiple reference frames, and even within a single slice, the reference index can vary from block to block. Furthermore, coding standards such as VVC, ECM, AVC, and HEVC support bidirectional prediction, which involves list 0 prediction and list 1 prediction, which are typical forward prediction and backward prediction, respectively. In such a scenario, in list 0 prediction, a reference frame with an index equal to 0 can be used for CABAC inheritance. In the following, "slice" can be used to refer to a slice or a frame where the slice is the entire frame. The "previous" slice that the context initialization table inherits for the current slice can also be called a "reference slice".
[0101] In the above examples, the context initialization value is inherited from a frame with a different QP value. If the context initialization table of a different QP value is directly inherited, the coding efficiency may be reduced. To avoid this loss, the conversion of the context initialization table can be realized based on the previous QP and the current QP.
[0102] In one embodiment, assume that the QPs of the reference slice and the current slice are QpY_prev and QpY_curr, respectively, and the m and n specified in the numbers 4 of the reference slice and the current slice are m_prev and n_prev, and m_curr and n_curr, respectively. Equation 5 can be rewritten as follows:
number
number
[0103] In this embodiment, prevCtxState(Qp_prev) and prevCtxState(Qp_curr) are not calculated from initValue defined in the VVC standard as shown in Equation 12 and Equation 13. Conversely, prevCtxState(Qp_prev) is set to the CABAC table CtxState(Qp_prev) of the previous slice inherited by the current slice, and is a known parameter. preCtxState(Qp_curr) is the CABAC table of the current slice, and can be obtained by converting CtxState(Qp_prev) using the quantization parameters QpY_prev and QpY_curr. From Equation 12 and Equation 13, it is as follows.
number
[0104] In another embodiment, the initial context value of the current slice is determined based on the QP values of the previous slice and the current slice, and the initial context value and the inherited context value of the previous slice. Figure 11 is a diagram showing examples of various values for the context initialization table transformation of this embodiment.
[0105] As shown in Figure 11, P i QP N is the slice QP value QP N The initial context value of frame N having the vector p(1) is the vector p(1) of the vector n. i QP N is the context value of the top-left CTU of frame N. f QP N is the context value at a fixed position inherited by the first CTU of slice M. As mentioned above, the fixed position is either the central CTU or the last CTU. For the top-left CTU of frame M, P i QP M is the slice QP value QP M is the initial context value for frame M with P f QP M QP M is the context value at a fixed position of either the central CTU or the last (bottom right) CTU of frame M having f QP M is the context value inherited by the first CTU of frame X. For the top-left CTU of frame X, P i QP X is the slice QP value QP X is the initial context value for frame X with P f QP X is the slice QP value SliceQP X is the context value at a fixed position of either the center CTU or the last CTU of frame X. In other words, P f QP X is the context value inherited by the first CTU of the frame following frame X.
[0106] In one example, P i QP M is derived as follows:
number
number
number
[0107] 12 illustrates an example process 1200 for decoding video encoded with a random access image coding structure via entropy coding with adaptive context initialization, according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video decoder 200) perform the operations illustrated in FIG. 12 by executing appropriate program code (e.g., program code implementing the entropy decoding module 216). For illustrative purposes, the process 1200 is described with reference to several examples depicted in the figure. However, other embodiments are possible.
[0108] At block 1202, the process 1200 includes accessing a video bitstream representing a video signal. The video bitstream is encoded by a video encoder using entropy coding with adaptive context initialization as presented herein. At block 1204, which includes blocks 1206-1212, the process 1200 includes reconstructing each frame of the video from the video bitstream. At block 1206, the process 1200 includes accessing a binary bit string representing a partition (e.g., a slice) of the frame from the video bitstream. In some examples, a slice may be an entire frame. At block 1208, the process 1200 includes determining an initial context value (e.g., p(1) in equation (1)) of an entropy coding model for the partition. This determination may be based on a context value stored for a CTU in a previous partition, an initial context value associated with the previous partition, a slice quantization parameter for the previous partition, and a slice quantization parameter for the partition.
[0109] As mentioned above, in one example, the initial context value may be determined to be the context value of a previous frame in coding order without considering the temporal layer and the QP value, as shown in the example of Figure 9. In another example, the initial context value may be determined to be the context value of a previous frame in a lower temporal layer, as shown in the example of Figure 10. In a further example, the initial context value may be determined to be the context value of a reference frame of the current frame based on the motion compensation and prediction structure. The context value of the previous frame may be the context value stored for the central CTU or the last CTU in the previous partition, as described above.
[0110] In each of the above examples, the initial context values are inherited from a partition with a different slice QP value. A context initialization table transformation based on the previous slice QP value and the current slice QP value is used to transform the inherited initial context values to fit the current partition with the current slice QP value. In one example, a transformation is performed according to equation (15) based on the default initial context values determined using the quantization parameters of the previous partition and the default initial context values determined using the slice quantization parameters of the current partition. In another example, a transformation is performed according to equation (16) based on the initial context values of the previous partition. The initial context values of the previous partition may be determined using the same methods described herein based on the partition previous to the previous partition.
[0111] At block 1210, the process 1200 includes decoding the partition by decoding the entropy coded portion of the binary string using the entropy coding model with the determined initial context values. The entropy decoded values may represent quantized and transformed residuals of the partition. At block 1212, the process 1200 includes reconstructing the frame based on the decoded partition. The reconstruction includes dequantizing and inverse transforming the entropy decoded values to reconstruct pixel samples of the partition, as described above with respect to FIG. 2. If the frame has multiple partitions, the operations of blocks 1206-1212 may be performed on other partitions of the frame to reconstruct the frame. At block 1214, the reconstructed frame may be output for display.
[0112] CIPF Buffer Management
[0113] In the current enhanced compression model (ECM) random access common test conditions, the CIPF described in FIG. 4 is applied as shown in FIG. 8. In this test condition, the same slice QP value is assigned to slices in the same time layer. Under the ECM low delay (LD) common test conditions (CTC), the CIPF described in FIG. 4 is applied as shown in FIG. 13. In this test condition, there is only one time layer, and multiple slice QP values are assigned within the time layer.
[0114] In the CIPF described in Figure 4, the total number of context values for the CIPF stored in the buffer is limited to 5. Figure 14 illustrates the behavior of the CIPF buffer for the example shown in Figure 8. In this example, QP1 to QP5 are defined as follows:
number
[0115] After the images of POC0 and POC32, which are I-frames, are processed, the entire CIPF buffer including buffer 1 to buffer 5 is emptied, and there is no need to store CABAC context values for inheritance. After the image of POC16, which is a B-frame, is processed, a CABAC context value having QP1 with Tid1 is stored in CIPF buffer 1. After the image of POC8 is processed, a CABAC context value having QP2 with Tid2 is stored in CIPF buffer 2. After the image of POC4 is processed, a CABAC context value having QP3 with Tid3 is stored in CIPF buffer 3. After the image of POC2 is processed, a CABAC context value having QP4 with Tid4 is stored in CIPF buffer 4. After the image of POC1 is processed, a CABAC context value having QP5 with Tid5 is stored in CIPF buffer 5.
[0116] During encoding or decoding the image of POC3, the CABAC context value with QP5 of Tid5 is used because POC3 has Tid5 and QP5. After encoding the image of POC1, this CABAC context value is stored in CIPF buffer 5. Therefore, after the image of POC3 is processed, the CABAC context value with QP5 of Tid5 in CIPF buffer 5 is updated.
[0117] During encoding or decoding of the image of POC6, the CABAC context value with QP4 of Tid4 stored in the CIPF buffer 4 after encoding the image of POC2 is used. After the image of POC6 is processed, the CABAC context value with QP4 of Tid4 in the CIPF buffer 4 is updated.
[0118] In real video encoders, rate control and quantization control for perceptual optimization are usually adopted. The rate control and quantization control allow different slice QPs in the same temporal layer. Figure 15 shows another example of RA test conditions. In this example, the GOP structure is the same as the example shown in Figure 8. However, for temporal layers 4 and 5, the QP values are not constant.
[0119] The behavior of the CIPF buffer for the example of Figure 15 is shown in Figure 16. In this example, QP1 to QP 5b is defined as follows:
number
[0120] During encoding or decoding of the POC3 image, the QP 5b The CABAC initialization value with QP 5b is new to the CIPF buffer, so the QP of Tid5 5a , QP of Tid4 4a The context values with QP3 in Tid3, QP2 in Tid2 are moved to CIPF buffers 4, 3, 2, and 1, respectively. Then, after the image in POC3 is processed, the QP 5b is loaded into CIPF buffer 5. As a result, the context value with QP1 for Tid1 is deleted from the buffer.
[0121] During encoding or decoding of the POC6 image, the QP 4b The CABAC initialization value with QP 4b is new to the CIPF buffer, so the QP of Tid5 5b , QP of Tid5 5a , QP of Tid4 4a , Tid3, and QP3 are moved to CIPF buffers 4, 3, 2, and 1, respectively. Then, after the image of POC6 is processed, the QP 4b The CABAC context value with QP2 of Tid2 is added to CIPF buffer 5. As a result, the context value with QP2 of Tid2 is removed from the buffer. As can be seen from the above description, the CABAC initialization value calculated using equation 5, rather than the CIPF, is used for coding the images from POC0 to POC6.
[0122] During encoding or decoding of the image of POC24, the CABAC context table with QP2 of Tid2 does not exist in the CIPF buffer, so CIPF cannot be applied. Since the image quality of the lower temporal layer affects the image quality of the image in the higher temporal layer, usually, a small QP value is applied to the image in the lower temporal layer, and a large QP value is applied to the image in the higher temporal layer. As a result, more bits are used for the image in the lower temporal layer, and relatively fewer bits are used for the image in the higher temporal layer. To improve the overall coding efficiency, the bit saving of the image during encoding in the lower temporal layer is more important. Therefore, in this example, the fact that the CABAC context table initialization cannot be applied to Tid2 significantly reduces the coding efficiency improvement to be achieved by CIPF.
[0123] To solve this problem, the number of buffers can be increased in some cases to accommodate different combinations of temporal layers and quantization parameters. For example, the number of buffers can be set to max(5, max number of sublayers-1) instead of 5. In this way, the buffers can handle the cases in Figures 8 and 13. For example, when the value of max sublayers (temporal layers)-1 is 0, i.e., only one temporal layer is included in the bitstream, the proposed buffer configuration allows multiple CIPF buffers to be assigned to a single Tid. The assigned multiple CIPF buffers can support conditions such as the LD CTC shown in Figure 13.
[0124] In another example, the number of CIPF buffers can be set equal to the number of hierarchical layers in the motion compensation and prediction structure. Inheriting the CABAC context value for the current slice can use the context value in the buffer with the same Tid. Inheritance is allowed even if the QP values of the previous slice and the current slice are different. The discrepancy between different QP values can be resolved by transforming the CABAC context value associated with the QP values of the previous frame and the current frame.
[0125] In one example, the conversion can be done using equations (16) and (17), or more generally, as follows:
number
[0126] In another example, the conversion can be performed as follows.
number
[0127] Note that QP(N) lies in the range of 0 to 63. If the QP value SliceQP for a particular frame is outside this range, then SliceQP should first be clipped accordingly before being applied to equation (20) or equation (21). As an example, the clipping function can be defined as follows:
number
[0128] Also, for Equations 20 and 21, if QP(N) and QP(N+1) are equal to 0, then the context model P f QP(N) is the initial value P for frame N+1 i QP(N+1) is used directly as QP(N+1). If QP(N) is equal to 0 and QP(N+1) is not equal to 0, then the CABAC context initialization value calculated using Equation 5 is applied.
[0129] FIG. 17 illustrates an example of the behavior of the proposed CIPF buffer configuration under the RA test conditions shown in FIG. 15 according to some embodiments of the present disclosure. In this example, QP1 to QP 5bis defined by Equation 19. After the I images of POC0 and POC32 are processed, the entire CIPF buffer is emptied since there is no need to store CABAC context values for inheritance. After the image of POC16 is processed, the CABAC context value with QP1 of Tid1 is stored in CIPF buffer 1. After the image of POC8 is processed, the CABAC context value with QP2 of Tid2 is stored in CIPF buffer 2. After the image of POC4 is processed, the CABAC context value with QP3 of Tid3 is stored in CIPF buffer 3. After the image of POC2 is processed, the CABAC context value with QP4 of Tid4 is stored in CIPF buffer 4. 4a After the image of POC1 is processed, the CABAC context value with QP 5a The CABAC context value having the following value is stored in the CIPF buffer 5:
[0130] During encoding or decoding of an image of POC3, which has a different QP value from that of an image of POC1, first the transformed CABAC context values are calculated. This transformation is performed by converting the CABAC context values in the CIPF buffer 5, the previous slice QP value QP according to Equation 10 or 11. 5a and the current slice QP value QP 5b The converted CABAC context values are applied during encoding or decoding. After encoding or decoding the image POC3, the CABAC context values at a selected position in the image of POC3 (e.g., the CTU position selected based on equation (8) or (11)) are stored in the CIPF buffer 5.
[0131] During encoding or decoding of an image of POC6, which has a different QP value from that of an image of POC2, first the transformed CABAC context value is calculated. This calculation is carried out by taking the CABAC context value in the CIPF buffer 4, the previous slice QP value QP according to Equation 10 or Equation 11. 4a and the current slice QP value QP 4bThe converted CABAC context value is applied during encoding or decoding. After encoding or decoding the image of POC6, the CABAC context value at a selected position in the image of POC6 (e.g., the CTU position selected based on Equation 8 or Equation 11) is stored in the CIPF buffer 4. Unlike FIG. 16 in which the CIPF cannot be applied to POC0 to POC6, the CABAC initialization value calculated by Equation 20 or Equation 21 can be used for coding at least these images. Furthermore, since the CABAC context table with QP2 of Tid2 is kept in the CIPF buffer and is available for coding the image of POC24, the CIPF can also be applied to the image of POC24.
[0132] With the proposed CIPF buffer management, the CIPF buffer always stores a set of CABAC context values from each temporal layer. As a result, the CIPF process can be applied to each eligible picture by using the CABAC context values stored in the buffer with the same temporal layer index. After coding the current picture, the existing CABAC context values in the buffer that have been used as the initial CABAC context values and have the same temporal layer index as the current picture are replaced with the new CABAC context values.
[0133] With the proposed CIPF buffer management, the CIPF proposed in Equation (10) or Equation (11) can be applied. Alternatively, the existing CABAC context value inheritance method (i.e., inheritance from a previous frame with the same slice QP value in the same temporal layer) can also be applied with the exception that if the slice QP value of the current image and the QP value of the CABAC context value in the buffer with the same temporal layer index are different, the deftar initialization value calculated using Equation (5) is applied instead.
[0134] Similarly, if the slice QP value of the current picture and the slice QP value stored in the buffer of the same temporal layer are different, the buffer management shown in Fig. 16 can be improved by applying the default context initialization of number 5. In this way, the CABAC context values of the lower temporal layer are not discarded and are available to the CIPF when coding the pictures in the lower temporal layer.
[0135] FIG. 18 illustrates an example process 1800 for decoding video encoded with a random access image coding structure via entropy coding with adaptive context initialization, according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video decoder 200) perform the operations illustrated in FIG. 18 by executing appropriate program code (e.g., program code implementing the entropy decoding module 216). For illustrative purposes, the process 1800 is described with reference to several examples depicted in the figure. However, other embodiments are possible.
[0136] In block 1802, the process 1800 includes accessing a video bitstream representing a video signal. The video bitstream is encoded by a video encoder using entropy coding with adaptive context initialization as presented herein. In block 1804, which includes blocks 1806-1814, the process 1800 includes reconstructing each frame of the video from the video bitstream. In block 1806, the process 1800 includes accessing, from the video bitstream, a binary bit string representing a partition (e.g., a slice) of the frame. In some examples, a slice may be an entire frame. In block 1808, the process 1800 includes determining an initial context value (e.g., PiQP(N+1) in Equations (20) and (21)) of an entropy coding model for the partition. The decoder accesses a buffer from a set of buffers corresponding to a temporal layer (i.e., sub-layer) of the frame to determine a context value (e.g., PiQP(N+1) in Equations (20) and (21) stored for a CTU in a previous frame. f QP(N)) can be obtained. As mentioned above, the stored context value may be for the central CTU or the last CTU in the previous frame.
[0137] As mentioned above, in one embodiment, the number of buffers is set to the number of temporal layers, and each temporal layer has one buffer for storing context values. The slice quantization parameters of frames in the same temporal layer may have different values. Therefore, it is necessary to store the context values of frames with different parameter values in the same buffer. If the slice quantization parameters of a frame and the slice quantization parameters of a previous frame are different, the context values retrieved from the buffer may be transformed before being used to derive the initial context value of the current frame. For example, the transformation may be performed according to number 20 or number 21. In another embodiment, the number of buffers may be set to the larger value of 5 and the maximum number of sub-layers in the video. In this way, one buffer is used to store data of one combination of temporal layer index and slice quantization parameter value. In this embodiment, no transformation is required as long as the combination of temporal layer index and slice quantization parameter is in the CIPF buffer.
[0138] At block 1810, the process 1800 includes decoding the partition by decoding the entropy coded portion of the binary string using an entropy coding model with the determined initial context value. The entropy decoded value may represent a quantized and transformed residual of the partition. At block 1812, the process 1800 includes replacing the context value stored in the buffer with the determined context value for the CTU in the frame during decoding. As described above, the CTU may be a CTU at the center of a slice or a last CTU in the frame depending on a value of a syntax element (e.g., sps_cipf_center_flag) indicating the location of the CTU in the CIPF.
[0139] At block 1814, the process 1800 includes reconstructing the frame based on the decoded partition. The reconstruction includes dequantizing and inverse transforming the entropy decoded values to reconstruct pixel samples of the partition, as described above with respect to FIG. 2. If the frame has multiple partitions, the operations of blocks 1806-1814 may be performed on other partitions of the frame to reconstruct the frame. At block 1816, the reconstructed frame may be output for display.
[0140] Examples of Computing Systems
[0141] Any suitable computing system may be used to perform the operations described herein. For example, FIG. 19 illustrates an example of a computing device 1900 that may implement the video encoder 100 of FIG. 1 or the video decoder 200 of FIG. 2. In some embodiments, the computing device 1900 may include a processor 1912 communicatively coupled to a memory 1914 and configured to execute computer-executable program code and / or access information stored in the memory 1914. The processor 1912 may include a microprocessor, an application specific integrated circuit (ASIC), a state machine, or other processing device. The processor 1912 may include any number of processing devices, including one. Such a processor may include or be in communication with a computer-readable medium that stores instructions. The instructions, when executed by the processor 1912, cause the processor to perform the operations described herein.
[0142] The memory 1914 may include any suitable non-transitory computer-readable medium. The computer-readable medium may include any electronic, optical, magnetic, or other storage device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include magnetic disks, memory chips, read only memory (ROM), random access memory (RAM), ASICs, configured processors, optical storage devices, magnetic tape or other magnetic storage devices, or other media from which a computer processor may read instructions. The instructions may include processor-specific instructions generated by compilation and / or interpretation from code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.
[0143] The computing device 1900 may also include a bus 1916. The bus 1916 may be communicatively coupled to one or more components of the computing device 1900. The computing device 1900 may also include multiple external or internal devices, such as input devices or output devices. For example, the illustrated computing device 1900 has an input / output (I / O) interface 1918 that may receive input from one or more input devices 1920 or provide output to one or more output devices 1922. The one or more input devices 1920 and the one or more output devices 1922 may be communicatively coupled to the I / O interface 1918. The communicative coupling may be realized in any suitable manner (e.g., connection via a printed circuit board, connection via a cable, communication via wireless transmission, etc.). Non-limiting examples of input devices 1920 include a touchscreen (e.g., one or more cameras for imaging a touch area, or a pressure sensor for detecting pressure changes caused by a touch), a mouse, a keyboard, or any other device that can be used to generate input events in response to physical movements of a user of a computing device. Non-limiting examples of output devices 1922 include an LCD screen, an external monitor, speakers, or any other device that can be used to display or otherwise present output generated by a computing device.
[0144] The computing device 1900 may execute program code that configures the processor 1912 to perform one or more of the operations described above with reference to Figures 1-18. The program code may include the video encoder 100 or the video decoder 200. The program code may reside in the memory 1914, or any suitable computer readable medium, and may be executed by the processor 1912, or any other suitable processor.
[0145] The computing device 1900 may also include at least one network interface device 1924. The network interface device 1924 may include any device or group of devices suitable for establishing a wired or wireless data connection with one or more data networks 1928. Non-limiting examples of the network interface device 1924 include Ethernet network adapters, modems, etc. The computing device 1900 may transmit messages via the network interface device 1924 as electronic or optical signals.
[0146] General Considerations
[0147] Numerous details are described herein to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will understand that the claimed subject matter may be practiced without these details. In other instances, methods, apparatus, or systems that may be known to those skilled in the art have not been described in detail so as not to obscure the claimed subject matter.
[0148] Unless otherwise explained, throughout this specification, discussions utilizing terms such as "processing," "computing," "calculating," "determining," and "identifying" are understood to refer to operations or processes of a computing device (e.g., one or more computers or similar electronic computing device(s)) that manipulates or transforms data represented as physical electronic or magnetic quantities in memories, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0149] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditional on one or more inputs. Suitable computing devices include general-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a specific computing device that implements one or more embodiments of the present subject matter. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained in the software used to program or configure a computing device herein.
[0150] The method embodiments disclosed herein may be performed in operation of such a computing device. The order of the blocks in the above examples may be changed, e.g., the blocks may be reordered, combined, and / or divided into sub-blocks. Some blocks or processes may be performed in parallel.
[0151] The use of "adapted to" or "configured to" herein does not exclude devices adapted or configured to perform additional tasks or steps, but is intended as an open and inclusive phrase. Additionally, the use of "based on" is intended to be open and inclusive, such that a process, step, calculation, or other action "based on" one or more recited conditions or values may in fact be based on additional conditions or values other than the recited conditions or values. The headings, lists, and numbering contained herein are for ease of description only and are not meant to be limiting.
[0152] Although the present subject matter has been described in detail with reference to specific embodiments thereof, it is understood that those skilled in the art, after understanding the foregoing, can easily provide modifications, variations, and equivalents to such embodiments. It is therefore to be understood that the present disclosure has been presented for purposes of illustration and not limitation, and is not intended to exclude the inclusion of modifications, variations, and / or additions that would be apparent to those skilled in the art.
Claims
1. 1. A method for decoding video from a video bitstream representing said video, said method comprising: accessing a binary string from the video bitstream, the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice previous to the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on the entropy coding model with the initial context values; reconstructing the frames of the video based at least in part on the decoded slices; and displaying the reconstructed frame together with other frames of the video; Including, 1. A method for decoding video from a video bitstream representing said video, comprising:
2. The CTUs in the previous slice are encoded in an encoding order, and the first CTU is encoded before the second CTU in the previous slice.
2. The method of claim 1 .
3. The location of the first CTU is determined by CTU location = min ((W + C) / 2 + 1, C); W is the number of CTUs in the CTU row of the previous slice, C is the total number of CTUs in the previous slice, and the second CTU is the last CTU in the previous slice.
3. The method of claim 2.
4. determining the initial context value extracting from the video bitstream a syntax element indicating a CTU position for obtaining the initial context value from the previous slice; in response to determining that the syntax element has a first value, determining the initial context value to be the first context value stored for the first CTU; in response to determining that the syntax element has a second value, determining the initial context value to be the second context value stored for the second CTU; Including, 2. The method of claim 1 .
5. determining the initial context value extracting from the video bitstream a second syntax element indicating whether to use the initial context value from the previous slice; In response to determining that the second syntax element has a first value, extracting the syntax element indicating the CTU position for obtaining the initial context value from the previous slice; in response to determining that the second syntax element has a second value, the initial context value is determined to be the default initial context value; the syntax element and the second syntax element are extracted from a picture header of the frame or a slice header of the slice; 5. The method of claim 4.
6. determining the initial context value extracting from the video bitstream a syntax element indicating a threshold for determining a CTU position for obtaining the initial context value from the previous slice; comparing a quantization parameter (QP) value of the previous slice with the threshold; in response to determining that the QP value is not greater than the threshold, determining the initial context value to be the first context value stored for the first CTU; in response to determining that the QP value is greater than the threshold, determining the initial context value to be the second context value stored for the second CTU; Including, 2. The method of claim 1 .
7. determining the initial context value extracting, from the video bitstream, a first syntax element indicating a first threshold for determining whether to use the initial context value from the previous slice and a second syntax element indicating a second threshold for determining a CTU position for obtaining the initial context value from the previous slice, wherein the second threshold is not greater than the first threshold; comparing a temporal layer index of the slice with the first threshold; in response to determining that the time stratum index is greater than the first threshold, determining the initial context value to be the default initial context value; responsive to determining that the time tier index is not greater than the first threshold, comparing the time tier index of the slice with the second threshold; in response to determining that the time layer index is not greater than the second threshold, determining the initial context value to be the first context value stored for the first CTU; in response to determining that the time layer index is greater than the second threshold, determining the initial context value to be the second context value stored for the second CTU; Including, 2. The method of claim 1 .
8. A non-transitory computer readable medium having stored thereon a program code and a bitstream, the program code being executable by one or more processing devices to perform the following operations to decode the bitstream, the operations comprising: accessing a binary string from a video bitstream representing a video, the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice previous to the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on the entropy coding model with the initial context values; reconstructing the frames of the video based at least in part on the decoded slices; and displaying the reconstructed frame together with other frames of the video; Including, 1. A non-transitory computer-readable medium comprising:
9. 1. A system comprising: a processor and a non-transitory computer-readable medium; The non-transitory computer-readable medium is communicatively coupled to the processing unit, and the processing unit is configured to execute program code stored on the non-transitory computer-readable medium, thereby: accessing a binary string from a video bitstream representing a video, the binary string representing a slice of a frame of the video; determining an initial context value of an entropy coding model of the slice to one of a first context value stored for a first CTU in a slice previous to the slice, a second context value stored for a second CTU in the previous slice, and a default initial context value independent of the previous slice; decoding the slice by decoding at least a portion of the binary string based on the entropy coding model with the initial context values; reconstructing the frames of the video based at least in part on the decoded slices; and displaying the reconstructed frame together with other frames of the video; Performing operations including A system characterized by:
10. 1. A method for decoding video from a video bitstream representing said video, said method comprising: accessing a binary string from the video bitstream, the binary string representing a partition of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored for CTUs in the previous partition based on initial context values associated with a partition previous to the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; reconstructing frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Including, 1. A method for decoding video from a video bitstream representing said video, comprising:
11. A non-transitory computer readable medium having stored thereon a program code and a bitstream, the program code being executable by one or more processing devices to perform the following operations to decode the bitstream, the operations comprising: accessing a binary string from a video bitstream of a video, the binary string representing a partition of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored for CTUs in the previous partition based on initial context values associated with a partition previous to the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; reconstructing frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Including, 1. A non-transitory computer-readable medium comprising:
12. 1. A system comprising: a processor and a non-transitory computer-readable medium; The non-transitory computer-readable medium is communicatively coupled to the processing unit, and the processing unit is configured to execute program code stored on the non-transitory computer-readable medium, thereby: accessing a binary string from a video bitstream of a video, the binary string representing a partition of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored for CTUs in the previous partition based on initial context values associated with a partition previous to the partition, a slice quantization parameter of the previous partition, and a slice quantization parameter of the partition; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; reconstructing frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Performing operations including A system characterized by:
13. 1. A method for decoding video from a video bitstream representing said video, said method comprising: accessing a binary string from the video bitstream, the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored in a buffer for CTUs in the previous frame based on initial context values associated with a frame previous to the frame, a quantization parameter of the previous frame, and a slice quantization parameter of the frame; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; replacing the context values stored in the buffer with context values of CTUs in the frame determined when decoding the partition; reconstructing the frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Including, 1. A method for decoding video from a video bitstream representing said video, comprising:
14. A non-transitory computer readable medium having stored thereon a program code and a bitstream, the program code being executable by one or more processing devices to perform the following operations to decode the bitstream, the operations comprising: accessing a binary string from a video bitstream of a video, the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored in a buffer for CTUs in the previous frame based on initial context values associated with a frame previous to the frame, a slice quantization parameter of the previous frame, and a slice quantization parameter of the frame; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; replacing the context values stored in the buffer with context values of CTUs in the frame determined when decoding the partition; reconstructing the frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Including, 1. A non-transitory computer-readable medium comprising:
15. 1. A system comprising: a processor and a non-transitory computer-readable medium; The non-transitory computer-readable medium is communicatively coupled to the processing unit, and the processing unit is configured to execute program code stored on the non-transitory computer-readable medium, thereby: accessing a binary string from a video bitstream of a video, the binary string representing a partition of a frame of the video; determining initial context values of an entropy coding model for the partition by transforming context values stored in a buffer for CTUs in the previous frame based on initial context values associated with a frame previous to the frame, a slice quantization parameter of the previous frame, and a quantization parameter of the frame; decoding the partition by decoding at least a portion of the binary string based on the entropy coding model with the initial context value; replacing the context values stored in the buffer with context values of CTUs in the frame determined when decoding the partition; reconstructing the frames of the video based at least in part on the decoded partitions; and displaying the reconstructed frames; Performing operations including A system characterized by: