History-based Rice parameter derivation for wavefront parallelism in video coding
Patent Information
- Application Number
- JP2024506578
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-04
- Filing Date
- 2022-08-26
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2042-08-26
Smart Images

Figure 0007920275000015 
Figure 0007920275000016 
Figure 0007920275000017
Abstract
Description
Technical Field
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 260,600, entitled "History-Based Rice Parameter Derivations for Wavefront Parallel Processing in Video Coding", filed on August 26, 2021, U.S. Provisional Application No. 63 / 262,078, entitled "History-Based Rice Parameter Derivations for Wavefront Parallel Processing in Video Coding", filed on October 4, 2021, and U.S. Provisional Application No. 63 / 251,385, entitled "Representation of Bit Depth Range for VVC Operation Range Extension", filed on October 1, 2021. The entire contents of each of the foregoing applications are incorporated herein by reference.
[0002] The present disclosure generally relates to computer-implemented methods and systems for video processing. Specifically, the present disclosure relates to history-based Rice parameter derivation for Wavefront Parallel Processing in video coding. Background Art
[0003] Camera-equipped devices such as smartphones, tablet computers, and computers are ubiquitous, making it easier than ever to capture videos or images. However, even short videos can have a considerable amount of data. Video encoding technology (including video encoding and video decoding) can compress video data to a smaller size, thereby facilitating the storage and transmission of various types of video. Video encoding is used in a wide range of applications (e.g., digital television broadcasting, video transmission over the internet and mobile networks, real-time applications (e.g., video chat, video conferencing), DVDs, Blu-ray discs, etc.). Improving the efficiency of video encoding schemes is expected to reduce the storage capacity required to store video and / or the network bandwidth consumed to transmit video. [Overview of the Initiative]
[0004] Some embodiments relate to history-based Rice parameter derivation for wavefront parallel processing in video coding. In one example, a method for decoding a video is provided, the method comprising accessing a binary string representing a partition of the video, the partition comprising a plurality of coding tree units (CTUs), the plurality of CTUs comprising one or more CTU rows, and for each of the plurality of CTUs in the partition, before decoding the CTU, determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row of one or more CTU rows in the partition, in response to determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating the color components of the Rice parameter to an initial value, decoding the CTU, determining the pixel value of the TU in the CTU based on the coefficient value, and outputting a decoded partition of the video, the decoded partition comprising a plurality of decoded CTUs in the partition, wherein decoding the CTU involves determining the transform unit (TU) in the CTU based on the history counter This includes calculating the rice parameter of the unit and decoding the binary string corresponding to the TU in the CTU into the coefficient value of the TU based on the calculated rice parameter.
[0005] In another example, a non-temporary computer-readable medium is provided which stores program code, the program code is executed by one or more processing devices to cause the processing devices to perform a plurality of operations, the plurality of operations being to access a binary string representing a partition of video, the partition being a plurality of coding tree units (CTUs). The method includes: including a CTU (transform unit), where one or more CTU rows are composed of multiple CTUs; for each of the multiple CTUs in the partition, before decoding the CTU, determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row among the one or more CTU rows in the partition; in response to determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row in the partition, setting a history counter to an initial value for calculating the color component of the Rice parameter; decoding the CTU; determining the pixel value of the TU in the CTU based on the coefficient value; and outputting a decoded partition of the video, the decoded partition includes multiple decoded CTUs in the partition, wherein decoding the CTU includes calculating the Rice parameter of the transform unit (TU) in the CTU based on the history counter; and decoding the binary string corresponding to the TU in the CTU to the coefficient value of the TU based on the calculated Rice parameter.
[0006] In another example, a system is provided, the system comprising a processing device and a non-temporary computer-readable medium communicably connected to the processing device. The processing device is configured to perform multiple operations by executing program code stored on a non-temporary computer-readable medium, the multiple operations including accessing a binary string representing a partition of video, wherein the partition comprises multiple coding tree units (CTUs), with one or more CTU rows being composed of multiple CTUs, and for each of the multiple CTUs in the partition, before decoding the CTU, the process includes determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row among one or more CTU rows in the partition, and in response to determining whether parallel coding is enabled and that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating the color components of the Rice parameter to an initial value, decoding the CTU, determining the pixel value of the TU in the CTU based on the coefficient value, and outputting a decoded partition of video, wherein the decoded partition comprises multiple decoded CTUs in the partition, and decoding the CTU includes determining the transform unit (TU) in the CTU based on the history counter. This includes calculating the rice parameter of the unit and decoding the binary string corresponding to the TU in the CTU into the coefficient value of the TU based on the calculated rice parameter.
[0007] In another example, a method for encoding video is provided, the method comprising accessing a partition of video, the partition comprising a plurality of encoding tree units (CTUs), wherein the plurality of CTUs constitute one or more rows of CTUs, and processing the partition of video to generate a binary representation of the partition, the processing comprising, for each of the plurality of CTUs in the partition, determining, before encoding the CTU, that parallel encoding is enabled and the CTU is the first CTU of the current row of one or more rows of CTUs in the partition, setting a history counter to a value for calculating the color component of the rice parameter in response to the determination that parallel encoding is enabled and the CTU is the first CTU of the current row of CTUs in the partition, encoding the CTU, and encoding the binary representation of the partition into a bitstream of video, wherein encoding the CTU comprises, based on the history counter, calculating the rice parameter of the transformation unit (TU) in the CTU, and based on the calculated rice parameter, encoding the coefficient value of the TU into a binary representation corresponding to the TU in the CTU.
[0008] In another example, a non-temporary computer-readable medium is provided which stores program code, the program code is executed by one or more processing devices to cause the processing devices to perform a plurality of operations, the plurality of operations being to access a partition of video, the partition comprising a plurality of coding tree units (CTUs), the plurality of CTUs comprising one or more CTU rows, and processing the partition of video to generate a binary representation of the partition, the processing being such that for each of the plurality of CTUs in the partition, parallel coding is enabled before encoding the CTU, and the CTU comprises one or more CTUs in the partition. The process includes determining that the current CTU row is the first CTU in the U row, setting a history counter to an initial value for calculating the color component of the rice parameter in response to the determination that parallel encoding is enabled and that the CTU is the first CTU in the current CTU row within the partition, encoding the CTU, and encoding the binary representation of the partition into a video bitstream, wherein encoding the CTU includes calculating the rice parameter of the transform unit (TU) within the CTU based on the history counter, and encoding the coefficient value of the TU into a binary representation corresponding to the TU within the CTU based on the calculated rice parameter.
[0009] In another example, a system is provided, the system comprising a processing device and a non-temporary computer-readable medium communicably connected to the processing device. The processing device is configured to perform a number of operations by executing program code stored on the non-temporary computer-readable medium, the number of operations being access to a partition of video, the partition comprising a number of coding tree units (CTUs), the number of CTUs comprising one or more CTU rows, and processing the partition of video to generate a binary representation of the partition, the processing being for each of the CTUs in the partition, with parallel coding enabled before encoding the CTU, and the CTU being the current CTU row of one or more CTU rows in the partition The process includes determining that it is the first CTU, initializing a history counter for calculating the color component of the rice parameter in response to the determination that parallel coding is enabled and that the CTU is the first CTU in the current CTU row in the partition, coding the CTU, and coding the binary representation of the partition into a video bitstream, wherein coding the CTU includes calculating the rice parameter of the transformation unit (TU) in the CTU based on the history counter, and coding the coefficient value of the TU into the binary representation corresponding to the TU in the CTU based on the calculated rice parameter.
[0010] These exemplary embodiments are intended not to limit or restrict the disclosure, but to provide examples that aid in understanding the disclosure. Specific embodiments discuss other embodiments, and further explanation is provided in the specific embodiments. [Brief explanation of the drawing]
[0011] [Figure 1] An exemplary block diagram of a video encoder configured to implement the embodiments presented herein is shown. [Figure 2]An exemplary block diagram of a video decoder configured to implement the embodiments presented herein is shown. [Figure 3] The following are examples of encoding tree unit partitioning of pictures in a video according to some embodiments of the present disclosure. [Figure 4] An example of encoding unit partitioning of an encoding tree unit according to some embodiments of this disclosure is shown. [Figure 5] An example of a coded block in which the processing order of the elements of the coded block is predetermined is shown. [Figure 6] An example of a template pattern for calculating a locally summable variable of coefficients located near the boundary of a transformation unit is shown. [Figure 7] An example of a tile where wavefront parallel processing is effective is shown. [Figure 8] Examples of frames in which history counters are calculated according to some embodiments of this disclosure, and examples of tiles and encoded tree units contained within such frames are shown. [Figure 9] The following are examples of video partition encoding processes according to some embodiments of this disclosure. [Figure 10] The following are examples of video partition decoding processes according to some embodiments of the present disclosure. [Figure 11] Here is another example of a video partition encoding process according to some embodiments of the present disclosure. [Figure 12] Here is another example of a video partition decoding process according to some embodiments of the present disclosure. [Figure 13] An example of a computing system for realizing some embodiments of this disclosure is shown. [Modes for carrying out the invention]
[0012] By referring to the drawings and reading the following specific embodiments, you will be able to better understand the features, examples, and advantages of this disclosure.
[0013] Each embodiment provides a history-based Rice parameter derivation for wavefront parallel processing in video coding. As mentioned above, an increasing amount of video data is being generated, stored, and transmitted. It is beneficial to improve the efficiency of video coding techniques, thereby representing video using less data without compromising the visual quality of the decoded video. One way to improve coding efficiency is to compress the processed video samples into a binary bitstream using as few bits as possible, through entropy coding. Furthermore, since video typically contains a large amount of data, it is beneficial to reduce the processing time during coding (coding and decoding). To this end, parallel processing can be employed for video coding and decoding.
[0014] In entropy coding, video samples are binarized into binary bins, which can then be further compressed into bits by coding algorithms such as Context-Adaptive Binary Arithmetic Coding (CABAC). Binarization requires the calculation of binarization parameters, including Rice parameters used in combination with truncated Rice (TR), as defined in the Versatile Video Coding (VVC) specification, and the k-th order Exp-Golomb (EGk) binarization process. History-based Rice parameter derivation is used to improve coding efficiency. In this history-based rice parameter derivation, the rice parameters of the transform units (TUs) within the current coding tree unit (CTU) of a partition (e.g., picture, slice, or tile) are derived based on a history counter (denoted as StatCoeff), which is calculated based on the coefficient of the previous TU among the current CTU and previous CTUs within the partition. Subsequently, the history counter is used to derive substitution variables (denoted as HistValue) in order to derive the rice parameters. The history counter may be updated when processing TUs. In some examples, updating the history counter does not change the substitution variables of the TUs.
[0015] The dependency between previous and current CTUs within a partition for calculating the history counter can conflict with the use of parallel processing, potentially leading to unstable or inefficient video encoding by limiting or hindering the use of parallel processing. Various embodiments described herein address these issues by reducing or eliminating dependencies between several CTUs within a partition, thereby enabling parallel processing to speed up the video processing process, or by detecting and avoiding collisions before they occur. Several embodiments are presented below, with non-limiting examples.
[0016] In one embodiment, dependency conflicts with parallel processing are eliminated by removing dependencies between CTUs of different CTU rows when calculating history counters. For example, the history counter may be reinitialized for each CTU row of a partition. The history counter may be set to an initial value before calculating the Rice parameter for the first CTU in a CTU row. A subsequent history counter may be calculated based on the history counter values of previous TUs within the same CTU row. In this way, in history-based Rice parameter derivation, CTU dependencies are limited within the same CTU row, do not interfere with parallel processing in different CTU rows, and at the same time, the benefit of coding gain achieved based on history-based Rice parameter derivation can be obtained. Furthermore, the history-based Rice parameter derivation process is simplified, and computational complexity is reduced.
[0017] In another embodiment, the dependency between CTUs when calculating a history counter matches the dependency between CTUs in parallel processing. For example, parallel encoding may be performed on CTU rows of a partition, and an N-CTU delay may exist between two consecutive CTU rows. That is, processing of a current CTU row is started after N CTUs of the immediately preceding CTU row have been processed. In this case, the history counter of the current CTU row can be calculated based on samples in the first N or fewer CTUs within the immediately preceding CTU row. This may be achieved by a storage synchronization process. After processing the last TU within the first CTU of a CTU row, the history counter may be stored in a storage variable. Thereafter, before processing the first TU within the first CTU of a subsequent CTU row, the history counter may be synchronized with the stored value in the storage variable.
[0018] In some examples, alternative history-based Rice parameter derivation is used. In this alternative history-based Rice parameter derivation, when the history counter StatCoeff is updated when processing a TU, the substitution variable HistValue is updated. In order to avoid dependency conflicts with parallel encoding, similarly, the dependency between CTUs when calculating the history counter can be limited to no more than N CTUs. Similarly, a storage synchronization process can be implemented. After processing the last TU in the first CTU of a CTU row, the history counter and the substitution variable can each be stored in storage variables. Thereafter, before processing the first TU in the first CTU of a subsequent CTU row, the history counter and the substitution variable can be synchronized with the stored values in the corresponding storage variables.
[0019] In this way, the dependency between CTUs of two consecutive CTU rows when calculating the history counter is restricted to be smaller than the dependency between CTUs when performing parallel encoding (that is, consistent with the dependency between CTUs when performing parallel encoding). Therefore, the history counter calculation does not interfere with parallel processing, and at the same time, the benefit of coding gain realized based on history-based Rice parameter derivation can be obtained.
[0020] Alternatively, coexistence of parallel processing and history-based Rice parameter derivation in a bitstream is prevented. For example, a video encoder can determine whether parallel processing is enabled. If parallel processing is enabled, history-based Rice parameter derivation is disabled; otherwise, history-based Rice parameter derivation is enabled. Similarly, if the video encoder determines that history-based Rice parameter derivation is enabled, parallel processing is disabled, and vice versa.
[0021] Using the Rice parameters determined as described above, the video encoder can binarize the prediction residual data (e.g., quantized transform coefficients of residuals) into binary bins and further compress the bins into bits contained in the video bitstream using an entropy coding algorithm. On the decoder side, the decoder can decode the bitstream back into binary bins, determine the Rice parameters using any of the above methods or a combination of methods, and then determine the coefficients based on the binary bins. Furthermore, inverse quantization and inverse transform can be performed on the coefficients to reconstruct the video blocks for display.
[0022] In some embodiments, the bit depth of a video sample (for example, the bit depth used to determine the initial value of the history counter StatCoeff) can be determined based on the Sequence Parameter Set (SPS) syntax element sps_bitdepth_minus8. The value of the SPS syntax element sps_bitdepth_minus8 ranges from 0 to 8. Similarly, the size of the decoded picture buffer (DPB) for storing the decoded pictures can be determined based on the Video Parameter Set (VPS) syntax element vps_ols_dpb_bitdepth_minus8. The value of the VPS syntax element vps_ols_dpb_bitdepth_minus8 ranges from 0 to 8. Based on the determined DPB size, storage capacity can be allocated to the DPB. The determined bit depth and DPB can be used throughout the entire process of decoding the video bitstream into pictures.
[0023] As described herein, several embodiments provide improvements in video coding efficiency and computational efficiency by coordinating history-based rice parameter derivation with parallel coding. In this way, the stability of the coding process can be improved by avoiding conflicts between history-based rice parameter derivation and parallel coding. Furthermore, by limiting the dependencies between CTUs in history-based rice parameter derivation to less than or equal to the dependencies in parallel coding, coding gains can be achieved by history-based rice parameter derivation without sacrificing the computational efficiency of the coding process. This technique has the potential to become an effective coding tool in future video coding standards.
[0024] Next, referring to the drawings, Figure 1 shows an exemplary block diagram of a video encoder configured to realize the embodiments submitted herein. In the example shown in Figure 1, the video encoder 100 comprises a partition module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, an in-loop filter module 120, an intra prediction module 126, an inter prediction module 124, a motion estimation module 122, a decoded picture buffer 130, and an entropy coding module 116.
[0025] The input to the video encoder 100 is an input video 102 containing a sequence of pictures (also called frames or images). In a block-based video encoder, for each picture, the video encoder 100 employs a partition module 112 to divide the picture into blocks 104, each block containing multiple pixels. A block may be a macroblock, a coding tree unit, a coding unit, a prediction unit, and / or a prediction block. A single picture may contain blocks of different sizes, and the block partitions of different pictures in a video may be different. Each block may be coded using a different prediction method (e.g., intra-prediction, inter-prediction, or a hybrid of intra- and inter-prediction).
[0026] Typically, the first picture in a video signal is an intra-prediction picture, and this picture is encoded using only intra-prediction. In intra-prediction mode, blocks of a picture are predicted using only data from the same picture. An intra-prediction picture can be decoded in the absence of information from other pictures. To perform intra-prediction, the video encoder 100 shown in Figure 1 can employ an intra-prediction module 126. The intra-prediction module 126 is configured to generate an intra-prediction block (prediction block 134) using reconstruction samples within reconstruction blocks 136 of adjacent blocks of the same picture. Intra-prediction is performed according to the intra-prediction mode selected for the block. Subsequently, the video encoder 100 calculates the difference between block 104 and intra-prediction block 134. This difference is called the residual block 106.
[0027] To further remove redundancy from the block, the transformation module 114 transforms the residual block 106 into a transformation domain by applying a transformation to the samples within the block. Examples of transformations include, but are not limited to, the discrete cosine transform (DCT) or the discrete sine transform (DST). The transformed values are called transformation coefficients, and these coefficients represent the residual block in the transformation domain. In some examples, the residual block can be quantized directly without being transformed by the transformation module 114. This is called transformation skip mode.
[0028] The video encoder 100 can further quantize the conversion coefficients using the quantization module 115 to obtain the quantization coefficients. Quantization involves dividing the sample by the quantization step size and then rounding, while inverse quantization involves multiplying the quantized value by the quantization step size. Such a quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of video samples (converted or unconverted), thereby representing the video samples using fewer bits.
[0029] The quantization of coefficients / samples within a block can be performed independently, and this quantization method is used in several existing video compression standards (e.g., H.264 and HEVC). For N×M blocks, the 2D coefficients of the block can be converted to a 1-D array in a specific scan order for coefficient quantization and encoding. Scan order information can be used for coefficient quantization within a block. For example, the quantization of a given coefficient within a block can be determined by the state of the previous quantized value along the scan order. Multiple quantizers can be used to further improve encoding efficiency. Which quantizer to use to quantize the current coefficient depends on the information that precedes the current coefficient in the encoding / decoding scan order. This quantization method is called dependent quantization.
[0030] The degree of quantization can be adjusted using the quantization step size. For example, in scalar quantization, different quantization step sizes can be applied to achieve finer or coarser quantization. Smaller quantization step sizes correspond to finer quantization, while larger quantization step sizes correspond to coarser quantization. The quantization step size can be represented by the quantization parameter (QP). By providing the quantization parameter in the video encoding bitstream, the same quantization parameter can be applied to the video decoder for decoding.
[0031] Subsequently, the quantized samples are encoded using the entropy coding module 116 to further reduce the size of the video signal. The entropy coding module 116 is configured to apply an entropy coding algorithm to the quantized samples. In some examples, the quantized samples are binarized into binary bins, and the coding algorithm further compresses the binary bins into bits. Examples of binarization methods include, but are not limited to, truncated Rice (TR) and k-th order Exp-Golomb (EGk) binarization. To improve coding efficiency, a history-based Rice parameter derivation method is used, where the Rice parameters derived for a transform unit (TU) are based on variables obtained or updated from the previous TU. Examples of entropy coding algorithms include, but are not limited to, variable-length coding (VLC) schemes, context-adaptive VLC (CAVLC) schemes, arithmetic coding schemes, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The entropy-coded data is added to the bitstream that outputs the coded video 132.
[0032] As described above, the intra-prediction of a picture block uses a reconstructed block 136 from an adjacent block. The generation of a block's reconstructed block 136 involves calculating the block's reconstructed residual. The reconstructed residual can be determined by applying inverse quantization and inverse transform to the block's quantized residual. The inverse quantization module 118 is configured to obtain dequantization coefficients by applying inverse quantization to the quantized samples. The inverse quantization module 118 applies the inverse scheme of the quantization scheme applied to the quantization module 115 by using the same quantization step size as the quantization module 115. The inverse transform module 119 is configured to apply the inverse transform (such as inverse DCT or inverse DST) of the transform applied to the transform module 114 to the dequantized samples. The output of the inverse transform module 119 is the reconstructed residual of the block in the pixel domain. The reconstructed block 136 in the pixel domain can be obtained by adding the reconstructed residual to the block's predicted block 134. For blocks where the transform was skipped, the inverse transform module 119 is not applied to those blocks. The dequantization sample is the reconstruction residual of the block.
[0033] Interpretation or intrapretation can be used to encode blocks of subsequent pictures following a first intrapredicted picture. In interpretation, the prediction of blocks within a picture is based on one or more previously encoded video pictures. To perform interpretation, the video encoder 100 uses an interpretation module 124. The interpretation module 124 is configured to perform motion compensation for blocks based on motion estimation provided by the motion estimation module 122.
[0034] The motion estimation module 122 performs motion estimation by comparing the current block 104 of the current picture with the decoded reference picture 108. The decoded reference picture 108 is stored in the decoded picture buffer 130. The motion estimation module 122 selects the reference block from the decoded reference picture 108 that best matches the current block. The motion estimation module 122 further identifies the offset between the position of the reference block (e.g., x, y coordinates) and the position of the current block. This offset is called the motion vector (MV) and is provided to the interprediction module 124. In some cases, multiple reference blocks are identified for blocks in multiple decoded reference pictures 108. Therefore, multiple motion vectors are generated and provided to the interprediction module 124.
[0035] The interprediction module 124 performs motion compensation using motion vectors and other interprediction parameters to generate a prediction for the current block (i.e., the interprediction block 134). For example, based on motion vectors, the interprediction module 124 can identify the prediction block pointed to by the motion vectors within the corresponding reference picture. If multiple prediction blocks exist, these prediction blocks are combined with several weights to generate the prediction block 134 for the current block.
[0036] In the case of an interprediction block, the video encoder 100 can generate a residual block 106 by subtracting the interprediction block 134 from block 104. The residual block 106 can be transformed, quantized, and entropy coded in the same manner as the residual of the intraprediction block described above. Similarly, the reconstruction block 136 of the interprediction block is obtained by performing inverse quantization and inverse transform on the residual and then combining it with the corresponding prediction block 134.
[0037] To obtain a decoded picture 108 for motion estimation, the reconstruction block 136 is processed by the loop filter module 120. The loop filter module 120 is configured to smooth the pixel transformation to improve video quality. The loop filter module 120 is configured to be implemented as one or more loop filters, such as a de-blocking filter, a sample-adaptive offset (SAO) filter, or an adaptive loop filter (ALF).
[0038] Figure 2 shows an example of a video decoder 200 configured to implement the embodiments presented herein. The video decoder 200 processes encoded video 202 in a bitstream and generates a decoded picture 208. In the example shown in Figure 2, the video decoder 200 comprises an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, a loop filter module 220, an intra-prediction module 226, an inter-prediction module 224, and a decoded picture buffer 230.
[0039] The entropy decoding module 216 is configured to perform entropy decoding of the encoded video 202. The entropy decoding module 216 decodes the encoding parameters, including quantization coefficients, intra-prediction parameters and inter-prediction parameters, and other information. In some examples, the entropy decoding module 216 decodes the bitstream of the encoded video 202 into a binary representation, and then converts the binary representation to the quantization level of the coefficients. The coefficients of the entropy decoding are then dequantized by the inverse quantization module 218 and subsequently inverse-transformed back into pixel domains by the inverse transform module 219. The functions of the inverse quantization module 218 and the inverse transform module 219 are analogous to the inverse quantization module 118 and the inverse transform module 119 described above with reference to Figure 1. The inverse-transformed residual blocks can be added to the corresponding prediction blocks 234 to generate the reconstruction blocks 236. For blocks where the transformation is skipped, the inverse transform module 219 is not applied to those blocks. The dequantized samples generated by the inverse quantization module 118 are used to generate the reconstruction block 236.
[0040] The prediction block 234 for a specific block is generated based on the block's prediction mode. If the block's coding parameters instruct the block to perform intra-prediction, the reconstruction block 236 of the reference block in the same picture is supplied to the intra-prediction module 226 to generate the block's prediction block 234. If the block's coding parameters instruct the block to perform inter-prediction, the prediction block 234 is generated by the inter-prediction module 224. The functions of the intra-prediction module 226 and the inter-prediction module 224 are similar to those of the intra-prediction module 126 and inter-prediction module 124 in Figure 1, respectively.
[0041] As described in Figure 1 above, one or more reference pictures are involved in the interpretation. The video decoder 200 generates a decoded picture 208 of the reference picture by applying the loop filter module 220 to the reconstruction block of the reference picture. The decoded picture 208 is stored in the decoded picture buffer 230 for use by the interpretation module 224 and is also used for output.
[0042] Referring next to Figure 3, Figure 3 shows an example of the division of coding tree units of a picture in a video according to some embodiments of the present disclosure. As described with respect to Figures 1 and 2 above, in order to encode a picture in a video, the picture is divided into blocks such as coding tree units (CTUs) 302 in VVC, as shown in Figure 3. For example, the CTU 302 may be a block of 128 × 128 pixels. The CTUs are processed according to a specific order (e.g., the order shown in Figure 3). In some examples, as shown in Figure 4, each CTU 302 in the picture can be divided into one or more coding units (CUs) 402, and the CUs 402 can be further divided into prediction units or transformation units (TUs) for prediction and transformation. Depending on the coding scheme, the CTUs 302 can be divided into CUs 402 in different ways. For example, in VVC, the CUs 402 may be rectangular or square and can be encoded without being divided into prediction units or transformation units. Each CU402 may be the same size as its root CTU302, or it may be a subdivision of the root CTU302 that is as small as a 4x4 block. As shown in Figure 4, the division of CTU302 into CU402 in VVC may be a quadtree, binary tree, or ternary tree. In Figure 4, solid lines represent quadtree divisions, and dashed lines represent binary or ternary tree divisions.
[0043] As described in Figures 1 and 2 above, quantization is used to reduce the dynamic range of elements in a block within a video signal, representing the video signal using fewer bits. In some examples, before quantization, the elements at a particular position in a block are called coefficients. After quantization, the quantized value of the coefficients is called the quantization level or level. Quantization typically involves division by the quantization step size followed by rounding, while inverse quantization involves multiplication by the quantization step size. This quantization process is also called scalar quantization. The quantization of coefficients within a block can be performed independently, and this independent quantization method is used in some existing video compression standards (e.g., H.264, HEVC, etc.). In other examples, such as VVC, dependent quantization is employed.
[0044] In the case of an N×M block, the 2-D coefficients of the block can be converted to a 1-D array in a specific scan order for coefficient quantization and coding, and the same scan order is used for coding and decoding. Figure 5 shows an example of a coded block (e.g., a transform unit (TU)) where the coefficients of the coded block are processed in a predetermined scan order. In this example, the size of the coded block 500 is 8×8, and the processing starts from the lower right corner at position L0 to the upper left corner L0. 63 The process ends there. If block 500 is a transformation block, the predetermined order shown in Figure 5 is from the highest frequency to the lowest frequency. In some examples, the processing of a block (e.g., quantization and binarization) begins with the first non-zero element of the block according to a predetermined scan order. For example, positions L0~L 17 The coefficients of all are 0, and L 18 If the coefficient is not 0, the process is L 18 Starting with the coefficient of L, according to the scan order, 18 Perform the same process on each coefficient after that.
[0045] About residual coding In video coding, residual coding is used to convert quantization levels into a bitstream. After quantization, in the case of an N×M transform unit (TU) coding block, there are N×M quantization levels. These N×M levels may be zero or non-zero. If a level is not in binary form, non-zero levels are further binarized into a binary bin. Contextual self-adaptive binary arithmetic coding (CABAC) can further compress the bin into bits. Furthermore, there are two coding methods based on context modeling. Specifically, one method is to self-adaptively update the context model based on adjacent coding information. Such a method is called a context coding method, and the bin coded in this method is called a context coding bin. In contrast, the other method assumes that the probability of being 1 or 0 is always 50%, and therefore always uses a fixed context model without self-adaptation. Such a method is called a bypass method, and the bin coded in this method is called a bypass bin.
[0046] In a VVC regular residual coding (RRC) block, the position of the last non-zero level is defined as the position of the last non-zero level along the coding scan order. The representation of the 2D coordinates of the last non-zero level (last_sig_coeff_x and last_sig_coeff_y) includes a total of four prefix and suffix syntax elements: last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. First, the syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix are coded using the context coding method. If last_sig_coeff_x_suffix and last_sig_coeff_y_suffix are present, they are coded using the bypass method. An RRC block can consist of multiple predefined subblocks. The syntax element sb_coded_flag is used to indicate whether all levels of the current subblock are equal to 0. If sb_coded_flag is equal to 1, the current subblock has at least one non-zero coefficient. If sb_coded_flag is equal to 0, all coefficients of the current subblock are 0. However, the sb_coded_flag of the last non-zero subblock that has the last non-zero level is derived as 1 from last_sig_coeff_x and last_sig_coeff_y according to the coding scan order, without being encoded into the bitstream. Furthermore, the sb_coded_flag of the top-left subblock containing the DC position is also derived as 1 without being encoded into the bitstream. The syntax element of sb_coded_flag in the bitstream is encoded by the context coding method. RRC encodes each subblock starting from the last non-zero subblock in the reverse coding scan order, as discussed with respect to Figure 5.
[0047] To ensure worst-case throughput, a predefined value, remBinsPassl, is used to limit the maximum number of context-encoded bins. Within a single subblock, RRC encodes the level at each position in reverse encoding scan order. If remBinsPassl is greater than 4, when encoding the current level, a flag named sig_coeff_flag is first encoded into the bitstream to indicate whether the level is 0 or not. If the level is not 0, abs_level_gtx_flag[n][0] is encoded to indicate whether the absolute level is 1 or greater than 1, where n is the index of the current position in the subblock along the scan order. If the absolute level is greater than 1, par_level_flag is encoded to indicate whether the level is odd or even in VVC, and abs_level_gtx_flag[n][l] exists. The flags par_level_flag and abs_level_gtx_flag[n][l] are used together to indicate whether the level is 2, 3, or greater than 3. After encoding each of the above syntax elements into context-encoded bins, the value of remBinsPassl is decreased by one.
[0048] If the absolute level is greater than 3, or if the value of remBinsPassl is 4 or less, after encoding the above bin using the context encoding method, the other two syntax elements abs_remainder and dec_abs_level can be encoded into bypass-encoded bin for the remaining levels. Furthermore, the codes of each level in the block are also encoded to represent the quantization level, and the codes of each level in the block are encoded into bypass-encoded bin.
[0049] Another residual coding method allows for the conditional parsing of syntax elements using abs_level_gtxX_flag and the remaining level for level coding of residual blocks, with the corresponding binarization of the absolute value of the level being shown in Table 1. Here, abs_level_gtxX_flag indicates whether the absolute value of the level is greater than X, where X is an integer such as 0, 1, 2, or N. If abs_level_gtxY_flag is 0, the flag abs_level_gtx(Y+1) does not exist, where Y is an integer from 0 to N-1. If abs_level_gtxY_flag is 1, the flag abs_level_gtx(Y+1) exists. Furthermore, if abs_level_gtxN_flag is 0, the remaining level does not exist. If abs_level_gtxN_flag is 1, the remaining level exists and is the value after removing (N+1) from the level. Typically, the abs_level_gtxX_flag is encoded using a context coding method, and the remaining levels are encoded using a bypass method. [Table 1]
[0050] For blocks encoded in transform skip residual coding mode (TSRC), TSRC starts with the top-left subblock and encodes each subblock in the order of the coding scan. Similarly, the syntax element sb_coded_flag is used to indicate whether all residuals in the current subblock are equal to 0. If certain conditions occur, all syntax elements of sb_coded_flag for all subblocks except the last subblock are encoded into the bitstream. If all sb_coded_flag for all subblocks before the last subblock are not equal to 1, the sb_coded_flag for the last subblock is derived as 1 without being encoded into the bitstream. To ensure worst-case throughput, a predefined value RemCcbs is used to limit the maximum context coding bin. If the current subblock has a non-zero level, TSRC encodes the level at each position in the coding scan order. If RemCcbs is greater than 4, the next syntax element is encoded using the context coding method. For each level, first, sig_coeff_flag is encoded into a bitstream to indicate whether the level is 0 or not. If the level is not 0, coeff_sign_flag is encoded to indicate whether the level is positive or negative. Then, abs_level_gtx_flag[n][0] is encoded to indicate whether the current absolute level at the current position is greater than 1, where n is the index of the current position in the subblock according to the scan order. If abs_level_gtx_flag[n][0] is not 0, par_level_flag is encoded. After encoding each of the above syntax elements using the context coding method, the value of RemCcbs is decreased by 1.
[0051] After encoding the above syntax elements for all positions in the current subblock, if RemCcbs is still greater than 4, the context encoding method is used to encode up to four other abs_level_gtx_flag[n][j], where n is the index along the scan order of the current position in the subblock and j is from 1 to 4. After encoding each abs_level_gtx_flag[n][j], the value of RemCcbs is decreased by one. If RemCcbs is 4 or less, the syntax element abs_remainder is encoded for the current position in the subblock using the bypass method, if necessary. For positions where the absolute level has been encoded using the syntax element abs_remainder entirely by the bypass method, the coeff_sign_flag is also encoded using the bypass method. In summary, to limit the total number of context encoding bins and ensure worst-case throughput, there is either a predefined counter remBinsPassl in the RRC or RemCcbs in the TSRC.
[0052] Regarding the derivation of the Rice parameter In the current RRC design of VVC, two syntax elements, abs_remainder and dec_abs_level, encoded in the bypass bin, may exist in the bitstream of the remaining levels. abs_remainder and dec_abs_level are binarized using a combination of truncated rice (TR) and restricted k-th exponential Golomb (EGk) binarization processes as defined in the VVC specification, which requires binarizing a given level with rice parameters. To obtain optimal rice parameters, a local summation method is employed as described below.
[0053] The array AbsLevel[xC][yC] represents an array of absolute values of the transformation coefficient levels of the current transformation block for the color component index cIdx. Given an array AbsLevel[x][y] with a transformation block for the color component index cIdx and the luminance position (x0,y0) at the upper left corner, the local sum variable locSumAbs is derived in the manner defined by the following dummy code.
[0054] locSumAbs = 0 If (xC < (1 << log2TbWidth) - 1), then { locSumAbs += AbsLevel[ xC + 1 ][ yC ] If (xC < (1 << log2TbWidth) - 2), locSumAbs += AbsLevel[ xC + 2 ][ yC ] If ( yC < ( 1 << log2TbHeight ) - 1 ), locSumAbs += AbsLevel[ xC + 1 ][ yC + 1 ] } If ( yC < ( 1 << log2TbHeight ) - 1 ), then { locSumAbs += AbsLevel[ xC ][ yC + 1 ] If ( yC < ( 1 << log2TbHeight ) - 2 ), locSumAbs += AbsLevel[ xC ][ yC + 2 ] } locSumAbs = Clip3( 0, 31, locSumAbs - baseLevel * 5 )
[0055] Here, log2TbWidth and log2TbHeight are the base-2 logarithms of the width and height of the transformation block, respectively. For abs_remainder and dec_abs_level, the variables baseLevel are 4 and 0, respectively. Given the local summation variable locSumAbs, the rice parameter cRiceParam is derived using the method specified in Table 2. [Table 2]
[0056] History-based Rice parameter derivation If a coefficient lies on the boundary of a TU, or if decoding is performed first using the Rice method, the template computation used for Rice parameter derivation may result in inaccurate coefficient estimations. For these coefficients, the template computation is biased towards zero because some templates may lie outside the TU and be interpreted or initialized as having a value of 0. Figure 6 shows an example of a template pattern for calculating locSumAbs for a coefficient located near the boundary of a TU. Figure 6 shows CTU602 divided into multiple CUs, each CU containing multiple TUs. For TU604, the current coefficient's position is shown by a black block, and the position of adjacent samples in the template pattern is shown by patterned blocks. The patterned blocks indicate a given region of the current coefficient for calculating the local sum variable locSumAbs.
[0057] In Figure 6, because the current coefficient 606 is close to the boundary of TU604, some of the adjacent samples of the current coefficient 606 in the template pattern (e.g., adjacent samples 608B and 608E) are located outside the TU boundary. In the above Rice parameter derivation, when calculating the local summation variable locSumAbs, these adjacent samples located outside the boundary are set to 0, resulting in an inaccurate Rice parameter derivation. For high-bit-depth samples (e.g., more than 10 bits), the number of adjacent samples located outside the TU boundary may be large. Setting these large numbers to 0 introduces even more error into the Rice parameter derivation.
[0058] To improve the accuracy of Rice estimation based on computation templates, we propose updating the local sum variable locSumAbs using historically derived values, rather than initializing it to zero, for template locations located outside the current TU. The implementation of this method is described below by an excerpt from the VVC specification text of Clause 9.3.3.2, where the proposed text is underlined.
[0059] To maintain the history of adjacent coefficient / sample values, a history counter StatCoeff[cIdx] for each color component is used, where cIdx = 0, 1, or 2, representing the three color components Y, U, and V, respectively. If the CTU is the first CTU in a partition (e.g., picture, slice, or tile), StatCoeff[cIdx] is initialized as follows: StatCoeff[idx]=2*Floor(Log2(BitDepth-10) (1)
[0060] Here, BitDepth defines the bit depth of the video luminance and chroma array samples, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. Before TU decoding and history counter updating, the substitution variable HistValue is initialized as follows: HistValue[cIdx]=1< <StatCoeff[cIdx] (2)
[0061] The substitution variable HistValue is used as an estimate of adjacent samples located outside the TU boundary (for example, adjacent samples having horizontal or vertical coordinates located outside the TU). The local summation variable locSumAbs is re-derived according to the scheme defined by the following dummy code, where the changes are underlined. locSumAbs = 0 If (xC < (1 << log2TbWidth) - 1), then { locSumAbs += AbsLevel[ xC + 1 ][ yC ] If (xC < (1 << log2TbWidth) - 2), locSumAbs += AbsLevel[ xC + 2 ][ yC ] Otherwise locSumAbs += HistValue If ( yC < ( 1 << log2TbHeight ) - 1 ), locSumAbs += AbsLevel[ xC + 1 ][ yC + 1 ] Otherwise locSumAbs += HistValue } Otherwise locSumAbs += 2 * HistValue If ( yC < ( 1 << log2TbHeight ) - 1 ), then { locSumAbs += AbsLevel[ xC ][ yC + 1 ] If ( yC < ( 1 << log2TbHeight ) - 2 ), locSumAbs += AbsLevel[ xC ][ yC + 2 ] Otherwise locSumAbs += HistValue} Otherwise locSumAbs += HistValue
[0062] The history counter StatCoeff is updated once per TU by an exponential moving average process from the transformation coefficient of the first non-zero Golomb-Rice coding (abs_remainder[cIdx] or dec_abs_level[cIdx]). When the transformation coefficient of the first non-zero Golomb-Rice coding in a TU is coded as abs_remainder, the history counter StatCoeff for the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (3)
[0063] When the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, update the history counter StatCoeff for the color component cIdx as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1 (4)
[0064] The updated StatCoeff can be used to calculate the substitution variable HistValue for the next TU according to equation (2) before decoding the next TU.
[0065] About Wavefront Parallel Processing (WPP) WPP is designed to provide a parallel coding algorithm. When WPP is enabled in VVC, each CTU row in a frame, tile, or slice constitutes a single partition. WPP is enabled / disabled by the SPS element sps_entropy_coding_sync_enabled_flag. Figure 7 shows an example of a tile with WPP enabled. In Figure 7, each CTU row in the tile is processed with a delay of one CTU relative to the immediately preceding CTU row. In this way, if palette coding is enabled at the end of each CTU row, it does not break dependencies between consecutive CTU rows at partition boundaries, except for the CABAC context variables and palette predictors. To mitigate potential losses in coding efficiency, the contents of the self-adapted CABAC context variables and palette predictors are propagated from the first coded CTU of the immediately preceding CTU row to the first CTU of the current CTU row. WPP does not alter the normal raster scan order of CTUs.
[0066] When WPP is enabled, multiple threads (up to the number of CTU rows in a partition (e.g., tile, slice, or frame)) operate in parallel to process each CTU row. By using WPP in the decoder, each decoding thread processes one CTU row in a partition. For each CTU, thread processing must be scheduled so that the decoding of the CTU adjacent to the top in the immediately preceding CTU row is complete. After completing the encoding of the first CTU in each CTU row except the last CTU row, a small additional overhead of WPP is added so that all CABAC context variables and the contents of the palette predictor can be stored.
[0067] When the history-based Rice parameter derivation discussed above is effective for high-bit-depth and high-bit-rate video coding, the last StatCoeff in the immediately preceding CTU row is passed to the first TU in the current CTU row. Therefore, if WPP is enabled simultaneously, this process interferes with WPP and disrupts its parallelism. In this disclosure, several solutions are proposed to resolve this problem when parallel coding (e.g., WPP) is enabled.
[0068] In one embodiment, interference of history-based Rice parameter derivation to parallel coding is eliminated by removing the interdependence between CTUs of different CTU rows when calculating the history counter StatCoeff. In this embodiment, instead of using the history counter StatCoeff value obtained from the immediately preceding CTU row, the first abs_remainder[cIdx] or dec_abs_level[cIdx] within each CTU row of a partition (e.g., frame or tile or slice) is coded using the initial value of StatCoeff[cIdx], where cIdx is the index of the color component.
[0069] As an example, the initial value of StatCoeff[cIdx] can be determined as follows: StatCoeff[idx]=2*Floor(Log2(BitDepth-10)) (5)
[0070] Here, BitDepth specifies the bit depth of the samples in the luminance or chroma array, and Floor(x) represents the largest integer less than or equal to x. As another example, the initial value of StatCoeff[cIdx] can be determined as follows: StatCoeff[idx]=Clip(MIN_Stat,MAX_Stat,(int)((19-QP) / 6))-l (6)
[0071] Here, MIN_Stat and MAX_Stat are two predefined integers, QP is the initial QP for each slice, and Clip() is an operation defined as follows: JPEG0007920275000003.jpg30170
[0072] Before encoding the first TU of each CTU row in a partition (e.g., frame, tile, or slice), the substitution variable HistValue is calculated as follows: HistValue[cIdx]=1< <StatCoeff[cIdx] (8)
[0073] HistValue can be used to calculate the local sum variable locSumAbs described above. HistValue can be updated once per TU by an exponential moving average process from the transformation coefficient of the first non-zero Golomb-Rice coding (abs_remainder[cIdx] or dec_abs_level[cIdx]). When the transformation coefficient of the first non-zero Golomb-Rice coding in a TU is coded as abs_remainder, the history counter StatCoeff[cIdx] of the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (9)
[0074] When the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, update the history counter StatCoeff[cIdx] for the color component cIdx as follows: StatCoeff[cIdx] = (StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1 (10)
[0075] The updated StatCoeff[cIdx] is used to calculate the substitution variable HistValue for the next TU of the current CTU or the first TU of the next CTU in the current CTU row, as shown in equation (8).
[0076] Figure 8 shows an example of frame 802 and the CTUs contained within it. In this example, frame 802 contains two tiles, tile 804A and tile 804B. Tile 804A contains four CTU rows, CTU row 1 to CTU row 4. The first CTU row contains CTU0 to CTU9, and the second CTU row contains CTU10 to CTU19. Similarly, tile 804B also contains four CTU rows, CTU row 1' to CTU row 4'. The first CTU row contains 10 CTUs, CTU0' to CTU9', the second CTU row contains CTU10' to CTU19', and so on.
[0077] According to this embodiment, the initial value of StatCoeff[cIdx] for tile 804A can be determined according to equation (5) or (6). Before encoding the first TU of each CTU row in CTU row 1 to CTU row 4, the substitution variable HistValue[cIdx] is calculated using equation (8) and the initial value of StatCoeff[cIdx]. For example, before encoding the first TU of CTU0, the variable HistValue is calculated using equation (8). This value of HistValue is used to determine the local sum variable locSumAbs of the coefficients of the first TU, and the local sum variable locSumAbs is further used to determine the rice parameter of each coefficient of the first TU. When processing the first TU of the current CTU0, the history counter StatCoeff can be updated according to equation (9) or (10). Before processing the second TU in CTU0, the current value of StatCoeff is used to determine the HistValue of the second TU according to equation (8). Then, a similar process is applied to the second TU, using HistValue to determine the rice parameter and update StatCoeff. For the first TU in CTU1, HistValue is calculated using the latest StatCoeff from the TU in CTU0 according to equation (8). This process can be repeated until the last CTU in the current CTU row 1 (i.e., CTU9) is processed.
[0078] For the second CTU row of tile 804A, the history counter StatCoeff is initialized according to formula (5) or (6) before encoding the first TU of CTU10 (i.e., the first CTU of the second CTU row). For the TUs within the CTUs of the second CTU row, a process similar to the one described above for CTU row 1 is performed. Similarly, before encoding the first TUs in CTU20 and CTU30, the variable StatCoeff is initialized again according to formula (5) or (6).
[0079] Tile 804B can be processed in a similar manner. Before encoding the first TU in each of the CTU rows l' to CTU row 4' (i.e., CTU0', CTU10', CTU20', and CTU30'), the value of StatCoeff[cIdx] is initialized according to equation (5) or (6), and the history counter HistValue is calculated using equation (8). The calculated history counter HistValue is used to calculate the locSumAbs and Rice parameters for the first CTU in each CTU row and for the TUs in the remaining CTUs. Furthermore, the history counter StatCoeff may be updated at most once per TU according to equation (9) or (10), and the updated value of StatCoeff is used to determine the HistValue of the next TU in the same CTU row.
[0080] Figure 8 illustrates frame 802 containing two tiles, 804A and 804B, but the same process applies to other scenarios such as slices containing multiple tiles, and frames containing multiple slices. In any of these scenarios, before encoding the first TU within each CTU row of a partition (e.g., frame, tile, or slice), the value of the history counter StatCoeff[cIdx] is reset to its initial value to eliminate the dependency of CTU rows in the rice parameter derivation.
[0081] The possible specification changes for VVC, indicated by the underlined parts, are as follows: [Table 3]
[0082] Regarding 9.3.2.1, other possible specification changes for VVC are as follows: [Table 4]
[0083] Video sample bit depth VVC version 2 supports input video bit depths exceeding 10 bits. Higher video bit depths result in lower compression distortion and provide higher visual quality to the decoded video. To support high bit depths for input video, the semantics of the corresponding Sequence Parameter Set (SPS) syntax element sps_bitdepth_minus8 and Video Parameter Set (VPS) syntax element vps_ols_dpb_bitdepth_minus8[i] can be modified as follows:
[0084] sps_bitdepth_minus8 specifies the bit depth BitDepth of the luminance and chroma array samples, and the value QpBdOffset, which is the offset for the luminance and chroma quantization parameter range, as follows: BitDepth = 8 + sps_bitdepth_minus8 (x1) QpBdOffset = 6 * sps_bitdepth_minus8 (x2)
[0085] sps_bitdepth_minus8 must be within the range of 0 to 8 (including the endpoint).
[0086] If sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer in the i-th multilayer OLS specified by the VPS (where i is in the range of 0 to NumMultiLayerOlss - 1 (including the endpoint)), the bitstream consistency requirement is that the value of sps_bitdepth_minus8 must be less than or equal to the value of vps_ols_dpb_bitdepth_minus8[i].
[0087] vps_ols_dpb_bitdepth_minus8[i] specifies the maximum allowable sps_bitdepth_minus8 for all SPS referenced by the CLVS in the CVS of the i-th multilayer OLS. The value of vps_ols_dpb_bitdepth_minus8[i] must be within the range of 0 to 8 (including endpoints).
[0088] Note 2 - To decode the i-th multilayer OLS, the decoder can safely allocate memory to the DPB based on the values of the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_minus8[i].
[0089] As can be seen from the above, the bit depth BitDepth of the luminance and chroma array samples can be derived based on the SPS syntax element sps_bitdepth_minus8 according to equation (x1). Using the determined BitDepth value, the history counter StatCoeff, substitution variable HistValue, and Rice parameter can be derived as described above.
[0090] The VPS syntax element vps_ols_dpb_bitdepth_minus8[i] can be used to derive the size of the decoded picture buffer (DPB). Multiple video layers can exist in the encoded bitstream. The video parameter set is used to specify the corresponding syntax element. For video decoding, the DPB can be used to store a reference picture, which is used to generate a prediction signal that previously encoded pictures use when encoding other pictures. The DPB can also be used to reorder the decoded pictures so that they are output and / or displayed in the correct order. The DPB can also be used for the output delay specified by the virtual reference decoder. The decoded picture is stored persistently for a predetermined period specified in the DPB for the virtual reference decoder, and is output after the predetermined period has elapsed.
[0091] To safely allocate memory to the DPB, the size of the DPB is determined by the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_minus8[i] as follows: picture_size1 (in bits) = vps_ols_dpb_pic_width[i] * vps_ols_dpb_pic_height[i] * (vps_ols_dpb_bitdepth_minus8[i]+8) (vps_ols_dpb_chroma_format[ i ] == 0) / / If it is monochrome, picture_size = picture_size1; Otherwise, (vps_ols_dpb_chroma_format[ i ]== 1) / / If 4:2:0, picture_size = 1.5 * picture_size1; Otherwise, (vps_ols_dpb_chroma_format[ i ] == 2 / / If 4:2:2, picture_size = 2 * picture_size1; Otherwise, (vps_ols_dpb_chroma_format[ i ] == 3 / / If it is 4:4:4, picture_size = 3 * picture_size1;
[0092] Therefore, the size of the DPB is determined by picture_size. In other words, the size of the DPB can be determined according to the chroma format of the sample. If the video frame is a monochrome frame, the size of the frame waiting to buffer is determined to be the base picture size picture_size1. If the color subsampling of the color video frame is 4:2:0, the size of the frame is determined to be 1.5 times the base picture size picture_sizel. If the color subsampling of the color video frame is 4:2:2, the size of the frame is determined to be 2 times the base picture size picture_sizel. If the color subsampling of the color video frame is 4:4:4, the size of the frame is determined to be 3 times the base picture size picture_sizel. According to the color subsampling, the size of the DPB can be determined to be the number of frames stored in the DPB multiplied by the size of the frames.
[0093] Figure 9 shows an example of a video partition encoding process 900 according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video encoder 100) achieve the operation shown in Figure 9 by executing appropriate program code (e.g., program code implementing the entropy encoding module 116). For illustrative purposes, the process 900 will be described with reference to some examples shown in the figure. However, other implementations are also possible.
[0094] In block 902, process 900 includes accessing a partition of the video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in the CTU row shown in Figure 8. As shown in Figure 6, each CTU contains one or more CTUs, and each CTU contains multiple TUs for encoding.
[0095] In block 904, process 900 includes processing each CTU in the set of CTUs within a partition in order to encode the partition into bits, and block 904 includes 906-914. In block 906, process 900 includes determining whether the parallel coding algorithm is enabled and whether the current CTU is the first CTU in the CTU row. In some examples, parallel coding is indicated by a flag, where a value of 0 indicates that parallel coding is disabled and a value of 1 indicates that parallel coding is enabled. If the parallel coding algorithm is enabled and it is determined that the current CTU is the first CTU in the CTU row, process 900 includes setting the history counter StatCoeff to its initial value in block 908. If the history-based Rice parameter derivation is enabled as described above, the initial value of the history counter can be set according to equation (5) or (6), otherwise the initial value of the history counter is set to 0.
[0096] If the parallel coding algorithm is not valid, or if it is determined that the current CTU is not the first CTU in the CTU row, or if the history counter has been set in block 908, process 900 includes calculating the rice parameter of the TUs in the CTU based on the history counter in block 910. As described in detail above with respect to Figures 6-8, if the history counter is reset in block 908, the rice parameter of the TUs in the CTU is calculated based on the reset history counter or the subsequently updated history counter. If the history counter is not reset in block 908, the rice parameter of the TUs in the CTU is calculated based on the history counter updated in the previous CTU or the history counter subsequently updated in the current CTU.
[0097] In block 912, process 900 includes encoding the TU in the CTU into a binary representation based on the calculated rice parameters, for example, by a combination of truncated rice (TR) and restricted k-th order EGk as defined in the VVC specification. In block 914, process 900 encodes the binary representation of the CTU into bits to be included in the video bitstream. For example, this can be done using context-self-adaptive binary arithmetic coding (CABAC) discussed above. In block 916, process 900 includes outputting the encoded video bitstream.
[0098] Figure 10 shows an example of a video partition decoding process 1000 according to some embodiments of the present disclosure. One or more computing devices can achieve the operation shown in Figure 10 by executing appropriate program code. For example, a computing device implementing a video decoder 200 can achieve the operation shown in Figure 10 by executing program code for the entropy decoding module 216, the inverse quantization module 218, and the inverse transform module 219. For illustrative purposes, the process 1000 will be described with reference to some examples shown in the figure. However, other implementations are also possible.
[0099] In block 1002, process 1000 includes accessing a binary string or binary representation representing a partition of the video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in the CTU row shown in Figure 8. As shown in Figure 6, each CTU contains one or more CTUs, and each CTU contains multiple TUs for encoding.
[0100] In block 1004, process 1000 includes processing the binary string of each CTU in the CTU set within the partition to generate a decoded sample of the partition, and block 1004 includes 1006-1014. In block 1006, process 1000 includes determining whether the parallel coding algorithm is enabled and whether the current CTU is the first CTU in the CTU row. Parallel coding is indicated by a flag, where a value of 0 indicates that parallel coding is disabled, and a value of 1 indicates that parallel coding is enabled. If the parallel coding algorithm is enabled and it is determined that the current CTU is the first CTU in the CTU row, process 1000 includes setting the history counter StatCoeff to its initial value in block 1008. If history-based Rice parameter derivation is enabled as described above, the initial value of the history counter can be set according to equation (5) or (6); otherwise, the initial value of the history counter is set to 0.
[0101] If the parallel coding algorithm is not valid, or if it is determined that the current CTU is not the first CTU in the CTU row, or if the history counter has been set in block 1008, process 1000 includes calculating the rice parameter of the TUs in the CTU based on the history counter in block 1010. As described in detail above with respect to Figures 6-8, if the history counter is reset in block 1008, the rice parameter of the TUs in the CTU is calculated based on the reset history counter or the subsequently updated history counter. If the history counter is not reset in block 1008, the rice parameter of the TUs in the CTU is calculated based on the history counter updated in the previous CTU or the subsequently updated history counter in the current CTU.
[0102] In block 1012, process 1000 includes decoding the binary string or binary representation of the TU in the CTU into coefficient values based on the calculated Rice parameters, for example, by a combination of truncated Rice (TR) and restricted k-th order EGK as defined in the VVC specification. In block 1014, process 1000 includes reconstructing the pixel values of the TU in the CTU, for example by inverse quantization and inverse transform as discussed above with respect to Figure 2. In block 1016, process 1000 includes outputting the decoded partition of the video.
[0103] In another embodiment, the dependencies between CTUs when calculating the history counter StatCoeff are the same as the dependencies between CTUs in a parallel coding algorithm (e.g., WPP). For example, the history counter StatCoeff of a partition's (e.g., frame, tile, or slice) CTU rows can be calculated based on the coefficient values of the first N or fewer CTUs in the immediately preceding CTU row, where N is the maximum delay between two consecutive CTU rows allowed in a parallel coding algorithm. In this way, the dependencies between two consecutive CTU rows when calculating the history counter StatCoeff are limited to be smaller than (and therefore the same as) the dependencies between CTUs when performing parallel processing.
[0104] This embodiment can be implemented using a storage synchronization process. For example, in the above WPP, the delay between two consecutive CTU rows is 1 CTU, so N=1. In the storage process, after encoding the last TU of the first CTU of each CTU row (except the last CTU row), StatCoeff[cIdx] can be stored in the storage variable StatCoeffWpp[cIdx], and for each CTU row other than the first CTU row, the synchronization process for Rice parameter derivation is applied before encoding the first TU. In the synchronization process, StatCoeff[cIdx] is synchronized with StatCoeffWpp[cIdx] stored from the previous CTU row.
[0105] As described above, before encoding the first TU within each CTU row, calculate the variable HistValue as follows: HistValue[cIdx] = 1 << StatCoeff[cIdx] (11)
[0106] If the current CTU row is the first CTU row in the partition, StatCoeff[cIdx] can be initialized according to equation (5) or (6). The calculated HistValue is used to determine the local sum variable locSumAbs, which is further used to determine the Rice parameter of the TU in the current CTU. StatCoeff is updated once per TU by an exponential moving average process from the conversion coefficients (abs_remainder[cIdx] or dec_abs_level[cIdx]) of the first non-zero Golomb-Rice coding, as described above with respect to equations (9) and (10).
[0107] After encoding the last TU of the first CTU in the first CTU row, in the storage step, StatCoeff[cIdx] can be stored as StatCoeffWpp[cIdx] as follows: StatCoeffWpp[cIdx] = StatCoeff[cIdx] (12)
[0108] The encoding of the remaining CTUs in the first CTU row can be performed in a manner similar to that described with respect to the first embodiment.
[0109] The StatCoeff[cIdx] can be obtained by a synchronization step before encoding the second CTU row and the first TU in any subsequent CTU row. StatCoeff[cIdx] = StatCoeffWpp[cIdx] (13)
[0110] Using the obtained StatCoeff[cIdx] value, calculate HistValue according to equation (11). The remaining process for the CTU row is the same as the process for the first CTU row.
[0111] The possible changes to the VVC specification are as follows (changes are underlined): [Table 5] JPEG0007920275000007.jpg174170
[0112] Alternative history-based rice parameter derivation History-based rice parameter derivation can be implemented in an alternative manner. In this alternative implementation, if CTU is the first CTU in a partition (e.g., picture, slice, or tile), HistValue is initialized using the initial value of StatCoeff[cIdx] as follows: HistValue=sps_persistent_Rice_adaptation_enabled_flag?1< <StatCoeff[cIdx]:0 (14)
[0113] The initial HistValue is used to encode the first abs_remainder[cIdx] or dec_abs_level[cIdx] until HistValue is updated according to the following rules: When the conversion coefficient of the first non-zero Golomb-Rice coding in the TU is encoded as abs_remainder, the history counter for the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (15)
[0114] When the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, update the history counter for the color component cIdx as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1(16)
[0115] When the history counter StatCoeff[cIdx] is updated, HistValue is updated as shown in equation (17), and the updated HistValue is used to derive the Rice parameters of the remaining syntax elements abs_remainder and dec_abs_level until the new StatCoeff[cIdx] and HistValue[cIdx] are updated again. HistValue[cIdx] = 1 << StatCoeff[cIdx] (17)
[0116] Based on the current VVC specifications, the possible specification changes are as follows:
[0117] Section 7.3.11.11 (Residual Coding Syntax) will be amended as follows (additional content is underlined): [Table 6] JPEG0007920275000009.jpg159170
[0118] To resolve the dependency conflict between parallel encoding and alternative history-based Rice parameter derivation, after encoding the last TU of the first CTU in each CTU row, StatCoeff[cIdx] and HistValue[cIdx] for each color component are stored. The stored values of StatCoeff[cIdx] and HistValue[cIdx] are used to initialize StatCoeff[cIdx] and HistValue[cIdx] before processing the first TU of the first CTU in the subsequent CTU row.
[0119] This embodiment can be further implemented using a storage synchronization process. For example, in the storage process, after processing the last TU of the first CTU in each CTU row, StatCoeff[cIdx] and HistValue[cIdx] can be stored in storage variables, for example, StatCoeffWpp[cIdx] and HistValueWpp[cIdx] shown in equations (18) and (19), respectively. StatCoeffWpp[cIdx] = StatCoeff[cIdx] (18) HistValueWpp[cIdx] = HistValue[cIdx] (19)
[0120] For each CTU row other than the first CTU row, the synchronization process for Rice parameter derivation is applied before encoding the first TU. For example, as shown in equations (20) and (21), StatCoeff[cIdx] and HistValue[cIdx] are synchronized with StatCoeffWpp[cIdx] and HistValueWpp[cIdx] saved from the previous CTU row, respectively. StatCoeff[cIdx] = StatCoeffWpp[cIdx] (20) HistValue[cIdx] = HistValueWpp[cIdx] (21)
[0121] The synchronization variable HistValue is used to encode the first abs_remainder[cIdx] or dec_abs_level[cIdx] until HistValue is updated.
[0122] As described above, StatCoeff[cIdx] can be updated once per TU from the conversion coefficient of the first non-zero Golomb-Rice coding (abs_remainder[cIdx] or dec_abs_level[cIdx]) as shown in equation (15) or (16). When the history counter StatCoeff[cIdx] is updated, HistValue is updated according to equation (17), and the updated HistValue is used to derive the Rice parameters of the remaining syntax elements abs_remainder and dec_abs_level until the new StatCoeff[cIdx] and HistValue are updated again.
[0123] Based on the current VVC specifications, the possible specification changes indicated by the underline are as follows: [Table 7] JPEG0007920275000011.jpg213170
[0124] Figure 11 shows an example of a video partition encoding process 1100 according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video encoder 100) achieve the operation shown in Figure 11 by executing appropriate program code (e.g., program code implementing the entropy encoding module 116). For illustrative purposes, the process 1100 will be described with reference to some examples shown in the figure. However, other implementations are also possible.
[0125] In block 1102, process 1100 includes accessing a partition of the video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in the CTU row shown in Figure 8. As shown in Figure 6, each CTU contains one or more CTUs, and each CTU contains multiple TUs for encoding.
[0126] In block 1104, process 1100 includes processing each CTU in the set of CTUs within a partition in order to encode the partition into bits, and block 1104 includes blocks 1106-1118. In block 1106, process 1100 includes determining whether the parallel coding algorithm is enabled and whether the current CTU is the first CTU in the CTU row. In some examples, parallel coding is indicated by a flag, where a value of 0 indicates that parallel coding is disabled and a value of 1 indicates that parallel coding is enabled. If the parallel coding algorithm is enabled and it is determined that the current CTU is the first CTU in the CTU row, process 1100 includes determining in block 1107 whether the current CTU row is the first CTU row in the partition. If so, process 1100 includes setting the history counter StatCoeff to its initial value in block 1108. The initial value of the history counter can be set according to equation (5) or (6) as described above. If the current CTU row is not the first CTU row in the partition, process 1100 sets the history counter StatCoeff in block 1109 to the value stored in the history counter storage variable as shown in equation (13) or (20). In some examples, for example, when an alternative Rice parameter derivation is used, the value of the substitution variable HistValue may be reset to the stored value as shown in equation (21).
[0127] If the parallel coding algorithm is not valid, or if it is determined that the current CTU is not the first CTU in the CTU row, or if the history counter value is set in block 1108 or 1109, process 1100 includes calculating the rice parameter of the TU in the CTU based on the history counter (and the substitution variable HistValue if HistValue has been reset) in block 1110. If the history counter value is reset in block 1108 or 1109 as described above (see, for example, Figure 8, or alternative rice parameter derivation), the rice parameter of the TU in the CTU is calculated based on the reset history counter or the subsequently updated history counter. If the history counter is not reset in block 1108 or 1109, the rice parameter of the TU in the CTU is calculated based on the history counter updated in the previous CTU or the subsequently updated history counter in the current CTU.
[0128] In block 1112, process 1100 includes encoding the TU within the CTU into a binary representation based on the calculated rice parameters, for example, by a combination of truncated rice (TR) and restricted k-th order EGK as defined in the VVC specification. In block 1114, process 1100 encodes the binary representation of the CTU into bits to be included in the video bitstream. For example, this can be done using context-adaptive binary arithmetic coding (CABAC) discussed above.
[0129] In block 1116, process 1100 includes determining whether parallel coding is enabled and whether the CTU is the first CTU in the current CTU row. If so, in block 1118, process 1100 includes storing the value of the history counter in the history counter storage variable as shown in equation (12) or (18). In some examples, for example, when an alternative Rice parameter derivation is used, the value of the substitution variable HistValue may also be stored in the storage variable as shown in equation (19). In block 1120, process 1100 includes outputting the encoded video bitstream.
[0130] In some scenarios, a CTU in a CTU row that is not the first may be located at a partition boundary. For example, the first CTU in the second CTU row may not have a CTU in the partition above it. In these scenarios, the history counter of the CTU may be set to an initial value instead of a stored value. In this case, a new block 1107' can be added between blocks 1107 and 1109 in Figure 11 to determine whether the CTU is located at a partition boundary (for example, whether the CTU does not have an upper adjacent CTU in the partition). If so, process 1100 proceeds to block 1108 to set the history counter to its initial value; otherwise, process 1100 proceeds to block 1109 to set the history counter to a stored value. The remaining blocks in Figure 11 remain unchanged.
[0131] Figure 12 shows an example of a video partition decoding process 1200 according to some embodiments of the present disclosure. One or more computing devices can achieve the operation shown in Figure 12 by executing appropriate program code. For example, a computing device implementing a video decoder 200 can achieve the operation shown in Figure 12 by executing program code for the entropy decoding module 216, the inverse quantization module 218, and the inverse transform module 219. For illustrative purposes, the process 1200 will be described with reference to some examples shown in the figure. However, other implementations are also possible.
[0132] In block 1202, process 1200 includes accessing a binary string or binary representation representing a partition of the video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in the CTU row shown in Figure 8. As shown in Figure 6, each CTU contains one or more CTUs, and each CTU contains multiple TUs for encoding.
[0133] In block 1204, process 1200 includes processing the binary string of each CTU in the set of CTUs in the partition to generate a decoded sample of the partition, and block 1204 includes 1206-1218. In block 1206, process 1200 includes determining whether the parallel coding algorithm is enabled and whether the current CTU is the first CTU in the CTU row. Parallel coding is indicated by a flag, where a value of 0 indicates that parallel coding is disabled and a value of 1 indicates that parallel coding is enabled. If the parallel coding algorithm is enabled and it is determined that the current CTU is the first CTU in the CTU row, process 1200 includes determining in block 1207 whether the current CTU row is the first CTU row in the partition. If so, process 1200 includes setting the history counter StatCoeff to its initial value in block 1208. The initial value of the history counter can be set according to equation (5) or (6) as described above. If the current CTU row is not the first CTU row in the partition, process 1200 sets the history counter StatCoeff in block 1209 to the value stored in the history counter storage variable as shown in equation (13) or (20). In some examples, for example, when an alternative Rice parameter derivation is used, the value of the substitution variable HistValue may be reset to the stored value as shown in equation (21).
[0134] If the parallel coding algorithm is not valid, or if it is determined that the current CTU is not the first CTU in the CTU row, or if the history counter has been set in block 1208 or 1209, process 1200 includes calculating the rice parameter of the TU in the CTU in block 1210 based on the history counter (and the substitution variable HistValue, if HistValue has also been set). As described above (for example, in the derivation of the alternative rice parameter with respect to Figure 8), if the value of the history counter has been reset in block 1208 or 1209, the rice parameter of the TU in the CTU is calculated based on the reset history counter or the subsequently updated history counter. If the history counter has not been reset in block 1208 or 1209, the rice parameter of the TU in the CTU is calculated based on the history counter updated in the previous CTU or the subsequently updated history counter in the current CTU.
[0135] In block 1212, process 1200 includes decoding the binary string or binary representation of the TU in the CTU into coefficient values based on the calculated Rice parameters, for example, by a combination of truncated Rice (TR) and restricted k-th order EGK as defined in the VVC specification. In block 1214, process 1200 includes reconstructing the pixel values of the TU in the CTU, for example by inverse quantization and inverse transform as discussed above with respect to Figure 2.
[0136] In block 1216, process 1200 includes determining whether parallel coding is enabled and whether CTU is the first CTU in the current CTU row. If so, in block 1218, process 1200 includes storing the value of the history counter in the history counter storage variable as shown in equation (12) or (18). In some examples, for example, when an alternative Rice parameter derivation is used, the value of the substitution variable HistValue may also be stored in the storage variable as shown in equation (19). In block 1220, process 1200 includes outputting the decoded partition of the video.
[0137] In another embodiment, WPP or other parallel coding algorithms are prevented from coexisting with history-based rice parameter derivation within the bitstream. For example, if WPP is enabled, history-based rice parameter derivation may be disabled. If WPP is disabled, history-based rice parameter derivation may be enabled. Similarly, if history-based rice parameter derivation is enabled, WPP may be disabled. As an example, syntax changes can be made as follows:
[0138] 7.3.2.22 Sequence parameter set range extension syntax (additional content is underlined) [Table 8]
[0139] As another example, the corresponding semantics are changed as follows (the changes are underlined): [Table 9] JPEG0007920275000014.jpg83170
[0140] In the above explanation, TU was illustrated and explained in a diagram (e.g., Figure 6), but the same technique can be applied to the transformation block (TB). In other words, in the embodiments presented above (including the diagram), TU can also represent TB.
[0141] An example of a computing system for implementing dependent quantization in video encoding. Any suitable computing system can be used to perform the operations described herein. For example, Figure 13 shows an example of a computing device 1300 that can implement the video encoder 100 of Figure 1 or the video decoder 200 of Figure 2. In some embodiments, the computing device 1300 may include a processor 1312 which is communicatively coupled to a memory 1314 and executes computer-executable program code and / or access information stored in the memory 1314. The processor 1312 may include a microprocessor, an application-specific integrated circuit (ASIC), a state machine, or other processing device. The processor 1312 may include any one of a plurality of processing devices (including one processing device). Such a processor may include or communicate with a computer-readable medium that stores instructions, and when an instruction is executed by the processor 1312, the instruction causes the processor to perform the operations described herein.
[0142] Memory 1314 may include any suitable non-temporary computer-readable medium. The computer-readable medium may include any electronic, optical, magnetic, or other memory device that can provide computer-readable instructions or other program code to the processor. Non-exclusive examples of computer-readable medium include magnetic disks, memory chips, ROM, RAM, ASICs, configured processors, optical memory, magnetic tape or other magnetic memory, or any other medium from which a computer processor can read instructions. Instructions include processor-specific instructions generated by a compiler and / or interpreter based on code written in any suitable computer programming language, which include, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.
[0143] The computing device 1300 further comprises a bus 1316, which communicatively couples one or more components of the computing device 1300. The computing device 1300 further comprises several external or internal devices, such as input or output devices. For example, the computing device 1300 is shown to comprise an input / output (I / O) interface 1318, which can receive input from one or more input devices 1320 or provide output to one or more output devices 1322. One or more input devices 1320 and one or more output devices 1322 are communicatively coupled to the I / O interface 1318. The communication coupling can be achieved by any suitable method (e.g., connection via a printed circuit board, connection via cables, communication via wireless transmission, etc.). Non-limiting examples of input devices 1320 include touchscreens (e.g., one or more cameras for capturing touch areas, or pressure sensors for detecting pressure changes caused by touch), mice, keyboards, or any other devices for generating input events in response to the physical movements of a user of a computing device. Non-limiting examples of output devices 1322 include LCD screens, external monitors, speakers, or any other devices for displaying or otherwise presenting outputs generated by a computing device.
[0144] The computing device 1300 can execute program code, which causes the processor 1312 to perform one or more operations in the operations described in Figures 1 to 12 above. The program code may include a video encoder 100 or a video decoder 200. The program code may reside in memory 1314 or any suitable computer-readable medium and may be executed by the processor 1312 or any other suitable processor.
[0145] The computing device 1300 may further include at least one network interface device 1324. The network interface device 1324 may include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks 1328. Non-limiting examples of the network interface device 1324 include Ethernet work adapters, wireless modems, and the like. The computing device 1300 may transmit messages as electronic or optical signals through the network interface device 1324.
[0146] General Considerations This specification includes many specific details in order to provide a complete understanding of the subject matter described in the claims. However, as a person skilled in the art can understand, the subject matter described in the claims can be carried out even without these specific details. In other examples, methods, apparatus, or systems known to a person skilled in the art are not described in detail so as not to obscure the subject matter described in the claims.
[0147] Unless otherwise specified, throughout this specification, terms such as “processing,” “computing,” “calculating,” “determining,” and “identifying” are used to describe the operation or process of computing equipment (e.g., one or more computers or one or more similar electronic computing devices) that manipulates or transforms data that is represented as physical electronic or magnetic quantities in the memory, registers or other information storage devices, transmission devices or display devices of a computing platform.
[0148] One or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a steady result conditional on one or more inputs. A suitable computing device includes a computer system based on a general-purpose microprocessor that accesses stored software, which programs or configures the computing device, transforms the computing device from a general-purpose computing device to a dedicated computing device, and thereby realizes one or more embodiments of the subject matter herein. Any suitable programming, scripting, or other type or combination of language may be used to realize the teachings contained herein in software configured to program or configure the computing device.
[0149] Embodiments of the methods disclosed herein can be implemented in the operation of such computing devices. The order of the blocks shown in the above examples can be changed; for example, the order of the blocks can be changed, combined, or divided into subblocks. Some blocks or processes can be executed in parallel.
[0150] The use of “applies to” or “configured to” in this specification means open and inclusive language and does not exclude equipment applied to or configured to perform additional tasks or steps. Furthermore, the use of “based on” means open and inclusive in that a process, step, calculation or other movement “based on” one or more enumerated conditions or values may actually be based on additional conditions or exceed the enumerated values. The headings, lists and numbering included herein are for illustrative purposes only and are not intended to limit the scope of the explanation.
[0151] While the subject matter of this specification is described in detail with respect to specific embodiments, those skilled in the art will understand that such embodiments can be easily modified, transformed, and equivalents can be produced by understanding the foregoing. Therefore, it should be understood that this disclosure is presented for illustrative purposes only, not limitation, and does not exclude such modifications, transformations, and / or additions to the subject matter of this specification that can be easily made by those skilled in the art.
Claims
1. A method for decrypting video, Accessing a binary string representing a partition of the video, wherein the partition includes a plurality of coding tree units (CTUs), and the plurality of CTUs constitute one or more CTU rows. For each of the multiple CTUs in the partition, Before decoding the CTU, it is determined that the CTU is the first CTU of the current CTU row among the one-term CTU rows in the partition, In response to the determination that the CTU is the first CTU in the current CTU row within the partition, a history counter for calculating the color components of the Rice parameter is set to an initial value, Decoding the aforementioned CTU and performing the following: Outputting a decoded partition of the video, wherein the decoded partition includes a plurality of decoded CTUs within the partition, Decoding the aforementioned CTU is Based on the history counter, the Rice parameter of the conversion unit (TU) within the CTU is calculated, Based on the calculated Rice parameters, the binary string corresponding to the TU within the CTU is decoded into the coefficient value of the TU, This includes determining the pixel value of the TU within the CTU based on the coefficient value, Setting the history counter for calculating the color component cIdx of the aforementioned rice parameter to an initial value is, In response to the determination that history-based Rice parameter derivation is valid, The initial value is calculated by the operation StatCoeff[cIdx] = 2 * Floor(Log2(BitDepth-10)), where StatCoeff represents the history counter, BitDepth defines the video brightness and the bit depth of the chroma array samples, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x, including, How to decrypt a video.
2. The partition is a frame, or a slice, or a tile. The method for decoding a video according to claim 1.
3. Calculating the Rice parameter of the TU within the CTU based on the history counter is: Based on the aforementioned history counter, the substitution variable HistValue is determined, Using the values of adjacent coefficients in a predetermined region of the coefficients and the substitution variable HistValue, the local sum variable locSumAbs of the coefficients within the TU of the CTU is calculated. This includes deriving the Rice parameter of the TU based on the local summation variable locSumAbs, The method for decoding a video according to claim 1.
4. Determining the substitution variable HistValue based on the history counter includes determining the substitution variable HistValue of the color component cIdx by the operation HistValue[cIdx] = 1 << StatCoeff[cIdx], StatCoeff represents the history counter, The video decoding method according to claim 3.
5. Calculating the local sum variable locSumAbs of the coefficients within the TU of the CTU is: Determining that the adjacent coefficients among the plurality of adjacent coefficients in the predetermined region of the coefficient are located outside the TU, This includes calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighbor coefficient located outside the TU, The video decoding method according to claim 3.
6. Decoding the aforementioned CTU further involves, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in TU is coded as abs_remainder, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(abs_remainder[cIdx])) + 2) >> 1, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1, StatCount represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. The method for decoding a video according to claim 1.
7. Processing equipment, Memory configured to store computer programs, A processor, which executes the computer program stored in the memory, Accessing a binary string representing a video partition, wherein the partition includes a plurality of coding tree units (CTUs), and the plurality of CTUs constitute one or more CTU rows. For each of the multiple CTUs in the partition, Before decoding the CTU, it is determined that the CTU is the first CTU of the current CTU row among the one-term CTU rows in the partition, In response to the determination that the CTU is the first CTU in the current CTU row within the partition, a history counter for calculating the color components of the Rice parameter is set to an initial value, Decoding the aforementioned CTU and performing the following: Outputting the decoded partition of the video, wherein the decoded partition includes a plurality of decoded CTUs within the partition, It is configured to perform, Decoding the aforementioned CTU is Based on the history counter, the Rice parameter of the conversion unit (TU) within the CTU is calculated, Based on the calculated Rice parameters, the binary string corresponding to the TU within the CTU is decoded into the coefficient value of the TU, This includes determining the pixel value of the TU within the CTU based on the coefficient value, The processor, when setting the history counter for calculating the color component cIdx of the Rice parameter to an initial value, executes the computer program, further In response to the determination that history-based Rice parameter derivation is valid, The initial value is calculated by the operation StatCoeff[cIdx] = 2 * Floor(Log2(BitDepth-10)), where StatCoeff represents the history counter, BitDepth defines the video brightness and the bit depth of the chroma array samples, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. Processing equipment.
8. The partition is a frame, or a slice, or a tile. The processing apparatus according to claim 7.
9. The aforementioned processor further, Based on the aforementioned history counter, the substitution variable HistValue is determined, Using the values of adjacent coefficients in a predetermined region of the coefficients and the substitution variable HistValue, the local sum variable locSumAbs of the coefficients within the TU of the CTU is calculated. The system is configured to derive the Rice parameter of the TU based on the local summation variable locSumAbs, The processing apparatus according to claim 7.
10. The processor is further configured to determine the substitution variable HistValue of the color component cIdx by the operation HistValue[cIdx] = 1 << StatCoeff[cIdx], StatCoeff represents the history counter, The processing apparatus according to claim 9.
11. The aforementioned processor further, Determining that the adjacent coefficients among the plurality of adjacent coefficients in the predetermined region of the coefficient are located outside the TU, The system is configured to perform the following: use the substitution variable HistValue as the value of the neighbor coefficient located outside the TU to calculate the local sum variable locSumAbs. The processing apparatus according to claim 9.
12. The aforementioned processor further, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in TU is coded as abs_remainder, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(abs_remainder[cIdx])) + 2) >> 1, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, the system is configured to update the history counter of the color component cIdx such that StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1, StatCount represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. The processing apparatus according to claim 7.
13. A method of encoding video, Accessing the video partition, wherein the partition includes a plurality of coding tree units (CTUs), and the plurality of CTUs constitute one or more CTU rows. This includes processing the partition of the video to generate a binary representation of the partition, The aforementioned process is, For each of the multiple CTUs in the partition, Before encoding the CTU, it is determined that parallel encoding is enabled and that the CTU is the first CTU of the current CTU row among the one or more CTU rows in the partition, In response to the fact that the parallel coding is enabled and that the CTU is determined to be the first CTU of the current CTU row in the partition, a history counter for calculating the color components of the Rice parameter is set to an initial value, Encoding the aforementioned CTU and performing the following: The process includes encoding the binary representation of the partition into a bitstream of the video, Encoding the aforementioned CTU is Based on the history counter, the rice parameters of the conversion unit (TU) within the CTU are calculated, This includes encoding the coefficient value of the TU into a binary representation corresponding to the TU in the CTU, based on the calculated Rice parameters. Setting the history counter for calculating the color component cIdx of the aforementioned rice parameter to an initial value is, In response to the determination that history-based Rice parameter derivation is valid, The initial value is calculated by the operation StatCoeff[cIdx] = 2 * Floor(Log2(BitDepth-10)), where StatCoeff represents the history counter, BitDepth defines the video brightness and the bit depth of the chroma array samples, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. Video encoding methods.
14. The partition is a frame, or a slice, or a tile. The video encoding method according to claim 13.
15. Calculating the Rice parameter of the TU within the CTU based on the history counter is: Based on the aforementioned history counter, the substitution variable HistValue is determined, Using the values of adjacent coefficients in a predetermined region of the coefficients and the substitution variable HistValue, the local sum variable locSumAbs of the coefficients within the TU of the CTU is calculated. This includes deriving the Rice parameter of the TU based on the local summation variable locSumAbs, The video encoding method according to claim 13.
16. Determining the substitution variable HistValue based on the history counter includes determining the substitution variable HistValue of the color component cIdx by the operation HistValue[cIdx] = 1 << StatCoeff[cIdx], StatCoeff represents the history counter, The video encoding method according to claim 15.
17. Calculating the local sum variable locSumAbs of the coefficients within the TU of the CTU is: Determining that the adjacent coefficients among the plurality of adjacent coefficients in the predetermined region of the coefficient are located outside the TU, This includes calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighbor coefficient located outside the TU, The video encoding method according to claim 15.
18. Encoding the aforementioned CTU further involves, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in TU is coded as abs_remainder, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(abs_remainder[cIdx])) + 2) >> 1, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1, StatCount represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. The video encoding method according to claim 13.
19. Processing equipment, Memory configured to store computer programs, A processor, which executes the computer program stored in the memory, Accessing a video partition, wherein the partition includes a plurality of coding tree units (CTUs), and the plurality of CTUs constitute one or more CTU rows. This includes processing the partition of the video to generate a binary representation of the partition, The aforementioned process is, For each of the multiple CTUs in the partition, Before encoding the CTU, it is determined that parallel encoding is enabled and that the CTU is the first CTU of the current CTU row among the one or more CTU rows in the partition, In response to the fact that the parallel coding is enabled and that the CTU is determined to be the first CTU of the current CTU row in the partition, a history counter for calculating the color components of the Rice parameter is set to an initial value, Encoding the aforementioned CTU and performing the following: Encoding the binary representation of the partition into the bitstream of the video, It is configured to perform, Encoding the aforementioned CTU is Based on the history counter, the rice parameters of the conversion unit (TU) within the CTU are calculated, This includes encoding the coefficient value of the TU into a binary representation corresponding to the TU in the CTU, based on the calculated Rice parameters. The processor, when setting the history counter for calculating the color component cIdx of the Rice parameter to an initial value, executes the computer program, further, In response to the determination that history-based Rice parameter derivation is valid, The initial value is calculated by the operation StatCoeff[cIdx] = 2 * Floor(Log2(BitDepth-10)), where StatCoeff represents the history counter, BitDepth defines the video brightness and the bit depth of the chroma array samples, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. Processing equipment.
20. The partition is a frame, or a slice, or a tile. The processing apparatus according to claim 19.
21. The aforementioned processor further, Based on the aforementioned history counter, the substitution variable HistValue is determined, Using the values of adjacent coefficients in a predetermined region of the coefficients and the substitution variable HistValue, the local sum variable locSumAbs of the coefficients within the TU of the CTU is calculated. The system is configured to derive the Rice parameter of the TU based on the local summation variable locSumAbs, The processing apparatus according to claim 19.
22. The processor is further configured to determine the substitution variable HistValue of the color component cIdx by the operation HistValue[cIdx] = 1 << StatCoeff[cIdx], StatCoeff represents the history counter, The processing apparatus according to claim 21.
23. The aforementioned processor further, Determining that the adjacent coefficients among the plurality of adjacent coefficients in the predetermined region of the coefficient are located outside the TU, The system is configured to perform the following: use the substitution variable HistValue as the value of the neighbor coefficient located outside the TU to calculate the local sum variable locSumAbs. The processing apparatus according to claim 21.
24. The aforementioned processor further, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in TU is coded as abs_remainder, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(abs_remainder[cIdx])) + 2) >> 1, In response to the determination that the first non-zero Golomb-Rice coding conversion coefficient in the TU is coded as dec_abs_level, the system is configured to update the history counter of the color component cIdx such that StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1, StatCount represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base-2 logarithm of x. The processing apparatus according to claim 19.
25. A computer-readable storage medium, A computer-readable storage medium that stores a computer program and a bitstream, wherein the computer program causes a processor to execute the video encoding method described in claims 13 to 18 to generate the bitstream.
26. A computer-readable storage medium, A computer-readable storage medium that stores a computer program and a bitstream, wherein the computer program causes a processor to execute the video decoding method described in claims 1 to 6 to decode the bitstream.
Citation Information
Patent Citations
Matched palette encoding
JP2017522839A
Using Low-Complexity History to Derive RICE Parameters for High-Bit-Depth Video Coding
JP2023554264A
Low complexity history usage for rice parameter derivation for high bit-depth video coding
US20220201332A1