History-Based Rice Parameter Derivation for Wavefront Parallel Processing in Video Coding

JP2024532697A5Pending Publication Date: 2025-10-31GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024506578
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-04
Filing Date
2022-08-26
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing video encoding techniques face inefficiencies in compressing large video data without compromising visual quality, particularly due to dependencies between coding tree units (CTUs) that hinder parallel processing.

Method used

Implement history-based Rician parameter derivation for wavefront parallel processing, where history counters are initialized or synchronized across CTU rows to minimize dependencies, allowing parallel encoding and decoding of video data.

Benefits of technology

This approach enhances video encoding efficiency by reducing computational complexity and maintaining coding gain, thus improving the stability and speed of the encoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In some embodiments, a video decoder decodes video from a video bitstream using history-based Rice parameter derivation and wavefront parallel processing (WPP). The video decoder accesses a binary string representing a partition of the video and processes each coding tree unit (CTU) in the partition to generate decoded coefficient values ​​in the CTU. The process includes determining whether the WPP is valid and the CTU is a first CTU in a current CTU row in the partition before decoding the CTU, and if so, setting a history counter to an initial value. The process further includes calculating Rice parameters for transform units (TUs) in the CTU based on the value of the history counter to decode the CTU, and decoding the binary string corresponding to the TU in the CTU into coefficient values ​​for the TU based on the calculated Rice parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 260,600, filed on August 26, 2021, and entitled "History-Based Rice Parameter Derivations for Wavefront Parallel Processing in Video Coding," U.S. Provisional Application No. 63 / 262,078, filed on October 4, 2021, and entitled "History-Based Rice Parameter Derivations for Wavefront Parallel Processing in Video Coding," and U.S. Provisional Application No. 63 / 251,385, filed on October 1, 2021, and entitled "Representation of Bit Depth Range for VVC Operation Range Extension," the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to computer-implemented methods and systems for video processing. In particular, this disclosure relates to history-based Rice Parameter derivation for Wavefront Parallel Processing in video coding. [Background technology]

[0003] Ubiquitous camera-equipped devices such as smartphones, tablet computers and computers have made it easier than ever to take videos or images. However, the amount of data even for a short video can be quite large. Video coding techniques (including video encoding and video decoding) can compress video data into smaller sizes, thereby storing and transmitting various videos. Video coding is used in a wide range of applications (e.g., digital television broadcasting, video transmission over the Internet and mobile networks, real-time applications (e.g., video chat, video conferencing), DVDs, Blu-ray discs, etc.). It is expected to improve the efficiency of video coding schemes to reduce the consumption of storage capacity for storing videos and / or network bandwidth for transmitting videos. Summary of the Invention

[0004] Some embodiments relate to history-based Rice parameter derivation for wave-front parallel processing in video coding. In one example, a method for decoding a video is provided, the method including: accessing a binary string representing a partition of the video, the partition including a plurality of coding tree units (CTUs) that constitute one or more CTU rows; for each CTU of the plurality of CTUs in the partition, determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row of the one or more CTU rows in the partition before decoding the CTU; in response to determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row in the partition, setting a history counter for calculating color components of Rice parameters to an initial value; performing decoding of the CTU; and determining pixel values ​​of TUs in the CTU based on coefficient values; and outputting a decoded partition of the video, the decoded partition including a plurality of decoded CTUs in the partition, the decoding of the CTU including: and decoding a binary string corresponding to a TU in the CTU into a coefficient value of the TU based on the calculated Rice parameter.

[0005] In another example, a non-transitory computer readable medium is provided having stored thereon program code that, when executed by one or more processing devices, causes the processing devices to perform a plurality of operations, the plurality of operations being to access a binary string representing a partition of a video, the partition being represented by a plurality of coding tree units (CTUs). the plurality of CTUs in the partition constitute one or more CTU rows; for each CTU of the plurality of CTUs in the partition, before decoding the CTU, determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row among the one or more CTU rows in the partition, in response to determining that parallel encoding is enabled and the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of Rice parameters to an initial value, performing decoding of the CTU, and determining pixel values ​​of TUs in the CTU based on coefficient values; and outputting a decoded partition of the video, the decoding partition including the plurality of decoded CTUs in the partition, wherein decoding the CTU includes calculating Rice parameters of transform units (TUs) in the CTU based on the history counter, and decoding a binary string corresponding to the TUs in the CTU into coefficient values ​​of the TUs based on the calculated Rice parameters.

[0006] In another example, a system is provided that includes a processing device and a non-transitory computer-readable medium communicatively coupled to the processing device. The processing device is configured to perform a plurality of operations by executing program code stored on a non-transitory computer-readable medium, the plurality of operations including: accessing a binary string representing a partition of the video, the partition including a plurality of coding tree units (CTUs) forming one or more CTU rows; for each CTU of the plurality of CTUs in the partition, determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row of the one or more CTU rows in the partition before decoding the CTU; in response to determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row in the partition, setting a history counter for calculating color components of Rice parameters to an initial value; performing decoding of the CTU; and determining pixel values ​​of TUs in the CTU based on coefficient values; and outputting a decoded partition of the video, the decoded partition including a plurality of decoded CTUs in the partition, the decoding of the CTU including: and decoding a binary string corresponding to a TU in the CTU into a coefficient value of the TU based on the calculated Rice parameter.

[0007] In another example, a method for encoding a video is provided, the method including: accessing a partition of the video, the partition including a plurality of coding tree units (CTUs), the plurality of CTUs constituting one or more CTU rows; and processing the partition of the video to generate a binary representation of the partition, the processing including: for each CTU of the plurality of CTUs in the partition, determining that parallel encoding is enabled and the CTU is a first CTU of a current CTU row of the one or more CTU rows in the partition before encoding the CTU; in response to determining that parallel encoding is enabled and the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of a Rice parameter to an initial value; encoding the CTU; and encoding the binary representation of the partition into a video bitstream, the encoding of the CTU including: calculating Rice parameters for transform units (TUs) in the CTU based on the history counter; and encoding coefficient values ​​of the TUs into a binary representation corresponding to the TUs in the CTU based on the calculated Rice parameters.

[0008] In another example, a non-transitory computer readable medium having stored thereon program code, the program code being executed by one or more processing devices to cause the processing devices to perform a plurality of operations, the plurality of operations including: accessing a partition of a video, the partition including a plurality of coding tree units (CTUs) forming one or more CTU rows; and processing the partition of the video to generate a binary representation of the partition, the processing including, for each CTU of the plurality of CTUs in the partition, before encoding the CTU, parallel encoding is enabled and the CTU is encoded by one or more CTs in the partition. The method includes: determining that the CTU is the first CTU in the current CTU row of U rows; in response to determining that parallel encoding is enabled and that the CTU is the first CTU in the current CTU row in the partition, setting a history counter for calculating color components of Rice parameters to an initial value; encoding the CTU; and encoding the binary representation of the partition into a video bitstream, wherein encoding the CTU includes calculating Rice parameters of transform units (TUs) in the CTU based on the history counter; and encoding coefficient values ​​of the TUs into a binary representation corresponding to the TUs in the CTU based on the calculated Rice parameters.

[0009] In another example, a system is provided that includes a processing device and a non-transitory computer-readable medium communicatively connected to the processing device. The processing device is configured to perform a plurality of operations by executing program code stored on the non-transitory computer-readable medium, the plurality of operations including: accessing a partition of a video, the partition including a plurality of coding tree units (CTUs) that constitute one or more CTU rows; and processing the partition of the video to generate a binary representation of the partition, the processing including, for each CTU of the plurality of CTUs in the partition, before encoding the CTU, parallel encoding is enabled and the CTU is a current CTU row of the one or more CTU rows in the partition. The method includes: determining that the CTU is the first CTU; in response to determining that parallel encoding is enabled and that the CTU is the first CTU in the current CTU row in the partition, setting a history counter for calculating color components of Rice parameters to an initial value; encoding the CTU; and encoding the binary representation of the partition into a video bitstream, wherein encoding the CTU includes calculating Rice parameters of transform units (TUs) in the CTU based on the history counter; and encoding coefficient values ​​of the TUs into a binary representation corresponding to the TUs in the CTU based on the calculated Rice parameters.

[0010] These illustrative examples are not intended to limit or restrict the present disclosure, but to provide examples to aid in understanding the present disclosure. Other examples are discussed and further description is provided in specific embodiments. [Brief description of the drawings]

[0011] [Figure 1] 1 illustrates an example block diagram of a video encoder configured to implement embodiments presented herein. [Diagram 2]1 illustrates an example block diagram of a video decoder configured to implement embodiments presented herein. [Diagram 3] 1 illustrates an example of coding tree unit partitioning of a picture in a video, according to some embodiments of the present disclosure. [Figure 4] 1 illustrates an example of a coding unit split of a coding tree unit according to some embodiments of the present disclosure. [Diagram 5] 1 illustrates an example of a coding block in which the processing order of the elements of the coding block is predetermined. [Figure 6] 13 shows an example of a template pattern for computing local sum variables of coefficients located near a boundary of a transform unit. [Figure 7] 1 shows an example of a tile for which wavefront parallel processing is effective. [Figure 8] 1 illustrates an example of a frame for which a history counter is calculated, tiles included in the frame, and coding tree units according to some embodiments of the present disclosure. [Figure 9] 1 illustrates an example of a process for encoding a partition of a video, according to some embodiments of the present disclosure. [Figure 10] 1 illustrates an example of a process for decoding a partition of a video, according to some embodiments of the present disclosure. [Figure 11] 4 illustrates another example of a process for encoding a partition of a video, according to some embodiments of the present disclosure. [Figure 12] 1 illustrates another example of a process for decoding a partition of a video, according to some embodiments of the present disclosure. [Figure 13] 1 illustrates an example of a computing system for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] The features, embodiments and advantages of the present disclosure will be better understood after reading the following specific embodiments with reference to the drawings.

[0013] Each embodiment provides a history-based Rice parameter derivation for wavefront parallel processing in video coding. As described above, more and more video data is generated, stored, and transmitted. It is beneficial to increase the efficiency of video coding techniques, thereby representing video using less data without compromising the visual quality of the decoded video. One way to improve coding efficiency is to compress the processed video samples into a binary bitstream using as few bits as possible by entropy coding. Furthermore, since video usually contains a large amount of data, it is beneficial to reduce the processing time during coding (encoding and decoding). To this end, parallel processing can be adopted for video coding and decoding.

[0014] In entropy coding, video samples are binarized into binary bins, and the bins can be further compressed into bits by coding algorithms such as Context-Adaptive Binary Arithmetic Coding (CABAC). Binarization requires the calculation of binarization parameters, such as the truncated Rice (TR) specified in the Versatile Video Coding (VVC) specification, and the Rice parameter used in combination with the restricted k-th order Exp-Golomb (EGk) binarization process. To improve coding efficiency, a history-based Rice parameter derivation is used. The history-based Rice parameter derivation derives a Rice parameter for a transform unit (TU) in a current coding tree unit (CTU) of a partition (e.g., picture, slice, or tile) based on a history counter (denoted StatCoeff), which is calculated based on coefficients of previous TUs in the current CTU and previous CTUs in the partition. The history counter is then used to derive a substitution variable (denoted HistValue) to derive the Rice parameter. The history counter may be updated as the TU is processed. In some examples, the substitution variable of the TU is not changed when the history counter is updated.

[0015] Dependencies between previous and current CTUs in a partition for calculating history counters may conflict with the use of parallel processing, limiting or preventing the use of parallel processing, which may result in unstable or inefficient video encoding. Various embodiments described herein address these issues by reducing or eliminating dependencies between some CTUs in a partition, thereby achieving parallel processing to speed up the video processing process, or by detecting and avoiding collisions before they occur. The following non-limiting examples are provided to introduce some embodiments.

[0016] In one embodiment, the dependency between CTUs of different CTU rows when calculating the history counters is removed to eliminate the dependency conflict with parallel processing. For example, the history counter can be initialized anew for each CTU row of a partition. Before calculating the Rice parameter of the first CTU of the CTU row, the history counter can be set to an initial value. Subsequent history counters can be calculated based on the history counter values ​​of previous TUs in the same CTU row. In this way, in the history-based Rice parameter derivation, the dependency of CTUs is limited within the same CTU row, and does not interfere with the parallel processing in different CTU rows, while at the same time, the coding gain realized based on the history-based Rice parameter derivation can be benefited. Furthermore, the history-based Rice parameter derivation process is simplified and the computational complexity is reduced.

[0017] In another embodiment, the dependency between CTUs in calculating the history counters is consistent with the dependency between CTUs in parallel processing. For example, parallel encoding can be performed on the CTU rows of a partition, and there can be an N-CTU delay between two consecutive CTU rows. That is, processing of a CTU row starts after processing N CTUs of the previous CTU row. In this case, the history counter of a CTU row can be calculated based on samples in the first N or fewer CTUs in the previous CTU row. This can be achieved by a storage synchronization process. After processing the last TU in the first CTU of a CTU row, the history counter can be stored in a storage variable. Then, before processing the first TU in the first CTU of a subsequent CTU row, the history counter can be synchronized with the stored value in the storage variable.

[0018] In some examples, an alternative history-based Rice parameter derivation is used, which updates the substitution variable HistValue when the history counter StatCoeff is updated when processing a TU. To avoid dependency conflicts with the parallel encoding, the dependency between CTUs when calculating the history counter can be similarly limited to N or fewer CTUs. Similarly, a storage synchronization process can be realized. After processing the last TU in the first CTU of a CTU row, the history counter and the substitution variable can be stored in storage variables, respectively. Then, before processing the first TU in the first CTU of a subsequent CTU row, the history counter and the substitution variable can be synchronized with the stored values ​​in the corresponding storage variables.

[0019] In this way, the inter-CTU dependency of two consecutive CTU rows when calculating the history counter is limited to be smaller than (i.e., consistent with) the inter-CTU dependency when performing parallel encoding, so that the history counter calculation does not interfere with parallel processing and at the same time can benefit from the coding gain realized based on the history-based Rice parameter derivation.

[0020] Alternatively, parallel processing and history-based Rice parameter derivation are prevented from coexisting in the bitstream. For example, the video encoder can determine whether parallel processing is enabled. If parallel processing is enabled, history-based Rice parameter derivation is disabled, and if not, history-based Rice parameter derivation is enabled. Similarly, if the video encoder determines that history-based Rice parameter derivation is enabled, parallel processing is disabled, and vice versa.

[0021] Using the Rice parameters determined as above, the video encoder can binarize the prediction residual data (e.g., quantized transform coefficients of the residuals) into binary bins and further compress the bins into bits included in the video bitstream using an entropy coding algorithm. At the decoder side, the decoder can decode the bitstream back into binary bins, determine the Rice parameters using any method or combination of methods described above, and then determine the coefficients based on the binary bins. Additionally, inverse quantization and inverse transform can be performed on the coefficients to reconstruct the video block for display.

[0022] In some embodiments, a bit depth of the video samples (e.g., a bit depth for determining an initial value of a history counter StatCoeff) can be determined based on a Sequence Parameter Set (SPS) syntax element sps_bitdepth_minus8. The value of the SPS syntax element sps_bitdepth_minus8 ranges from 0 to 8. Similarly, a size of a decoded picture buffer (DPB) for storing decoded pictures can be determined based on a Video Parameter Set (VPS) syntax element vps_ols_dpb_bitdepth_minus8. The value of the VPS syntax element vps_ols_dpb_bitdepth_minus8 ranges from 0 to 8. Based on the determined DPB size, storage capacity can be allocated to the DPB. The determined bit depth and DPB can be used in the entire process of decoding the video bitstream into pictures.

[0023] As described herein, some embodiments provide improved video coding efficiency and computational efficiency by coordinating history-based Rice parameter derivation and parallel coding. In this manner, the stability of the coding process can be improved by avoiding conflicts between history-based Rice parameter derivation and parallel coding. Furthermore, by restricting the dependencies between CTUs in the history-based Rice parameter derivation to be equal to or less than those in the parallel coding, coding gain can be achieved by the history-based Rice parameter derivation without sacrificing the computational efficiency of the coding process. This technique may be an effective coding tool in future video coding standards.

[0024] Referring now to the drawings, Figure 1 shows an example block diagram of a video encoder configured to realize the embodiments presented herein. In the example shown in Figure 1, the video encoder 100 comprises a partition module 112, a transform module 114, a quantization module 115, an inverse quantization module 118, an inverse transform module 119, an in-loop filter module 120, an intra prediction module 126, an inter prediction module 124, a motion estimation module 122, a decoded picture buffer 130 and an entropy coding module 116.

[0025] The input of the video encoder 100 is an input video 102 including a sequence of pictures (also called frames or images). In a block-based video encoder, for each picture, the video encoder 100 employs a partition module 112 to divide the picture into blocks 104, each block including a number of pixels. A block may be a macroblock, a coding tree unit, a coding unit, a prediction unit, and / or a prediction block. A picture may include blocks of different sizes, and the block partitions of different pictures of a video may be different. Each block may be coded using different predictions (e.g., intra prediction or inter prediction or hybrid intra and inter prediction).

[0026] Typically, the first picture of a video signal is an intra-predicted picture, which is coded using only intra-prediction. In intra-prediction mode, blocks of a picture are predicted using only data from the same picture. An intra-predicted picture can be decoded in the absence of information from other pictures. To perform intra-prediction, the video encoder 100 shown in FIG. 1 may employ an intra-prediction module 126. The intra-prediction module 126 is configured to generate an intra-predicted block (prediction block 134) using reconstructed samples in a reconstructed block 136 of a neighboring block of the same picture. The intra-prediction is performed according to the intra-prediction mode selected for the block. Then, the video encoder 100 calculates the difference between the block 104 and the intra-predicted block 134. The difference is called a residual block 106.

[0027] To further remove redundancy from the block, the transform module 114 transforms the residual block 106 into a transform domain by applying a transform to the samples in the block. Examples of transforms include, but are not limited to, a discrete cosine transform (DCT) or a discrete sine transform (DST). The transformed values ​​are called transform coefficients, and the transform coefficients represent the residual block in the transform domain. In some examples, the residual block can be directly quantized without being transformed by the transform module 114. This is called a transform skip mode.

[0028] The video encoder 100 may further quantize the transform coefficients using a quantization module 115 to obtain quantized coefficients. Quantization involves dividing the samples by a quantization step size followed by rounding, and inverse quantization involves multiplying the quantized values ​​by the quantization step size. Such a quantization process is called scalar quantization. Quantization is used to reduce the dynamic range of video samples (transformed or untransformed), thereby representing the video samples using fewer bits.

[0029] Quantization of coefficients / samples within a block can be done independently, and this quantization method is used in some existing video compression standards (e.g., H.264 and HEVC). For an N×M block, the 2D coefficients of the block can be converted into a 1-D array in a specific scan order for coefficient quantization and encoding. Quantization of coefficients within a block can use scan order information. For example, the quantization of a given coefficient in a block can be determined by the state of the previous quantization value along the scan order. To further improve the coding efficiency, multiple quantizers can be used. Which quantizer is used to quantize the current coefficient depends on the information located before the current coefficient in the encoding / decoding scan order. This quantization method is called dependent quantization.

[0030] The quantization step size can be used to adjust the degree of quantization. For example, in the case of scalar quantization, different quantization step sizes can be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, and a larger quantization step size corresponds to coarser quantization. The quantization step size can be indicated by a quantization parameter (QP). Providing the quantization parameter in the encoded bitstream of the video allows a video decoder to apply the same quantization parameter for decoding.

[0031] The quantized samples are then encoded using an entropy encoding module 116 to further reduce the size of the video signal. The entropy encoding module 116 is configured to apply an entropy encoding algorithm to the quantized samples. In some examples, the quantized samples are binarized into binary bins, and the encoding algorithm further compresses the binary bins into bits. Examples of binarization methods include, but are not limited to, truncated Rice (TR) and k-th order Exp-Golomb (EGk) binarization, etc. To improve the encoding efficiency, a history-based Rice parameter derivation method is used, where the Rice parameters derived for a transform unit (TU) are based on variables obtained or updated from a previous TU. Examples of entropy coding algorithms include, but are not limited to, variable length coding (VLC) schemes, context adaptive VLC schemes (CAVLC), arithmetic coding schemes, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding techniques. The entropy coded data is added to a bitstream that outputs the coded video 132.

[0032] As described above, the intra prediction of a block of a picture uses a reconstructed block 136 from neighboring blocks. The generation of the reconstructed block 136 of a block involves the calculation of a reconstructed residual of the block. The reconstructed residual can be determined by applying inverse quantization and inverse transformation to the quantized residual of the block. The inverse quantization module 118 is configured to apply inverse quantization to the quantized samples to obtain dequantized coefficients. The inverse quantization module 118 applies an inverse scheme of the quantization scheme applied to the quantization module 115 by using the same quantization step size as the quantization module 115. The inverse transform module 119 is configured to apply an inverse transform (such as an inverse DCT or an inverse DST) of the transform applied to the transform module 114 to the dequantized samples. The output of the inverse transform module 119 is the reconstructed residual of the block in the pixel domain. The reconstructed block 136 in the pixel domain can be obtained by adding the reconstructed residual to the prediction block 134 of the block. For blocks whose transforms have been skipped, the inverse transform module 119 is not applied to those blocks. The dequantized samples are the reconstructed residuals of the block.

[0033] Inter prediction or intra prediction may be used to encode blocks of a subsequent picture following a first intra predicted picture. In inter prediction, prediction of a block in a picture is made based on one or more previously encoded video pictures. To perform inter prediction, video encoder 100 employs inter prediction module 124. Inter prediction module 124 is configured to perform motion compensation on the block based on a motion estimate provided by motion estimation module 122.

[0034] The motion estimation module 122 performs motion estimation by comparing a current block 104 of a current picture with a decoded reference picture 108. The decoded reference picture 108 is stored in the decoded picture buffer 130. The motion estimation module 122 selects a reference block from the decoded reference picture 108 that best matches the current block. The motion estimation module 122 further identifies an offset between the location (e.g., x, y coordinates) of the reference block and the location of the current block. The offset is referred to as a motion vector (MV) and is provided to the inter prediction module 124. In some cases, multiple reference blocks are identified for blocks in multiple decoded reference pictures 108. Thus, multiple motion vectors are generated and provided to the inter prediction module 124.

[0035] The inter prediction module 124 performs motion compensation using the motion vector and other inter prediction parameters to generate a prediction of the current block (i.e., the inter prediction block 134). For example, based on the motion vector, the inter prediction module 124 can identify a prediction block pointed to by the motion vector in a corresponding reference picture. If there are multiple prediction blocks, these prediction blocks are combined with some weights to generate the prediction block 134 of the current block.

[0036] For an inter-predicted block, video encoder 100 may subtract inter-predicted block 134 from block 104 to generate residual block 106. Residual block 106 may be transformed, quantized, and entropy coded in the same manner as the residual for an intra-predicted block described above. Similarly, the reconstructed block 136 for the inter-predicted block is obtained by inverse quantizing, inverse transforming, and combining the residual with the corresponding predicted block 134.

[0037] To obtain the decoded picture 108 for motion estimation, the reconstruction block 136 is processed by a loop filter module 120. The loop filter module 120 is configured to smooth pixel transformations to improve video quality. The loop filter module 120 is configured to implement one or more loop filters, such as a de-blocking filter, or a sample-adaptive offset (SAO) filter, or an adaptive loop filter (ALF), etc.

[0038] 2 shows an example of a video decoder 200 configured to realize the embodiments presented herein. The video decoder 200 processes encoded video 202 in a bitstream and generates decoded pictures 208. In the example shown in FIG. 2, the video decoder 200 includes an entropy decoding module 216, an inverse quantization module 218, an inverse transform module 219, a loop filter module 220, an intra prediction module 226, an inter prediction module 224, and a decoded picture buffer 230.

[0039] The entropy decoding module 216 is configured to perform entropy decoding of the encoded video 202. The entropy decoding module 216 decodes the coding parameters and other information, including quantized coefficients, intra-prediction parameters and inter-prediction parameters. In some examples, the entropy decoding module 216 decodes the bitstream of the encoded video 202 into a binary representation and then converts the binary representation into quantization levels of the coefficients. The entropy decoded coefficients are then inverse quantized by the inverse quantization module 218 and subsequently inverse transformed to the pixel domain by the inverse transform module 219. The functions of the inverse quantization module 218 and the inverse transform module 219 are similar to the inverse quantization module 118 and the inverse transform module 119, respectively, described with reference to FIG. 1 above. The inverse transformed residual blocks can be added to the corresponding prediction blocks 234 to generate the reconstruction blocks 236. For blocks whose transforms have been skipped, the inverse transform module 219 is not applied to those blocks. The dequantized samples generated by the inverse quantization module 118 are used to generate a reconstruction block 236 .

[0040] A prediction block 234 for a particular block is generated based on the prediction mode of the block. If the coding parameters of the block indicate that intra prediction is to be performed for the block, a reconstructed block 236 of a reference block in the same picture is provided to the intra prediction module 226 to generate the prediction block 234 for the block. If the coding parameters of the block indicate that inter prediction is to be performed for the block, the prediction block 234 is generated by the inter prediction module 224. The functions of the intra prediction module 226 and the inter prediction module 224 are similar to the intra prediction module 126 and the inter prediction module 124 of FIG. 1, respectively.

[0041] As described above with respect to FIG. 1, inter prediction involves one or more reference pictures. The video decoder 200 generates a decoded picture 208 of the reference picture by applying a loop filter module 220 to a reconstructed block of the reference picture. The decoded picture 208 is stored in a decoded picture buffer 230 for use by the inter prediction module 224 and also for output.

[0042] Referring now to FIG. 3, FIG. 3 illustrates an example of a division of coding tree units of a picture in a video according to some embodiments of the present disclosure. As described above with respect to FIG. 1 and FIG. 2, to code a picture of a video, the picture is divided into blocks, such as coding tree units (CTUs) 302 in VVC, as shown in FIG. 3. For example, the CTUs 302 may be blocks of 128×128 pixels. The CTUs are processed according to a specific order (for example, the order shown in FIG. 3). In some examples, as shown in FIG. 4, each CTU 302 in a picture may be divided into one or more coding units (CUs) 402, and the CUs 402 may be further divided into prediction units or transform units (TUs) for prediction and transformation. Depending on the coding scheme, the CTUs 302 may be divided into CUs 402 in different manners. For example, in VVC, the CUs 402 may be rectangular or square, and may be coded without division into prediction units or transform units. Each CU 402 may be the same size as its root CTU 302, or may be a subdivision as small as a 4x4 block of the root CTU 302. As shown in Figure 4, the division of CTUs 302 into CUs 402 in VVC may be a quadtree or binary or ternary tree division. In Figure 4, the solid lines indicate a quadtree division and the dashed lines indicate a binary or ternary tree division.

[0043] As described above with respect to FIG. 1 and FIG. 2, quantization is used to reduce the dynamic range of elements of a block in a video signal, and to represent the video signal using fewer bits. In some examples, before quantization, elements at a particular position of a block are called coefficients. After quantization, the quantized value of a coefficient is called a quantization level or levels. Quantization typically involves division by a quantization step size followed by rounding, whereas inverse quantization involves multiplication by the quantization step size. This quantization process is also called scalar quantization. Quantization of coefficients in a block can be done independently, and this independent quantization method is used in some existing video compression standards (e.g., H.264, HEVC, etc.). In other examples, for example, VVC employs dependent quantization.

[0044] For an N×M block, the 2-D coefficients of the block can be transformed into a 1-D array in a particular scan order for quantization and encoding of the coefficients, and the same scan order is used for encoding and decoding. Figure 5 shows an example of a coding block (e.g., a transform unit (TU)) that processes the coefficients of the coding block in a predefined scan order. In this example, the coding block 500 is 8×8 in size, and the process starts from the lower right corner at location L0 and ends at the upper left corner L1. 63 If the block 500 is a transform block, the predetermined order shown in FIG. 5 is from highest frequency to lowest frequency. In some examples, processing of the block (e.g., quantization and binarization) starts with the first non-zero element of the block according to a predetermined scan order. For example, from positions L0 to L 17 All coefficients of are 0 and L 18 If the coefficient of is not 0, the process is 18 Starting with the coefficients of,L,, in the scan order, 18 The process is then performed on each coefficient after .

[0045] About Residual Coding In video coding, residual coding is used to convert quantization levels into a bitstream. After quantization, for an N×M transform unit (TU) coding block, there are N×M quantization levels. These N×M levels may be zero or non-zero values. If the levels are not in binary form, the non-zero levels are further binarized into binary bins. Context self-adaptive binary arithmetic coding (CABAC) can further compress the bins into bits. In addition, there are two coding methods based on context modeling. Specifically, one method is to self-adaptively update the context model based on neighboring coding information. Such a method is called a context coding method, and the bins coded in this manner are called context coded bins. In contrast, the other method assumes that the probability of 1 or 0 is always 50%, and therefore always uses fixed context modeling without self-adaptation. Such a method is called a bypass method, and the bins coded in this manner are called bypass bins.

[0046] For regular residual coding (RRC) blocks in VVC, the position of the last non-zero level is defined as the position of the last non-zero level along the coding scan order. The representation of the 2D coordinates (last_sig_coeff_x and last_sig_coeff_y) of the last non-zero level includes a total of four prefix and suffix syntax elements: last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_y_suffix. First, code the syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix using the context coding method. If last_sig_coeff_x_suffix and last_sig_coeff_y_suffix exist, code last_sig_coeff_x_suffix and last_sig_coeff_y_suffix using the bypass method. An RRC block may be composed of multiple predefined sub-blocks. The syntax element sb_coded_flag is used to indicate whether all levels of the current sub-block are equal to 0 or not. If sb_coded_flag is equal to 1, there is at least one non-zero coefficient in the current sub-block. If sb_coded_flag is equal to 0, all coefficients of the current sub-block are zero. However, the sb_coded_flag of the last non-zero sub-block with the last non-zero level is derived as 1 from last_sig_coeff_x and last_sig_coeff_y according to the coding scan order without being coded in the bitstream. In addition, the sb_coded_flag of the top-left sub-block containing the DC position is also derived as 1 without being coded in the bitstream. The syntax element sb_coded_flag in the bitstream is coded by the context coding method. RRC codes sub-block by sub-block starting from the last non-zero sub-block in the reverse coding scan order as discussed with respect to FIG. 5.

[0047] To guarantee worst case throughput, a predefined value remBinsPassl is used to limit the maximum number of context coding bins. Within one subblock, RRC codes the levels for each position in the reverse coding scan order. If remBinsPassl is greater than 4, when coding the current level, first, a flag named sig_coeff_flag is coded into the bitstream to indicate whether the level is 0 or not. If the level is not 0, abs_level_gtx_flag[n][0] is coded to indicate whether the absolute level is 1 or greater than 1, where n is the index of the current position in the subblock along the scan order. If the absolute level is greater than 1, par_level_flag is coded to indicate whether the level is odd or even in VVC, and abs_level_gtx_flag[n][l] is present. The flags par_level_flag and abs_level_gtx_flag[n][l] are used together to indicate whether the level is 2, or 3, or greater than 3. After encoding each of the above syntax elements into a context-coded bin, the value of remBinsPassl is decreased by one.

[0048] If the absolute level is greater than 3 or the value of remBinsPassl is less than or equal to 4, after encoding the above bins by the context encoding method, the other two syntax elements abs_remainder and dec_abs_level can be encoded into bypass encoding bins for the remaining levels. In addition, the code of each level of the block is also encoded to represent a quantization level, and the code of each level of the block is encoded into a bypass encoding bin.

[0049] In another residual coding method, for level coding of the residual block, the syntax elements can be conditionally parsed using abs_level_gtxX_flag and the remaining levels, and the corresponding binarization of the absolute value of the levels is as shown in Table 1. Here, abs_level_gtxX_flag indicates whether the absolute value of the level is greater than X, where X is an integer such as 0, 1, 2 or N. If abs_level_gtxY_flag is 0, the flag abs_level_gtx(Y+1) does not exist, where Y is an integer from 0 to N-1. If abs_level_gtxY_flag is 1, the flag abs_level_gtx(Y+1) exists. Furthermore, if abs_level_gtxN_flag is 0, the remaining levels do not exist. If abs_level_gtxN_flag is 1, the remaining levels exist and are the values ​​after removing (N+1) from the levels. Typically, the context coding method is used to code abs_level_gtxX_flag, and the remaining levels are coded using the bypass method. [Table 1]

[0050] For blocks coded in transform skip residual coding mode (TSRC), TSRC starts with the top-left subblock and codes each subblock along the coding scan order. Similarly, the syntax element sb_coded_flag is used to indicate whether all residuals of the current subblock are equal to 0 or not. If a certain condition occurs, all syntax elements of sb_coded_flag of all subblocks except the last subblock are coded into the bitstream. If all sb_coded_flag of all subblocks before the last subblock are not equal to 1, the sb_coded_flag of the last subblock is derived as 1 without coding it into the bitstream. To guarantee the worst-case throughput, a predefined value RemCcbs is used to limit the maximum context coding bin. If there is a non-zero level in the current subblock, TSRC codes the level of each position in the coding scan order. If RemCcbs is greater than 4, the context coding method is used to code the next syntax element. For each level, first, sig_coeff_flag is coded into the bitstream to indicate whether the level is 0 or not. If the level is not 0, then coeff_sign_flag is coded to indicate whether the level is positive or negative. Then, abs_level_gtx_flag[n][0] is coded to indicate whether the current absolute level at the current position is greater than 1 or not, where n is the index of the current position in the subblock along the scan order. If abs_level_gtx_flag[n][0] is not 0, then par_level_flag is coded. After coding each of the above syntax elements using the context coding method, the value of RemCcbs is decremented by one.

[0051] If, after encoding the above syntax elements for all positions in the current subblock, RemCcbs is still greater than 4, then encode up to four other abs_level_gtx_flag[n][j] using the context encoding method, where n is the index along the scan order of the current position in the subblock, and j is from 1 to 4. After encoding each abs_level_gtx_flag[n][j], the value of RemCcbs is decreased by one. If RemCcbs is less than or equal to 4, then encode the syntax element abs_remainder using the bypass method, if necessary, for the current position in the subblock. For positions whose absolute levels have been encoded using the syntax element abs_remainder entirely by the bypass method, also encode coeff_sign_flag using the bypass method. In summary, to limit the total number of context encoding bins and guarantee worst-case throughput, there is a predefined counter remBinsPassl in the RRC or RemCcbs in the TSRC.

[0052] On the derivation of Rice parameters In the current RRC design of VVC, two syntax elements abs_remainder and dec_abs_level, coded in the bypass bin, may be present in the remaining level bitstream. We binarize abs_remainder and dec_abs_level by a combination of truncated Rice (TR) and restricted exponential-golomb (EGk) binarization processes specified in the VVC specification, which require binarizing a given level with the Rice parameters. To obtain the optimal Rice parameters, we adopt a local summation method as described below.

[0053] The array AbsLevel[xC][yC] represents an array of absolute values ​​of the transform coefficient levels of the current transform block with color component index cIdx. Given an array AbsLevel[x][y] with a transform block with color component index cIdx and luminance position (x0, y0) in the upper left corner, derive the local sum variable locSumAbs by the method specified in the following dummy code.

[0054] locSumAbs = 0 If ( xC < ( 1 << log2TbWidth ) - 1 ), then locSumAbs += AbsLevel[ xC + 1 ][ yC ] If ( xC < ( 1 << log2TbWidth ) - 2 ), then locSumAbs += AbsLevel[ xC + 2 ][ yC ] If ( yC < ( 1 << log2TbHeight ) - 1 ) , then locSumAbs += AbsLevel[ xC + 1 ][ yC + 1 ] } If ( yC < ( 1 << log2TbHeight ) - 1 ), then { locSumAbs += AbsLevel[ xC ][ yC + 1 ] If ( yC < ( 1 << log2TbHeight ) - 2 ) , then locSumAbs += AbsLevel[ xC ][ yC + 2 ] } locSumAbs = Clip3( 0, 31, locSumAbs - baseLevel * 5 )

[0055] where log2TbWidth and log2TbHeight are the base 2 logarithms of the width and height of the transform block, respectively. For abs_remainder and dec_abs_level, the variables baseLevel are 4 and 0, respectively. Given the local sum variable locSumAbs, we derive the Rice parameter cRiceParam according to the scheme specified in Table 2. [Table 2]

[0056] On History-Based Derivation of Rice Parameters If a coefficient is located on the boundary of a TU or is decoded using the Rice method first, the template computation used for Rice parameter derivation may cause inaccurate estimation of the coefficient. For these coefficients, some templates may be located outside the TU and may be interpreted or initialized as 0, so the template computation is biased towards 0. FIG. 6 shows an example of a template pattern for calculating locSumAbs of a coefficient located near the boundary of a TU. FIG. 6 shows a CTU 602 divided into multiple CUs, each CU including multiple TUs. For a TU 604, the position of the current coefficient is indicated by a black block, and the position of the neighboring samples in the template pattern is indicated by a patterned block. The patterned block indicates a predetermined area of ​​the current coefficient for calculating the local sum variable locSumAbs.

[0057] In FIG. 6, because the current coefficient 606 is close to the boundary of the TU 604, some neighboring samples (e.g., neighboring samples 608B and 608E) of the current coefficient 606 in the template pattern are located outside the TU boundary. In the above Rice parameter derivation, when calculating the local summation variable locSumAbs, these neighboring samples located outside the boundary are set to 0, which makes the Rice parameter derivation inaccurate. For high bit-depth samples (e.g., more than 10 bits), the neighboring samples located outside the TU boundary may be a large number. Setting these large numbers to 0 will cause more errors in the Rice parameter derivation.

[0058] To improve the accuracy of the Ricean estimation based on computational templates, we propose to update the local sum variable locSumAbs using a history derived value, rather than initializing it to 0, for template positions located outside the current TU. In the following, we explain the implementation of the method by means of an excerpt of the VVC specification text from clause 9.3.3.2, where the proposed text is underlined.

[0059] To maintain the history of neighboring coefficient / sample values, a history counter StatCoeff[cIdx] for each color component is utilized, where cIdx=0, 1, 2, respectively representing the three color components Y, U, and V. If the CTU is the first CTU in a partition (e.g., a picture, slice, or tile), StatCoeff[cIdx] is initialized as follows: StatCoeff[idx]=2*Floor(Log2(BitDepth-10) (1)

[0060] where BitDepth specifies the bit depth of the samples in the video luma and chroma arrays, Floor(x) denotes the largest integer less than or equal to x, and Log2(x) is the base 2 logarithm of x. Before TU decoding and history counter update, the replacement variable HistValue is initialized as follows: HistValue[cIdx]=1< <StatCoeff[cIdx] (2)

[0061] The replacement variable HistValue is used as an estimate for neighboring samples that lie outside the TU boundary (e.g., neighboring samples have horizontal or vertical coordinates that lie outside the TU). The local sum variable locSumAbs is again derived according to the scheme specified by the following dummy code, where the changes are underlined: locSumAbs = 0 If ( xC < ( 1 << log2TbWidth ) - 1 ), then locSumAbs += AbsLevel[ xC + 1 ][ yC ] If ( xC < ( 1 << log2TbWidth ) - 2 ), then locSumAbs += AbsLevel[ xC + 2 ][ yC ] If not locSumAbs += HistValue If ( yC < ( 1 << log2TbHeight ) - 1 ), then locSumAbs += AbsLevel[ xC + 1 ][ yC + 1 ] If not locSumAbs += HistValue } If not locSumAbs += 2 * HistValue If ( yC < ( 1 << log2TbHeight ) - 1 ), then locSumAbs += AbsLevel[ xC ][ yC + 1 ] If ( yC < ( 1 << log2TbHeight ) - 2 ), then locSumAbs += AbsLevel[ xC ][ yC + 2 ] If not locSumAbs += HistValue} If not locSumAbs += HistValue

[0062] The history counter StatCoeff is updated once per TU from the first non-zero Golomb-Rice coded transform coefficient (abs_remainder[cIdx] or dec_abs_level[cIdx]) by an exponential moving average process. When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as abs_remainder, the history counter StatCoeff of color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (3)

[0063] When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as dec_abs_level, the history counter StatCoeff of the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1 (4)

[0064] The updated StatCoeff may be used to calculate the replacement variable HistValue of the next TU according to equation (2) before decoding the next TU.

[0065] About Wavefront Parallel Processing (WPP) WPP is designed to provide a parallel coding algorithm. When WPP is enabled in VVC, each CTU row of a frame or tile or slice constitutes a single partition. WPP is enabled / disabled by the SPS element sps_entropy_coding_sync_enabled_flag. Figure 7 shows an example of a tile with WPP enabled. In Figure 7, each CTU row of a tile is processed with a delay of one CTU with respect to its previous CTU row. In this way, we do not break the dependencies between consecutive CTU rows at partition boundaries, except for the CABAC context variables and palette predictor, when palette coding is enabled at the end of each CTU row. To mitigate potential loss of coding efficiency, the self-adapted CABAC context variables and palette predictor contents are propagated from the first coding CTU of the previous CTU row to the first CTU of the current CTU row. WPP does not change the normal raster scan order of CTUs.

[0066] When WPP is enabled, multiple threads (the number of which can be up to the number of CTU rows in a partition (e.g., tile, slice, or frame)) work in parallel to process each CTU row. By using WPP in the decoder, each decoding thread processes one CTU row of the partition. Thread processing must be scheduled such that for each CTU, the decoding of the upper adjacent CTU in the previous CTU row must be completed. An additional small overhead of WPP is added to store all CABAC context variables and palette predictor contents after completing the encoding of the first CTU in each CTU row except the last CTU row.

[0067] When the history-based Rice parameter derivation discussed above is enabled for high bit depth and high bit rate video coding, the last StatCoeff in the previous CTU row is passed to the first TU in the current CTU row. Thus, if WPP is enabled at the same time, this process will interfere with WPP and destroy the parallelism of WPP. In this disclosure, some solutions are proposed to solve this problem when parallel coding (e.g., WPP) is enabled.

[0068] In one embodiment, the interference of history-based Rice parameter derivation for parallel encoding is eliminated by removing the dependency between CTUs of different CTU rows when calculating the history counter StatCoeff. In the embodiment, instead of using the history counter StatCoeff value obtained from the previous CTU row, an initial value of StatCoeff[cIdx] is used to encode the first abs_remainder[cIdx] or dec_abs_level[cIdx] in each CTU row of a partition (e.g., a frame or a tile or a slice), where cIdx is the index of a color component.

[0069] As an example, the initial value of StatCoeff[cIdx] can be determined as follows: StatCoeff[idx]=2*Floor(Log2(BitDepth-10)) (5)

[0070] where BitDepth specifies the bit depth of the samples in the luma or chroma array, and Floor(x) represents the largest integer less than or equal to x. As another example, the initial value of StatCoeff[cIdx] can be determined as follows: StatCoeff[idx]=Clip(MIN_Stat,MAX_Stat,(int)((19-QP) / 6))-l (6)

[0071] Here, MIN_Stat, MAX_Stat are two predefined integers, QP is the initial QP of each slice, and Clip() is an operation defined as follows: JPEG2024532697000004.jpg30170

[0072] Before encoding the first TU of each CTU row of a partition (eg, a frame, a tile, or a slice), the replacement variable HistValue is calculated as follows: HistValue[cIdx]=1< <StatCoeff[cIdx] (8)

[0073] HistValue can be used to calculate the above local sum variable locSumAbs. HistValue can be updated once per TU from the first non-zero Golomb-Rice coded transform coefficient (abs_remainder[cIdx] or dec_abs_level[cIdx]) by an exponential moving average process. When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as abs_remainder, the history counter StatCoeff[cIdx] of color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (9)

[0074] When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as dec_abs_level, the history counter StatCoeff[cIdx] of the color component cIdx is updated as follows: StatCoeff[cIdx] = (StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1 (10)

[0075] The updated StatCoeff[cIdx] is used to calculate the substitution variable HistValue for the next TU of the current CTU or the first TU of the next CTU in the current CTU row, as shown in equation (8).

[0076] 8 shows an example of a frame 802 and the CTUs included in the frame. In this example, frame 802 includes two tiles, tile 804A and tile 804B. Tile 804A includes four CTU rows, CTU row 1 to CTU row 4. The first CTU row includes CTU0 to CTU9, and the second CTU row includes CTU10 to CTU19. Similarly, tile 804B includes four CTU rows, CTU row 1' to CTU row 4'. The first CTU row includes ten CTUs, CTU0' to CTU9', and the second CTU row includes CTU10' to CTU19', and so on.

[0077] According to this embodiment, the initial value of StatCoeff[cIdx] of the tile 804A can be determined according to equation (5) or (6). Before encoding the first TU of each CTU row in CTU row 1 to CTU row 4, the substitution variable HistValue[cIdx] is calculated using the initial value of StatCoeff[cIdx] using equation (8). For example, before encoding the first TU of CTU0, the variable HistValue is calculated using equation (8). The value of HistValue is used to determine the local sum variable locSumAbs of the coefficients of the first TU, and the local sum variable locSumAbs is further used to determine the Rice parameter of each coefficient of the first TU. When processing the first TU of the current CTU0, the history counter StatCoeff can be updated according to equation (9) or (10). Before processing the second TU in CTU0, the current value of StatCoeff is used to determine the HistValue of the second TU according to equation (8). Then, a similar process is adopted for the second TU, where HistValue is used to determine the Rice parameter and StatCoeff is updated. For the first TU in CTU1, HistValue is calculated using the latest StatCoeff from the TU in CTU0 according to equation (8). The process can be repeated until the last CTU in the current CTU row 1 (i.e., CTU9) is processed.

[0078] For the second CTU row of tile 804A, before encoding the first TU of CTU10 (i.e., the first CTU of the second CTU row), the history counter StatCoeff is initialized according to equation (5) or (6). For the TUs in the CTUs in the second CTU row, a process similar to that described above for CTU row 1 is performed. Similarly, before encoding the first TU in each of CTU20 and CTU30, the variable StatCoeff is again initialized according to equation (5) or (6).

[0079] Tile 804B can be processed in a similar manner. Before encoding each of the first TUs in CTU row l' to CTU row 4' (i.e., CTU0', CTU10', CTU20', and CTU30'), initialize the value of StatCoeff[cIdx] according to equation (5) or (6), and calculate the history counter HistValue using equation (8). The calculated history counter HistValue is used to calculate the locSumAbs and Rice parameters of the TUs in the first CTU and the remaining CTUs in each CTU row. Furthermore, the history counter StatCoeff may be updated at most once per TU according to equation (9) or (10), and the updated value of StatCoeff is used to determine the HisValue of the next TU in the same CTU row.

[0080] 8 is illustrated as a frame 802 including two tiles 804A and 804B, the same process applies to other scenarios such as a slice including multiple tiles, a frame including multiple slices, etc. In any of these scenarios, before encoding the first TU in each CTU row of a partition (e.g., frame, tile or slice), the value of the history counter StatCoeff[cIdx] is reset to an initial value to eliminate the dependency of the CTU row in the Rice parameter derivation.

[0081] The possible specification changes of the VVC are underlined and are specified below. [Table 3]

[0082] Another possible specification change for VVC relative to 9.3.2.1 is to provide the following: [Table 4]

[0083] Video sample bit depth The bit depth of the input video supported by VVC version 2 can exceed 10 bits. Higher video bit depths can provide lower compression distortion and higher visual quality for the decoded video. To support higher bit depths of the input video, the semantics of the corresponding Sequence Parameter Set (SPS) syntax element sps_bitdepth_minus8 and Video Parameter Set (VPS) syntax element vps_ols_dpb_bitdepth_minus8[i] can be modified as follows:

[0084] sps_bitdepth_minus8 specifies the bit depth of the samples in the luma and chroma arrays, BitDepth, and the value of the luma and chroma quantization parameter range offset, QpBdOffset, as follows: BitDepth = 8 + sps_bitdepth_minus8 (x1) QpBdOffset = 6 * sps_bitdepth_minus8 (x2)

[0085] sps_bitdepth_minus8 must be in the range 0 to 8 (including the endpoints).

[0086] If sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer contained in the i-th (i in the range of 0 to NumMultiLayerOlss - 1 (endpoints included)) multi-layer OLS specified by the VPS, the bitstream consistency requirement is that the value of sps_bitdepth_minus8 must be less than or equal to the value of vps_ols_dpb_bitdepth_minus8[i].

[0087] vps_ols_dpb_bitdepth_minus8[i] specifies the maximum allowed value of sps_bitdepth_minus8 for all SPSs referenced by CLVSs in the CVS of the i-th multi-layer OLS. The value of vps_ols_dpb_bitdepth_minus8[i] must be in the range 0 to 8 (inclusive).

[0088] NOTE 2 - To decode the i-th multi-layer OLS, a decoder can safely allocate memory in the DPB based on the values ​​of the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i] and vps_ols_dpb_bitdepth_minus8[i].

[0089] As can be seen above, the bit depth BitDepth of the samples of the luma and chroma arrays can be derived based on the SPS syntax element sps_bitdepth_minus8 according to formula (x1). The determined BitDepth value can be used to derive the history counter StatCoeff, the substitution variable HistValue and the Rice parameter as described above.

[0090] The VPS syntax element vps_ols_dpb_bitdepth_minus8[i] can be used to derive the size of the decoded picture buffer (DPB). There can be multiple video layers in the encoded bitstream. The video parameter set is used to specify the corresponding syntax elements. For video decoding, the DPB can be used to store reference pictures so that previously encoded pictures are used to generate prediction signals to use when encoding other pictures. The DPB can further reorder the decoded pictures so that they can be output and / or displayed in the correct order. The DPB can also be used for output delays specified by the hypothetical reference decoder. The decoded pictures are persistently stored for a predefined period specified in the DPB for the hypothetical reference decoder and are output after the predefined period has elapsed.

[0091] To safely allocate memory for a DPB, the size of the DPB is determined by the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i] and vps_ols_dpb_bitdepth_minus8[i] as follows: picture_size1(in bits) = vps_ols_dpb_pic_width[i] * vps_ols_dpb_pic_height[i] * (vps_ols_dpb_bitdepth_minus8[i]+8) (vps_ols_dpb_chroma_format[ i ]== 0) / / If monochrome, picture_size = picture_size1; else if (vps_ols_dpb_chroma_format[ i ]== 1) / / 4:2:0, picture_size = 1.5 * picture_size1; Otherwise, if (vps_ols_dpb_chroma_format[ i ]== 2 / / 4:2:2, picture_size = 2 * picture_size1; Otherwise, if (vps_ols_dpb_chroma_format[ i ]== 3 / / 4:4:4, picture_size = 3 * picture_size1;

[0092] Therefore, the size of the DPB is determined by picture_size. In other words, the size of the DPB can be determined according to the chroma format of the sample. If the video frame is a monochrome frame, determine that the size of the frame waiting to be buffered is the base picture size picture_size1. If the color subsampling of the color video frame is 4:2:0, determine that the size of the frame is 1.5 times the base picture size picture_sizel, if the color subsampling of the color video frame is 4:2:2, determine that the size of the frame is 2 times the base picture size picture_sizel, if the color subsampling of the color video frame is 4:4:4, determine that the size of the frame is 3 times the base picture size picture_sizel. According to the color subsampling, the size of the DPB can be determined to be the number of frames stored in the DPB multiplied by the size of the frame.

[0093] 9 illustrates an example of a process 900 for encoding a partition of a video, according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video encoder 100) perform the operations illustrated in FIG. 9 by executing appropriate program code (e.g., program code implementing the entropy encoding module 116). For illustrative purposes, the process 900 is described with reference to some examples shown in the figure. However, other implementations are possible.

[0094] At block 902, the process 900 includes accessing a partition of a video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in a CTU row as shown in Figure 8. As shown in Figure 6, each CTU includes one or more CTUs, and each CTU includes multiple TUs for encoding.

[0095] At block 904, the process 900 includes processing each CTU of the set of CTUs in the partition to encode the partition into bits, and the block 904 includes blocks 906-914. At block 906, the process 900 includes determining whether a parallel encoding algorithm is enabled and whether the current CTU is the first CTU in the row of CTUs. In some examples, the parallel encoding is indicated by a flag, where a value of the flag of 0 indicates that the parallel encoding is disabled and a value of the flag of 1 indicates that the parallel encoding is enabled. If the parallel encoding algorithm is enabled and the current CTU is the first CTU in the row of CTUs, the process 900 includes setting a history counter StatCoeff to an initial value at block 908. As described above, if the history-based Rice parameter derivation is enabled, the initial value of the history counter may be set according to equation (5) or (6), and if not, the initial value of the history counter is set to 0.

[0096] If it is determined that the parallel encoding algorithm is not enabled or the current CTU is not the first CTU in the row of CTUs, or after setting the history counter in block 908, the process 900 includes calculating Rice parameters of the TUs in the CTU based on the history counter in block 910. As described in detail above with respect to Figures 6-8, if the history counter is reset in block 908, the Rice parameters of the TUs in the CTU are calculated based on the reset history counter or a subsequently updated history counter. If the history counter is not reset in block 908, the Rice parameters of the TUs in the CTU are calculated based on the history counter updated in the immediately preceding CTU or a subsequently updated history counter in the current CTU.

[0097] At block 912, the process 900 includes encoding the TUs in the CTU into a binary representation based on the calculated Rice parameters, for example, by a combination of truncated Rice (TR) and restricted k-th order EGk as defined in the VVC specification. At block 914, the process 900 encodes the binary representations of the CTUs into bits for inclusion in a video bitstream, for example, using the context self-adaptive binary arithmetic coding (CABAC) discussed above. At block 916, the process 900 includes outputting the encoded video bitstream.

[0098] FIG. 10 illustrates an example of a process 1000 for decoding a partition of a video, according to some embodiments of the present disclosure. One or more computing devices may implement the operations illustrated in FIG. 10 by executing appropriate program code. For example, a computing device implementing the video decoder 200 may implement the operations illustrated in FIG. 10 by executing program code for the entropy decoding module 216, the inverse quantization module 218, and the inverse transform module 219. For illustrative purposes, the process 1000 is described with reference to some examples shown in the figure. However, other implementations are possible.

[0099] At block 1002, the process 1000 includes accessing a binary string or binary representation representing a partition of a video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in a CTU row as shown in FIG. 8. As shown in FIG. 6, each CTU includes one or more CTUs, and each CTU includes multiple TUs for encoding.

[0100] In block 1004, the process 1000 includes processing the binary string of each CTU in the set of CTUs in the partition to generate a decoded sample for the partition, and the block 1004 includes 1006 to 1014. In block 1006, the process 1000 includes determining whether a parallel encoding algorithm is enabled and whether the current CTU is the first CTU in the row of CTUs. The parallel encoding is indicated by a flag, where a value of the flag of 0 indicates that the parallel encoding is disabled, and a value of the flag of 1 indicates that the parallel encoding is enabled. If the parallel encoding algorithm is enabled and the current CTU is the first CTU in the row of CTUs, the process 1000 includes setting a history counter StatCoeff to an initial value in block 1008. As described above, if the history-based Rice parameter derivation is enabled, the initial value of the history counter may be set according to equation (5) or (6), and if not, the initial value of the history counter is set to 0.

[0101] If it is determined that the parallel encoding algorithm is not enabled or the current CTU is not the first CTU in the CTU row, or after setting the history counter in block 1008, process 1000 includes calculating Rice parameters of the TUs in the CTU based on the history counter in block 1010. As described in detail above with respect to Figures 6-8, if the history counter is reset in block 1008, the Rice parameters of the TUs in the CTU are calculated based on the reset history counter or a subsequently updated history counter. If the history counter is not reset in block 1008, the Rice parameters of the TUs in the CTU are calculated based on the history counter updated in the immediately preceding CTU or a subsequently updated history counter in the current CTU.

[0102] At block 1012, process 1000 includes decoding binary strings or binary representations of the TUs in the CTU into coefficient values ​​based on the calculated Rice parameters, for example, by a combination of Truncated Rice (TR) and Restricted k-th order EGK as defined in the VVC specification. At block 1014, process 1000 includes reconstructing pixel values ​​of the TUs in the CTU, for example, by inverse quantization and inverse transform as discussed above with respect to FIG. 2. At block 1016, process 1000 includes outputting the decoded partition of the video.

[0103] In another embodiment, the dependency between CTUs when calculating the history counter StatCoeff is consistent with the dependency between CTUs in a parallel encoding algorithm (e.g., WPP). For example, the history counter StatCoeff of a CTU row of a partition (e.g., a frame, a tile, or a slice) can be calculated based on the coefficient values ​​of the first N or fewer CTUs in the previous CTU row, where N is the maximum delay between two consecutive CTU rows allowed in the parallel encoding algorithm. In this way, the dependency between CTUs of two consecutive CTU rows when calculating the history counter StatCoeff is limited to be smaller than (and thus consistent with) the dependency between CTUs when performing parallel processing.

[0104] This embodiment can be realized using a storage synchronization process. For example, in the above WPP, the delay between two consecutive CTU rows is 1 CTU, so N=1. In the storage process, after encoding the last TU of the first CTU of each CTU row (except the last CTU row), StatCoeff[cIdx] can be stored in a storage variable StatCoeffWpp[cIdx], and for each CTU row other than the first CTU row, a synchronization process of Rice parameter derivation is applied before encoding the first TU. In the synchronization process, StatCoeff[cIdx] is synchronized with the StatCoeffWpp[cIdx] stored from the previous CTU row.

[0105] As described above, before encoding the first TU in each CTU row, the variable HistValue is calculated as follows: HistValue[cIdx] = 1 << StatCoeff[cIdx] (11)

[0106] If the current CTU row is the first CTU row of a partition, StatCoeff[cIdx] can be initialized according to equation (5) or (6). The calculated HistValue is used to determine the local sum variable locSumAbs, which is further used to determine the Rice parameters of the TUs in the current CTU. StatCoeff is updated once per TU from the first non-zero Golomb-Rice coding transform coefficient (abs_remainder[cIdx] or dec_abs_level[cIdx]) by the exponential moving average process as described above with respect to equations (9) and (10).

[0107] After encoding the last TU of the first CTU in the first CTU row, in the storage step, StatCoeff[cIdx] can be saved as StatCoeffWpp[cIdx] as follows: StatCoeffWpp[cIdx] = StatCoeff[cIdx] (12)

[0108] The encoding of the remaining CTUs in the first CTU row may be performed in a manner similar to that described with respect to the first embodiment.

[0109] Before encoding the second CTU row and the first TU in any subsequent CTU rows, StatCoeff[cIdx] can be obtained by a synchronization step. StatCoeff[cIdx] = StatCoeffWpp[cIdx] (13)

[0110] The obtained StatCoeff[cIdx] value is used to calculate HistValue according to equation (11). The remaining process of the CTU row is the same as that of the first CTU row.

[0111] Possible changes to the VVC specification are specified below (changes are underlined): [Table 5] JPEG2024532697000008.jpg174170

[0112] An alternative history-based derivation of the Rice parameters History-based Rice parameter derivation can be implemented in an alternative manner, in which if the CTU is the first CTU in a partition (e.g., a picture, slice, or tile), the initial value of StatCoeff[cIdx] is used to initialize HistValue as follows: HistValue=sps_persistent_Rice_adaptation_enabled_flag?1< <StatCoeff[cIdx]:0 (14)

[0113] The initial HistValue is used to code the first abs_remainder[cIdx] or dec_abs_level[cIdx] until the HistValue is updated according to the following rule: When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as abs_remainder, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_remainder[cIdx]))+2)>>1 (15)

[0114] When the first non-zero Golomb-Rice coded transform coefficient in a TU is coded as dec_abs_level, the history counter of the color component cIdx is updated as follows: StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1(16)

[0115] When the history counter StatCoeff[cIdx] is updated, we update HistValue as shown in equation (17), and the updated HistValue is used to derive the Rice parameters of the remaining syntax elements abs_remainder and dec_abs_level until the new StatCoeff[cIdx] and HistValue[cIdx] are updated again. HistValue[cIdx] = 1 << StatCoeff[cIdx] (17)

[0116] Based on the current VVC specifications, possible specification changes are as specified below.

[0117] Clause 7.3.11.11 (Residual Coding Syntax) is modified to read as follows (additions are underlined): [Table 6] JPEG2024532697000010.jpg159170

[0118] To resolve dependency conflicts between parallel encoding and the alternative history-based Rice parameter derivation, we save StatCoeff[cIdx] and HistValue[cIdx] for each color component after encoding the last TU of the first CTU in each CTU row. The saved StatCoeff[cIdx] and HistValue[cIdx] values ​​are used to initialize StatCoeff[cIdx] and HistValue[cIdx] before processing the first TU of the first CTU in the subsequent CTU row.

[0119] The embodiment can further be implemented using a storage synchronization process. For example, in the storage process, after processing the last TU of the first CTU of each CTU row, StatCoeff[cIdx] and HistValue[cIdx] can be saved into storage variables, for example, StatCoeffWpp[cIdx] and HistValueWpp[cIdx] as shown in formulas (18) and (19), respectively. StatCoeffWpp[cIdx] = StatCoeff[cIdx] (18) HistValueWpp[cIdx] = HistValue[cIdx] (19)

[0120] In each CTU row except the first CTU row, before encoding the first TU, a synchronization process of Rice parameter derivation is applied. For example, as shown in equations (20) and (21), StatCoeff[cIdx] and HistValue[cIdx] are synchronized with StatCoeffWpp[cIdx] and HistValueWpp[cIdx] saved from the previous CTU row, respectively. StatCoeff[cIdx] = StatCoeffWpp[cIdx] (20) HistValue[cIdx] = HistValueWpp[cIdx] (21)

[0121] The synchronization variable HistValue is used to encode the first abs_remainder[cIdx] or dec_abs_level[cIdx] until HistValue is updated.

[0122] As described above, StatCoeff[cIdx] can be updated once per TU from the first non-zero Golomb-Rice encoded transform coefficient (abs_remainder[cIdx] or dec_abs_level[cIdx]) as shown in equation (15) or (16). When the history counter StatCoeff[cIdx] is updated, HistValue is updated according to equation (17), and the updated HistValue is used to derive the Rice parameters of the remaining syntax elements abs_remainder and dec_abs_level until new StatCoeff[cIdx] and HistValue are updated again.

[0123] Based on the current VVC specifications, possible specification changes, as underlined, are as follows: [Table 7] JPEG2024532697000012.jpg213170

[0124] 11 illustrates an example of a process 1100 for encoding a partition of a video, according to some embodiments of the present disclosure. One or more computing devices (e.g., computing devices implementing the video encoder 100) perform the operations illustrated in FIG. 11 by executing appropriate program code (e.g., program code implementing the entropy encoding module 116). For illustrative purposes, the process 1100 is described with reference to some examples shown in the figure. However, other implementations are possible.

[0125] At block 1102, the process 1100 includes accessing a partition of a video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in a CTU row as shown in Figure 8. As shown in Figure 6, each CTU includes one or more CTUs, and each CTU includes multiple TUs for encoding.

[0126] In block 1104, the process 1100 includes processing each CTU of the set of CTUs in the partition to encode the partition into bits, and the block 1104 includes blocks 1106-1118. In block 1106, the process 1100 includes determining whether a parallel encoding algorithm is enabled and whether the current CTU is the first CTU of the CTU row. In some examples, the parallel encoding is indicated by a flag, where a value of the flag of 0 indicates that the parallel encoding is disabled and a value of the flag of 1 indicates that the parallel encoding is enabled. If the parallel encoding algorithm is enabled and the current CTU is the first CTU of the CTU row, the process 1100 includes determining whether the current CTU row is the first CTU row in the partition in block 1107. If so, the process 1100 includes setting a history counter StatCoeff to an initial value in block 1108. As described above, the initial value of the history counter can be set according to equation (5) or (6). If the current CTU row is not the first CTU row in the partition, process 1100 sets the history counter StatCoeff to the value stored in the history counter storage variable as shown in equation (13) or (20) at block 1109. In some examples, for example, when an alternative Rice parameter derivation is utilized, the value of the substitution variable HistValue may be reset to the stored value as shown in equation (21).

[0127] If it is determined that the parallel encoding algorithm is not enabled or the current CTU is not the first CTU in the CTU row, or after setting the value of the history counter at block 1108 or 1109, process 1100 includes calculating the Rice parameters of the TUs in the CTU based on the history counter (and the replacement variable HistValue if HistValue was reset) at block 1110. As described above (see, e.g., FIG. 8, or alternative Rice parameter derivations), if the value of the history counter is reset at block 1108 or 1109, the Rice parameters of the TUs in the CTU are calculated based on the reset history counter or a subsequently updated history counter. If the history counter is not reset at block 1108 or 1109, the Rice parameters of the TUs in the CTU are calculated based on the history counter updated in the immediately preceding CTU or a subsequently updated history counter in the current CTU.

[0128] At block 1112, the process 1100 includes encoding the TUs in the CTU into a binary representation based on the calculated Rice parameters, for example, by a combination of truncated Rice (TR) and restricted k-th order EGK as defined in the VVC specification. At block 1114, the process 1100 encodes the binary representations of the CTUs into bits for inclusion in a video bitstream, for example, using the context self-adaptive binary arithmetic coding (CABAC) discussed above.

[0129] At block 1116, the process 1100 includes determining whether parallel encoding is enabled and the CTU is the first CTU in the current CTU row. If so, the process 1100 includes storing the value of the history counter in a history counter storage variable as shown in equation (12) or (18) at block 1118. In some examples, for example, when an alternative Rice parameter derivation is utilized, the value of the substitution variable HistValue may also be stored in a storage variable as shown in equation (19). At block 1120, the process 1100 includes outputting the encoded video bitstream.

[0130] In some scenarios, a CTU in a non-first CTU row may be located on a partition boundary. For example, the first CTU in the second CTU row has no CTU in the partition that is located above the CTU. In these scenarios, the history counter of the CTU can be set to an initial value rather than a stored value. In this case, a new block 1107' can be added between blocks 1107 and 1109 in FIG. 11 to determine whether the CTU is located on a partition boundary (e.g., the CTU has no upper adjacent CTU in the partition). If so, process 1100 proceeds to block 1108 to set the history counter to the initial value, otherwise process 1100 proceeds to block 1109 to set the history counter to the stored value. The remaining blocks in FIG. 11 remain unchanged.

[0131] FIG. 12 illustrates an example of a process 1200 for decoding a partition of a video, according to some embodiments of the present disclosure. One or more computing devices may implement the operations illustrated in FIG. 12 by executing appropriate program code. For example, a computing device implementing a video decoder 200 may implement the operations illustrated in FIG. 12 by executing program code for an entropy decoding module 216, an inverse quantization module 218, and an inverse transform module 219. For illustrative purposes, the process 1200 is described with reference to some examples shown in the figure. However, other implementations are possible.

[0132] At block 1202, the process 1200 includes accessing a binary string or binary representation representing a partition of a video signal. The partition may be a video frame, slice, or tile, or any type of partition that is processed as a unit by the video encoder when performing encoding. The partition includes a set of CTUs arranged in a CTU row as shown in FIG. 8. As shown in FIG. 6, each CTU includes one or more CTUs, and each CTU includes multiple TUs for encoding.

[0133] In block 1204, the process 1200 includes processing the binary string of each CTU in the set of CTUs in the partition to generate a decoded sample of the partition, and the block 1204 includes 1206 to 1218. In block 1206, the process 1200 includes determining whether the parallel encoding algorithm is enabled and the current CTU is the first CTU in the CTU row. The parallel encoding is indicated by a flag, where a value of the flag of 0 indicates that the parallel encoding is disabled, and a value of the flag of 1 indicates that the parallel encoding is enabled. If the parallel encoding algorithm is enabled and the current CTU is determined to be the first CTU in the CTU row, the process 1200 includes determining whether the current CTU row is the first CTU row in the partition, in block 1207. If so, the process 1200 includes setting a history counter StatCoeff to an initial value, in block 1208. As described above, the initial value of the history counter can be set according to equation (5) or (6). If the current CTU row is not the first CTU row in the partition, process 1200 sets the history counter StatCoeff to the value stored in the history counter storage variable as shown in equation (13) or (20) at block 1209. In some examples, for example, when an alternative Rice parameter derivation is utilized, the value of the substitution variable HistValue may be reset to the stored value as shown in equation (21).

[0134] If it is determined that the parallel encoding algorithm is not enabled or the current CTU is not the first CTU in the CTU row, or after setting the history counter at block 1208 or 1209, process 1200 includes calculating the Rice parameters of the TUs in the CTU based on the history counter (and the substitution variable HistValue if HistValue is also set) at block 1210. As described above (e.g., with respect to FIG. 8, in the alternative Rice parameter derivation), if the value of the history counter is reset at block 1208 or 1209, the Rice parameters of the TUs in the CTU are calculated based on the reset history counter or a subsequently updated history counter. If the history counter is not reset at block 1208 or 1209, the Rice parameters of the TUs in the CTU are calculated based on the history counter updated in the immediately preceding CTU or a subsequently updated history counter in the current CTU.

[0135] At block 1212, the process 1200 includes decoding the binary strings or binary representations of the TUs in the CTU into coefficient values ​​based on the calculated Rice parameters, for example, by a combination of Truncated Rice (TR) and Restricted k-th order EGK as defined in the VVC specification. At block 1214, the process 1200 includes reconstructing pixel values ​​of the TUs in the CTU, for example, by inverse quantization and inverse transform as discussed above with respect to FIG.

[0136] At block 1216, the process 1200 includes determining whether parallel encoding is enabled and the CTU is the first CTU in the current CTU row. If so, the process 1200 includes storing the value of the history counter in a history counter storage variable as shown in equation (12) or (18) at block 1218. In some examples, for example, if an alternative Rice parameter derivation is utilized, the value of the substitution variable HistValue may also be stored in a storage variable as shown in equation (19). At block 1220, the process 1200 includes outputting the decoded partition of the video.

[0137] In another embodiment, WPP or other parallel encoding algorithms are prevented from coexisting in the bitstream with history-based Rice parameter derivation. For example, if WPP is enabled, history-based Rice parameter derivation may not be enabled. If WPP is not enabled, history-based Rice parameter derivation may be enabled. Similarly, if history-based Rice parameter derivation is enabled, WPP may not be enabled. As an example, syntax changes may be made as follows:

[0138] 7.3.2.22 Sequence Parameter Set Range Extension Syntax (additions are underlined) [Table 8]

[0139] As another example, change the corresponding semantics to the following (changes are underlined): [Table 9] JPEG2024532697000015.jpg83170

[0140] In the above description, TU is illustrated and described in the figures (e.g., FIG. 6), but the same technique can be applied to transform block (TB). In other words, in the embodiments presented above (including the figures), TU can also represent TB.

[0141] An example of a computing system for implementing dependent quantization for video coding Any suitable computing system may be used to perform the operations described herein. For example, FIG. 13 illustrates an example of a computing device 1300 that may implement the video encoder 100 of FIG. 1 or the video decoder 200 of FIG. 2. In some embodiments, the computing device 1300 may include a processor 1312 that is communicatively coupled to a memory 1314 and executes computer-executable program code and / or accesses information stored in the memory 1314. The processor 1312 may include a microprocessor, an application specific integrated circuit (ASIC), a state machine, or other processing device. The processor 1312 may include any one of a plurality of processing devices (including one processing device). Such a processor may include or be in communication with a computer readable medium that stores instructions, and when executed by the processor 1312, the instructions cause the processor to perform the operations described herein.

[0142] The memory 1314 may include any suitable non-transitory computer-readable medium. The computer-readable medium may include any electronic, optical, magnetic or other memory device capable of providing computer-readable instructions or other program code to a processor. Non-limiting examples of computer-readable media include magnetic disks, storage chips, ROM, RAM, ASICs, configured processors, optical memory, magnetic tape or other magnetic memory, or any other medium from which a computer processor can read instructions. The instructions include processor-specific instructions generated by a compiler and / or interpreter based on code written in any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, and ActionScript.

[0143] The computing device 1300 further comprises a bus 1316. The bus 1316 communicatively couples one or more components of the computing device 1300. The computing device 1300 further comprises a number of external or internal devices, such as input or output devices, e.g., the computing device 1300 is shown to comprise an input / output (I / O) interface 1318, which can receive input from one or more input devices 1320 or provide output to one or more output devices 1322. The one or more input devices 1320 and the one or more output devices 1322 are communicatively coupled to the I / O interface 1318. The communicative coupling can be achieved in any suitable manner (e.g., connection via a printed circuit board, connection via a cable, communication via wireless transmission, etc.). Non-limiting examples of input devices 1320 include a touchscreen (e.g., one or more cameras for photographing the touch area, or a pressure sensor for detecting pressure changes caused by a touch), a mouse, a keyboard, or any other device for generating input events in response to physical movements of a user of the computing device. Non-limiting examples of output devices 1322 include an LCD screen, an external monitor, speakers, or any other device for displaying or otherwise presenting output generated by the computing device.

[0144] The computing device 1300 may execute program code that configures the processor 1312 to perform one or more of the operations described in Figures 1-12 above. The program code may include the video encoder 100 or the video decoder 200. The program code may reside in the memory 1314 or any suitable computer-readable medium and may be executed by the processor 1312 or any other suitable processor.

[0145] The computing device 1300 may further comprise at least one network interface device 1324. The network interface devices 1324 may include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks 1328. Non-limiting examples of the network interface devices 1324 include Ethernet network adapters, wireless modems, etc. The computing device 1300 may transmit messages as electronic or optical signals through the network interface devices 1324.

[0146] General Considerations Numerous specific details are described herein to provide a thorough understanding of the claimed subject matter. However, as will be understood by those skilled in the art, the claimed subject matter may be practiced without these specific details. In other instances, methods, devices, or systems known to those skilled in the art have not been described in detail so as not to obscure the claimed subject matter.

[0147] Unless otherwise indicated, terms such as "processing," "computing," "calculating," "determining," "identifying," and the like are used throughout this specification to refer to operations or processes of a computing device (e.g., one or more computers or one or more similar electronic computing devices) that manipulate or transform data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage, transmission, or display devices of the computing platform.

[0148] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that produce a result conditional on one or more inputs. Suitable computing devices include computer systems based on general purpose microprocessors that access stored software that programs or configures the computing system, changing it from a general purpose computing device to a special purpose computing device, thereby implementing one or more embodiments of the subject matter herein. Any suitable programming, scripting, or other type or combination of languages ​​can be used to implement the teachings contained herein in software that is configured to program or configure a computing device.

[0149] The method embodiments disclosed herein may be implemented in operation with such a computing device. The order of the blocks shown in the above examples may be changed, e.g., the blocks may be reordered, combined, or divided into sub-blocks. Some blocks or processes may be executed in parallel.

[0150] The use of "applied to" or "configured to" herein means open and inclusive language and does not exclude equipment adapted or configured to perform additional tasks or steps. Additionally, the use of "based on" means open and inclusive in that a process, step, calculation or other action "based on" one or more recited conditions or values ​​may in fact be based on additional conditions or beyond the recited values. Headings, lists and numbering contained herein are for ease of description only and are not intended to be limiting.

[0151] Although the subject matter of this specification has been described in detail with respect to specific examples thereof, it will be understood that those skilled in the art, upon understanding the foregoing, may readily make modifications, variations, and equivalents to such examples. It is therefore to be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter of this specification that may be readily made by those skilled in the art.

Claims

1. 1. A method for decoding a video, comprising: accessing a binary string representing a partition of the video, the partition including a plurality of coding tree units (CTUs) forming one or more CTU rows; For each CTU of the plurality of CTUs in the partition, Before decoding the CTU, determining that the CTU is a first CTU in a current CTU row of the one or more CTU rows in the partition; In response to determining that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of a Rice parameter to an initial value; Decoding the CTU; and performing outputting a decoded partition of the video, the decoded partition including a plurality of decoded CTUs in the partition; Decoding the CTU includes: Calculating the Rice parameters for transform units (TUs) within the CTU based on the history counters; decoding the binary string corresponding to the TU in the CTU into coefficient values ​​of the TU based on the calculated Rice parameter; determining pixel values ​​of the TUs within the CTU based on the coefficient values; Video decoding method.

2. The partition is a frame, a slice, or a tile.

10. The video decoding method of claim 1.

3.

4. Calculating a Rice parameter for a TU within the CTU based on the history counter includes: determining a substitution variable HistValue based on the history counter; Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU using values ​​of neighboring coefficients in a predetermined region of the coefficient and the substitution variable HistValue; deriving the Rice parameter for the TU based on the local sum variable locSumAbs.

10. The video decoding method of claim 1.

5. determining a substitution variable HistValue based on the history counter includes determining a substitution variable HistValue for a color component cIdx by the operation HistValue[cIdx]=1<<StatCoeff[cIdx]; StatCoeff represents the history counter; 5. The video decoding method of claim 4.

6. Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU includes: determining that a neighboring coefficient within a plurality of neighboring coefficients in the predetermined region of the coefficient is located outside the TU; and calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighboring coefficient located outside the TU.

5. The video decoding method of claim 4.

7. Decoding the CTU further comprises: responsive to determining that the first non-zero Golomb-Rice coded transform coefficient in the TU is coded as abs_reminder, updating a history counter for color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_reminder[cIdx]))+2)>>1; in response to determining that a first non-zero Golomb-Rice coded transform coefficient in the TU is coded as dec_abs_level, updating the history counter for the color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1; StatCoeff represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base 2 logarithm of x.

10. The video decoding method of claim 1.

8. 1. A processing device comprising: a memory configured to store a computer program; a processor, wherein the processor executes the computer program stored in the memory to accessing a binary string representing a partition of a video, the partition including a plurality of coding tree units (CTUs) forming one or more CTU rows; For each CTU of the plurality of CTUs in the partition, Before decoding the CTU, determining that the CTU is a first CTU in a current CTU row of the one or more CTU rows in the partition; In response to determining that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of a Rice parameter to an initial value; Decoding the CTU; and performing outputting a decoded partition of the video, the decoded partition including a plurality of decoded CTUs in the partition; configured to run Decoding the CTU includes: Calculating the Rice parameters for transform units (TUs) within the CTU based on the history counters; decoding the binary string corresponding to the TU in the CTU into coefficient values ​​of the TU based on the calculated Rice parameter; determining pixel values ​​of the TUs within the CTU based on the coefficient values; Processing equipment.

9. The partition is a frame, a slice, or a tile.

9. The processing device of claim 8.

10.

11. The processor further comprises: determining a substitution variable HistValue based on the history counter; Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU using values ​​of neighboring coefficients in a predetermined region of the coefficient and the substitution variable HistValue; deriving the Rice parameter for the TU based on the local sum variable locSumAbs.

9. The processing device of claim 8.

12. the processor is further configured to determine a substitution variable HistValue for the color component cIdx by the operation HistValue[cIdx]=1<<StatCoeff[cIdx]; StatCoeff represents the history counter; 12. The processing device of claim 11.

13. The processor further comprises: determining that a neighboring coefficient within a plurality of neighboring coefficients in the predetermined region of the coefficient is located outside the TU; and calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighboring coefficient located outside the TU.

12. The processing device of claim 11.

14. The processor further comprises: responsive to determining that the first non-zero Golomb-Rice coded transform coefficient in the TU is coded as abs_reminder, updating a history counter for color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_reminder[cIdx]))+2)>>1; in response to determining that a first non-zero Golomb-Rice coded transform coefficient in the TU is to be coded as dec_abs_level, updating the history counter for the color component cIdx such that StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1; StatCoeff represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base 2 logarithm of x.

9. The processing device of claim 8.

15. 1. A method for encoding video, comprising: accessing a partition of the video, the partition including a plurality of coding tree units (CTUs), the plurality of CTUs configuring one or more CTU rows; processing the partition of the video to generate a binary representation of the partition; The process comprises: For each CTU of the plurality of CTUs in the partition, Before encoding the CTU, determining that parallel encoding is enabled and that the CTU is a first CTU in a current CTU row of the one or more CTU rows in the partition; In response to determining that the parallel encoding is enabled and that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of a Rice parameter to an initial value; encoding the CTU; and encoding the binary representation of the partition into a bitstream of the video; Encoding the CTU includes: Calculating a Rice parameter for a transform unit (TU) within the CTU based on the history counter; encoding coefficient values ​​of the TU into a binary representation corresponding to the TU within the CTU based on the calculated Rice parameter; Video encoding method.

16. The partition is a frame, a slice, or a tile.

16. The video encoding method of claim 15.

17.

18. Calculating a Rice parameter for a TU within the CTU based on the history counter includes: determining a substitution variable HistValue based on the history counter; Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU using values ​​of neighboring coefficients in a predetermined region of the coefficient and the substitution variable HistValue; deriving the Rice parameter for the TU based on the local sum variable locSumAbs.

16. The video encoding method of claim 15.

19. determining a substitution variable HistValue based on the history counter includes determining a substitution variable HistValue for a color component cIdx by the operation HistValue[cIdx]=1<<StatCoeff[cIdx]; StatCoeff represents the history counter; 20. The video encoding method of claim 18.

20. Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU includes: determining that a neighboring coefficient within a plurality of neighboring coefficients in the predetermined region of the coefficient is located outside the TU; and calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighboring coefficient located outside the TU.

20. The video encoding method of claim 18.

21. Encoding the CTU further comprises: responsive to determining that the first non-zero Golomb-Rice coded transform coefficient in the TU is coded as abs_reminder, updating a history counter for color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_reminder[cIdx]))+2)>>1; in response to determining that a first non-zero Golomb-Rice coded transform coefficient in the TU is coded as dec_abs_level, updating the history counter for the color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(dec_abs_level[cIdx])))>>1; StatCoeff represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base 2 logarithm of x.

16. The video encoding method of claim 15.

22. 1. A processing device comprising: a memory configured to store a computer program; a processor, wherein the processor executes the computer program stored in the memory to accessing a partition of a video, the partition including a plurality of coding tree units (CTUs) that make up one or more CTU rows; processing the partition of the video to generate a binary representation of the partition; The process comprises: For each CTU of the plurality of CTUs in the partition, Before encoding the CTU, determining that parallel encoding is enabled and that the CTU is a first CTU in a current CTU row of the one or more CTU rows in the partition; In response to determining that the parallel encoding is enabled and that the CTU is the first CTU of the current CTU row in the partition, setting a history counter for calculating color components of a Rice parameter to an initial value; encoding the CTU; and encoding the binary representation of the partition into a bitstream of the video; configured to run Encoding the CTU includes: Calculating a Rice parameter for a transform unit (TU) within the CTU based on the history counter; encoding coefficient values ​​of the TU into a binary representation corresponding to the TU within the CTU based on the calculated Rice parameter; Processing equipment.

23. The partition is a frame, a slice, or a tile.

23. The processing device of claim 22.

24.

25. The processor further comprises: determining a substitution variable HistValue based on the history counter; Calculating a local sum variable locSumAbs of coefficients within a TU of the CTU using values ​​of neighboring coefficients in a predetermined region of the coefficient and the substitution variable HistValue; deriving a Rice parameter for the TU based on the local sum variable locSumAbs.

23. The processing device of claim 22.

26. the processor is further configured to determine a substitution variable HistValue for the color component cIdx by the operation HistValue[cIdx]=1<<StatCoeff[cIdx]; StatCoeff represents the history counter; 26. The processing device of claim 25.

27. The processor further comprises: determining that a neighboring coefficient within a plurality of neighboring coefficients in the predetermined region of the coefficient is located outside the TU; and calculating the local sum variable locSumAbs using the substitution variable HistValue as the value of the neighboring coefficient located outside the TU.

26. The processing device of claim 25.

28. The processor further comprises: responsive to determining that the first non-zero Golomb-Rice coded transform coefficient in the TU is coded as abs_reminder, updating a history counter for color component cIdx such that StatCoeff[cIdx]=(StatCoeff[cIdx]+Floor(Log2(abs_reminder[cIdx]))+2)>>1; in response to determining that a first non-zero Golomb-Rice coded transform coefficient in the TU is to be coded as dec_abs_level, updating the history counter for the color component cIdx such that StatCoeff[cIdx] = (StatCoeff[cIdx] + Floor(Log2(dec_abs_level[cIdx]))) >> 1; StatCoeff represents the history counter, Floor(x) represents the largest integer less than or equal to x, and Log2(x) is the base 2 logarithm of x.

23. The processing device of claim 22.

29. A computer-readable storage medium, comprising: A computer-readable storage medium having stored thereon a computer program and a bitstream, the computer program causing a processor to execute the video encoding method of claims 15 to 21 to generate the bitstream.

30. A computer-readable storage medium, comprising: A computer-readable storage medium having stored thereon a computer program and a bitstream, said computer program causing a processor to execute the video decoding method of claims 1 to 7 to decode said bitstream.