Methods, computing devices, storage media, and program products for video decoding

By introducing Rice parameter control for transform skipping stripes and an improved residual encoding/decoding method into video encoding/decoding, combined with multi-type tree structures and dependency scalar quantization, the problem of balancing compression efficiency and quality in existing technologies is solved, achieving more efficient video encoding/decoding.

CN119052480BActive Publication Date: 2025-11-11BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410828406.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2021-09-23
Publication Date
2025-11-11
Estimated Expiration
2041-09-23

AI Technical Summary

Technical Problem

Existing video codec technologies struggle to achieve the optimal balance between compression efficiency and quality, especially in high-efficiency video codec standards such as VVC, where there is room for improvement in codec tools to further enhance compression efficiency.

Method used

We optimize the video encoding and decoding process by employing Rice parameter control with transform skipping stripes and an improved residual encoding and decoding method, combined with multi-type tree structures, dependency scalar quantization, and improved context modeling.

Benefits of technology

It improves the compression efficiency and quality of video encoding and decoding. By optimizing the transform skip mode and residual encoding and decoding, it reduces the bit rate requirement and improves the encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052480B_ABST
    Figure CN119052480B_ABST
Patent Text Reader

Abstract

This disclosure relates to methods, computing devices, storage media, and program products for video decoding. The method for video decoding includes: receiving, by a decoder, control flags at the strip header level and control flags at the sequence picture set (SPS) level, wherein the strip header-level control flags and the SPS-level control flags are used to determine whether at least one syntax element associated with a Rice parameter is signaled for a transform skip stripe; in response to determining that the at least one syntax element is signaled, receiving the at least one syntax element at the strip header level, wherein the at least one syntax element is used to determine the Rice parameter; and performing entropy decoding of a video bitstream by the decoder based on the strip header-level control flags, the SPS-level control flags, and the at least one syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is a divisional application of Chinese Patent Application No. 2021800650417, which is the Chinese national phase application of International Patent Application PCT / US2021 / 051700 filed on September 23, 2021. The International Patent Application claims priority to U.S. Patent Application No. 63 / 085,966 filed on September 30, 2020 and U.S. Patent Application No. 63 / 082,452 filed on September 23, 2020. Technical Field

[0003] This disclosure relates to video encoding / decoding and compression. More specifically, this disclosure relates to improvements and simplifications in residual and coefficient encoding / decoding for video encoding / decoding. Background Technology

[0004] Various video codec techniques can be used to compress video data. Video codecs are performed according to one or more video codec standards. For example, video codec standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High-Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Expert Group (MPEG) codecs. Video codecs typically use prediction methods (e.g., inter-frame prediction, intra-frame prediction), which utilize redundancy present in video images or sequences. A key goal of video codec techniques is to compress video data to a lower bitrate while avoiding or minimizing video quality degradation. Summary of the Invention

[0005] Examples of this disclosure provide methods and apparatus for residual and coefficient encoding / decoding in video encoding / decoding.

[0006] According to a first aspect of this disclosure, a method for video decoding is provided. The method includes a decoder receiving a video bitstream. The decoder may further receive control flags at the strip header level. The control flags may signal whether a rice parameter is enabled for a transform skip stripe. The decoder may also receive at least one syntax element at the strip header level. The at least one syntax element is signaled for the transform skip stripe and indicates the rice parameter. The decoder may perform further entropy decoding on the video bitstream based on the control flags and the at least one syntax element.

[0007] It should be understood that the general description above and the detailed description below are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0008] Examples consistent with this disclosure are illustrated in conjunction with the accompanying drawings, which are included in and form part of this specification, and together with the description, serve to explain the principles of this disclosure.

[0009] Figure 1 This is a block diagram of an encoder based on an example of this disclosure.

[0010] Figure 2 This is a block diagram of a decoder based on an example of this disclosure.

[0011] Figure 3A This is a diagram illustrating block partitioning in a multi-type tree structure according to an example of this disclosure.

[0012] Figure 3B This is a diagram illustrating block partitioning in a multi-type tree structure according to an example of this disclosure.

[0013] Figure 3C This is a diagram illustrating block partitioning in a multi-type tree structure according to an example of this disclosure.

[0014] Figure 3D This is a diagram illustrating block partitioning in a multi-type tree structure according to an example of this disclosure.

[0015] Figure 3E This is a diagram illustrating block partitioning in a multi-type tree structure according to an example of this disclosure.

[0016] Figure 4 This is an illustration of an image with 18×12 luminance coding tree units (CTUs) according to an example of this disclosure.

[0017] Figure 5 This is an illustration of an image with 18×12 luminance CTUs, according to an example of this disclosure.

[0018] Figure 6A This is an illustration of examples of ternary tree (TT) and binary tree (BT) partitions not permitted in a VVC test model (VTM) according to the examples of this disclosure.

[0019] Figure 6B This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0020] Figure 6C This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0021] Figure 6D This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0022] Figure 6E This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0023] Figure 6F This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0024] Figure 6G This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0025] Figure 6H This is an illustration of an example of TT and BT partitions not permitted in a VTM according to the example of this disclosure.

[0026] Figure 7 This is a diagram illustrating a residual encoding / decoding structure for a transform block, according to an example of this disclosure.

[0027] Figure 8 This is an illustration of a residual codec structure for transforming skip blocks, based on an example of this disclosure.

[0028] Figure 9 The illustrations are of two scalar quantizers according to examples of this disclosure.

[0029] Figure 10A This is a diagram illustrating state transitions based on examples of this disclosure.

[0030] Figure 10B This is a diagram illustrating the quantizer selection based on the examples of this disclosure.

[0031] Figure 11 This is a diagram illustrating a template for selecting a probability model according to this disclosure.

[0032] Figure 12 This is an illustration of an example of a coded block according to the palette pattern of this disclosure.

[0033] Figure 13 It is a diagram illustrating the use of palette prediction values ​​to represent palette entries with signals according to this disclosure.

[0034] Figure 14A This is a diagram illustrating a horizontal traversal scan according to this disclosure.

[0035] Figure 14B This is a diagram of a vertical traversal scan according to the present disclosure.

[0036] Figure 15A This is an illustration of a sub-block-based index mapping scan of a color palette according to this disclosure.

[0037] Figure 15B This is an illustration of a sub-block-based index mapping scan of a color palette according to this disclosure.

[0038] Figure 16 This is an example of a method for decoding video signals according to this disclosure.

[0039] Figure 17 This is an example of a method for decoding video signals according to this disclosure.

[0040] Figure 18 This is a diagram illustrating a computing environment coupled to a user interface according to an example of this disclosure. Detailed Implementation

[0041] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. All descriptions herein refer to the accompanying drawings, in which, unless otherwise stated, the same reference numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with the aspects related to this disclosure as set forth in the appended claims.

[0042] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein is intended to represent and include any or all possible combinations of one or more of the associated enumerated items.

[0043] It should be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various types of information, such information should not be limited by these terms. These terms are merely used to distinguish one type of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information; and similarly, second information may be referred to as first information. As used herein, depending on the context, the term “if” may be understood to mean “when,” “in,” or “in response to a judgment.”

[0044] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceived quality compared to its predecessor, H.264 / MPEG-AVC. While HEVC offered significant codec improvements over its predecessor, evidence suggested that even better codec efficiency could be achieved using additional codec tools. Building on this, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG established the Joint Video Exploration Team (JVET) to conduct major research into advanced technologies that could significantly improve codec efficiency. JVET maintains a reference software called the joint exploration model (JEM) by integrating multiple additional codec tools on top of the HEVC test model (HM).

[0045] In October 2017, ITU-T and ISO / IEC released a joint call for proposals (CfP) on video compression capabilities exceeding HEVC. In April 2018, the 10th JVET meeting received and evaluated 23 CfP responses, demonstrating a compression efficiency improvement of approximately 40% compared to HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard called Universal Video Codec (VVC). That same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0046] Like HEVC, VVC is built on a block-based hybrid video codec framework.

[0047] Figure 1 A general diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode decision 116, block prediction value 140, adder 128, transform 130, quantization 132, prediction related information 142, intra-frame prediction 118, image buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, loop filter 122, entropy coding 138, and bitstream 144.

[0048] In encoder 100, video frames are partitioned into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.

[0049] The prediction residual, representing the difference between the current video block (a portion of video input 110) and its predicted value (a portion of block prediction value 140), is sent from adder 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then fed to entropy coding 138 to generate a compressed video bitstream. Figure 1 As shown, prediction-related information 142 from the intra / inter-frame mode decision 116 (such as video block partitioning information, motion vector (MV), reference image index, and intra-frame prediction mode) is also fed into the compressed bitstream 144 via entropy coding 138. The compressed bitstream 144 includes the video bitstream.

[0050] In encoder 100, additional circuitry associated with the decoder is required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is then combined with block prediction values ​​140 to generate the unfiltered reconstructed pixels for the current video block.

[0051] Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video frame as the current video block to predict the current video block.

[0052] Timing prediction (also known as "inter-frame prediction") uses reconstructed pixels from encoded video frames to predict the current video block. Timing prediction reduces the temporal redundancy inherent in the video signal. The timing prediction signal for a given coding unit (CU) or coding block is typically represented by one or more MV signals indicating the amount and direction of motion between the current CU and its timing reference. Further, if multiple reference frames are supported, an additional reference frame index is sent to identify which reference frame in the reference frame storage the timing prediction signal originates from.

[0053] Motion estimation 114 acquires the video input 110 and the signal from the image buffer 120, and outputs the motion estimation signal to motion compensation 112. Motion compensation 112 acquires the video input 110, the signal from the image buffer 120, and the motion estimation signal from motion estimation 114, and outputs the motion compensation signal to intra / inter-frame mode decision 116.

[0054] After performing spatial and / or temporal prediction, the intra / inter-frame mode decision 116 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Then, the block prediction value 140 is subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantization residual coefficients are dequantized by inverse quantization 134, and then inversely transformed by inverse transform 136 to form a reconstruction residual. This reconstruction residual is then added back to the prediction block to form the reconstructed CU signal. Further, before placing the reconstructed CU in the reference image storage device of image buffer 120 for encoding and decoding future video blocks, loop filters 122, such as deblocking filters, sample adaptive offset (SAO), and / or adaptive in-loop filters (ALF), can be applied to the reconstructed CU. In order to form the output video bitstream 144, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are sent to the entropy coding unit 138 for further compression and packing to form the bitstream.

[0055] Figure 1A block diagram of a general block-based hybrid video coding system is presented. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128×128 pixels. However, unlike HEVC, which only partitions blocks based on quadtrees, in VVC, a coding tree unit (CTU) is divided into multiple CUs to accommodate different local characteristics based on quadtrees, binary trees, or tritrees. By definition, a coding tree block (CTB) is an N×N sample block of some value N, such that dividing the components into CTBs is partitioning. A CTU includes a luminance sample CTB of an image with three sample arrays, two corresponding chrominance sample CTBs, or a sample CTB of a coded monochrome image or image using three separate color planes, and a syntax structure for encoding the samples. Furthermore, the concept of multiple partitioning unit types in HEVC has been removed; that is, VVC no longer separates CUs, prediction units (PUs), and transform units (TUs). Instead, each CU always serves as the basic unit for both prediction and transformation, without further partitioning. In a multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary tree and ternary tree structures. Figure 3A , Figure 3B , Figure 3C , Figure 3D and Figure 3E As shown, there are five partitioning types: quad partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0056] Figure 3A A diagram illustrating block quadrilateral partitioning in a multi-type tree structure is shown. Figure 3B A diagram illustrating block vertical binary partitioning in a multi-type tree structure is shown. Figure 3C A diagram illustrating block-level binary partitioning in a multi-type tree structure according to this disclosure is shown. Figure 3D A diagram illustrating block vertical ternary partitioning in a multi-type tree structure is shown. Figure 3E A diagram illustrating block-level ternary partitioning in a multi-type tree structure is shown.

[0057] exist Figure 1In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded adjacent blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Similarly, if multiple reference pictures are supported, an additional reference picture index is sent to identify which reference picture in the reference picture storage device the temporal prediction signal comes from. After performing spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are dequantized and inversely transformed to form the reconstructed residuals, which are then added back to the prediction block to form the reconstructed CU signal. Further, before placing the reconstructed CU in the reference image storage device and using it for encoding and decoding future video blocks, loop filtering such as deblocking filters, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF) can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and the quantized residual coefficients are sent to the entropy coding unit for further compression and packing to form the bitstream.

[0058] Figure 2 A general block diagram of a video decoder for VVC is shown. Specifically, Figure 2 A block diagram of a typical decoder 200 is shown. The decoder 200 has a bitstream 210, entropy decoding 212, dequantization 214, inverse transform 216, adder 218, intra / inter-frame mode selection 220, intra-frame prediction 222, memory 230, loop filter 228, motion compensation 224, image buffer 226, prediction-related information 234, and video output 232.

[0059] Decoder 200 is similar to that which exists Figure 1The reconstruction-related part is located in the encoder 100. In the decoder 200, the incoming video bitstream 210 is first decoded by entropy decoding 212 to obtain quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by inverse quantization 214 and inverse transform 216 to obtain the reconstruction prediction residual. The block prediction mechanism implemented in the intra / inter-frame mode selector 220 is configured to perform intra-frame prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstruction prediction residual from inverse transform 216 with the prediction output generated by the block prediction mechanism using summer 218.

[0060] Before being stored in the image buffer 226, which serves as a reference image storage device, the reconstructed blocks can be further passed through the loop filter 228. The reconstructed video in the image buffer 226 can be sent to drive the display device and used to predict future video blocks. With the loop filter 228 enabled, filtering operations are performed on these reconstructed pixels to obtain the final reconstructed video output 232.

[0061] Figure 2 A general block diagram of a block-based video decoder is presented. First, entropy decoding of the video bitstream is performed at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (in the case of intra-frame coding) or the temporal prediction unit (in the case of inter-frame coding) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. Then, the prediction blocks and residual blocks are added together. Before being stored in the reference image storage device, the reconstructed blocks can be further filtered through a loop. Finally, the reconstructed video in the reference image storage device is sent to drive the display device and is used to predict future video blocks.

[0062] Typically, the basic intra prediction scheme used in VVC is the same as that in HEVC, except that several modules are further extended and / or improved, such as intra sub-partition (ISP) coding mode, extended intra prediction using wide-angle intra direction, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation.

[0063] Partitioning of images, tile groups, tiles, and CTUs in VVC

[0064] In VVC, a tile is defined as a rectangular CTU region within a specific tile column and row in an image. A tile group is a group of an image containing an integer number of tiles that are contained within a single Network Abstraction Layer (NAL) unit. Essentially, the concept of a tile group is the same as that of a stripe defined in HEVC. For example, an image is divided into tile groups and tiles. A tile is a series of CTUs that cover a rectangular region of the image. A tile group contains multiple tiles of the image. Two modes of tile groups are supported: raster scan tile group mode and rectangular tile group mode. In raster scan tile group mode, a tile group contains a series of tiles from a raster scan of the image. In rectangular tile group mode, a tile group contains multiple tiles of the image that together form a rectangular region of the image. The tiles within a rectangular tile group are arranged according to the raster scan order of the tile group.

[0065] Figure 4 An example of raster scan tile grouping of an image is shown, in which the image is divided into 12 tiles and 3 raster scan tile groups. Figure 4 This includes tiles 410, 412, 414, 416, and 418. Each tile has 18 CTUs. More specifically, Figure 4 An image with 18×12 luminance CTUs is shown, which is divided into 12 tiles and 3 tile groups (informative). The three tile groups are as follows: (1) the first tile group includes tiles 410 and 412, (2) the second tile group includes tiles 414, 416, 418, 420 and 422, and (3) the third tile group includes tiles 424, 426, 428, 430 and 432.

[0066] Figure 5 An example of rectangular tile grouping for an image is shown, where the image is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tile groups. Figure 5 This includes tiles 510, 512, 514, 516, 518, 520, 522, 524, 526, 528, 530, 532, 534, 536, 538, 540, 542, 544, 546, 548, 550, 552, 554, and 556. More specifically, Figure 5An image with 18×12 luminance CTUs is shown, which is divided into 24 tiles and 9 tile groups (informative). Tile groups contain tiles, and tiles contain CTUs. The 9 rectangular tile groups include: (1) two tiles 510 and 512, (2) two tiles 514 and 516, (3) two tiles 518 and 520, (4) four tiles 522, 524, 534 and 536, (5) four tile groups 526, 528, 538 and 540, (6) four tiles 530, 532, 542 and 544, (7) two tiles 546 and 548, (8) two tiles 550 and 552, and (9) two tiles 554 and 556.

[0067] Large-size transformation with high-frequency zeroing in VVC

[0068] In VTM4, large block size transforms up to 64×64 are enabled, primarily for higher resolution video, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed out, leaving only low-frequency coefficients. For example, for an M×N transform block (where M is the block width and N is the block height), when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the first 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without zeroing out any values.

[0069] Virtual pipeline data units (VPDUs) in VVC

[0070] Virtual Pipeline Data Units (VPDUs) are defined as non-overlapping units in an image. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so keeping the VPDU size small is important. In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, ternary tree (TT) and binary tree (BT) partitioning can lead to an increase in VPDU size.

[0071] To maintain the VPDU size at 64×64 luminance samples, the following partition restrictions (with syntax signaling modifications) are applied in VTM5:

[0072] For a CU with a width or height equal to 128, or both width and height equal to 128, TT splitting is not allowed. For a CU of size 128xN with N≤64 (i.e., width equal to 128 and height less than 128), horizontal BT splitting is not allowed. For a CU of size Nx128 with N≤64 (i.e., height equal to 128 and width less than 128), vertical BT splitting is not allowed.

[0073] Figure 6A , Figure 6B , Figure 6C , Figure 6D , Figure 6E , Figure 6F , Figure 6G and Figure 6H Examples of TT and BT partitions that are not allowed in VTM are shown.

[0074] Transform coefficient encoding and decoding in VVC

[0075] Transform coefficient encoding and decoding in VVC is similar to HEVC because they both use non-overlapping coefficient groups (also called CGs or sub-blocks). However, there are some differences between the two encoding and decoding methods. In HEVC, each CG of a coefficient has a fixed size of 4×4. In VVC draft 6, the CG size becomes dependent on the size of the TB. Therefore, various CG sizes (1×16, 2×8, 8×2, 2×4, 4×2, and 16×1) are provided in VVC. CGs within a coded block and transform coefficients within a CG are encoded and decoded according to a predefined scan order.

[0076] To limit the maximum number of bits for context coding per pixel, the maximum number of bits for context-coded bin (CCB) of a TB is determined using the area of ​​the TB and the type of video components (e.g., luma and chroma components). The maximum number of bits for context coding is equal to Bozize * 1.75. Here, TB_zosize indicates the number of samples within the TB after the coefficients are zeroed. Note that for CCB counting, coded_sub_block_flag (which indicates whether the CG contains non-zero coefficients) is not considered.

[0077] Coefficient zeroing is an operation performed on a transform block that forces coefficients located in a specific region of the transform block to be zero. For example, in the current VVC, the 64×64 transform has an associated zeroing operation. Therefore, transform coefficients located outside the top-left 32×32 region within the 64×64 transform block are forced to be zero. In fact, in the current VVC, for any transform block with a size exceeding 32 along a certain dimension, a coefficient zeroing operation is performed along that dimension to force coefficients located outside the top-left 32×32 region to be zero.

[0078] In VVC transform coefficient encoding and decoding, the variable `remBinsPass1` is initially set to the maximum number of allowed context-coded bins (MCCB). During encoding and decoding, this variable is decremented by one each time context-coded bins are represented by signals. When `remBinsPass1` is greater than or equal to four, coefficients are first represented by signals using the syntaxes `sig_coeff_flag`, `abs_level_gtl_flag`, `par_level_flag`, and `abs_level_gt3_flag`. All these syntax elements use context-coded bins in the first pass. The remaining level information of the coefficients is encoded using the `absremainder` syntax element in the second pass, using Golomb-Rice codes and bypass-coded bins. When `remBinsPass1` becomes less than four during the first pass, the current coefficient is not encoded in the first pass but is directly encoded in the second pass using Golomb-Rice codes and bypass-coded bins via the `decabslevel` syntax element. The rice parameter derivation process for dec_abs_level[] is as specified in Table 3. After encoding all the levels mentioned above, the sign (sign_flag) of all scan positions where sig_coeff_flag equals 1 is finally encoded as bypass binary bits. This process... Figure 7 The following is a description of the process. RemBinsPass1 is reset for each TB. The transformation from context-encoded bits for sig_coeff_flag, abs_level_gtl_flag (abs_level_gtx_flag[0]), par_level_flag, and abs_level_gt3_flag (abs_level_gtx_flag[1]) to bypass-encoded bits for the remaining coefficients occurs at most once per TB. For a coefficient subblock, if remBinsPass1 is less than 4 before encoding its first coefficient, the entire coefficient subblock is encoded using bypass-encoded bits.

[0079] Figure 7 A diagram of the residual encoding / decoding structure used for transform blocks is shown.

[0080] The syntax for representing `abs_remainder` and `dec_abs_level` using signals is derived using a unified (identical) `ricePara` parameter. The only difference is that the base level `baseLevel` is set to 4 and 0 respectively for encoding `abs_remainder` and `dec_abs_level`. The `ricePara` parameter is determined not only based on the sum of the absolute levels of the five adjacent transform coefficients in the local template but also based on the corresponding base level, as shown below:

[0081] RicePara=RiceParTable[max(min(31,sumAbs–5*baseLevel),0)]

[0082] The syntax and related semantics of residual encoding and decoding in the current VVC draft specification are illustrated in Tables 1 and 2, respectively. How to read Table 1 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0083] Table 1. Syntax of Residual Encoding and Decoding

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] Table 2. Semantics of Residual Encoding and Decoding

[0092]

[0093]

[0094]

[0095]

[0096] Table 3. Derivation of Rice parameters for abs_remainder[] and dec_abs_level[]

[0097]

[0098] Table 4. Specification of cRiceParam based on locSumAbs

[0099] locSumAbs 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 cRiceParam 0 0 0 0 0 0 0 1 1 1 1 1 1 1 2 2 locSumAbs 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 cRiceParam 2 2 2 2 2 2 2 2 2 2 2 2 3 3 3 3

[0100] Residual encoding and decoding of transform skip mode in VVC

[0101] Unlike HEVC, which uses a single residual codec scheme to encode and decode both transform coefficients and transform skip coefficients, VVC uses two separate residual codec schemes for the transform coefficients and transform skip coefficients (i.e., the residuals).

[0102] In transform skip mode, the statistical properties of the residual signal differ from those of the transform coefficients, and no energy accumulation is observed around the low-frequency components. The residual encoding / decoding is modified to account for the different signal characteristics of the (spatial) transform skip residuals, including:

[0103] The final x / y position is not indicated by a signal;

[0104] When all previous flags are equal to 0, the coded_sub_block_flag of each sub-block except for the DC sub-block is encoded;

[0105] Model using the context of sig_coeff_flag of two adjacent coefficients;

[0106] Use only one context model's par_level_flag;

[0107] Additional markers of 5, 7, or 9 or more;

[0108] Derivation of the rice parameter for the modification of remainder binarization; and

[0109] The context modeling of the symbol flag is determined based on the adjacent coefficient values ​​on the left and above, and the symbol flag is parsed after sig_coeff_flag to keep all context-encoded bits together.

[0110] like Figure 8As shown (as described below), in the first pass, the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag (abs_level_gtx_flag[0]) and par_level_flag are encoded in an interleaved manner on a residual sample basis. In the second pass, the bit plane of abs_level_gtX_flag (abs_level_gtx_flag[0], abs_level_gtx_flag[1], abs_level_gtx_flag[2] and abs_level_gtx_flag[3]) is encoded, and in the third pass, abs_remainder is encoded.

[0111] Pass 1: sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag(abs_level_gtx_flag[0]) and par_level_flag.

[0112] Pass 2: abs_level_gt3_flag(abs_level_gtx_flag[0]), abs_level_gt5_flag(abs_level_gtx_flag[1]), abs_level_gt7_flag(abs_level_gtx_flag[2]) and abs_level_gt9_flag(abs_level_gtx_flag[3]).

[0113] The 3rd pass: abs_remainder.

[0114] Figure 8 A diagram of the residual codec structure used for transforming skip blocks is shown.

[0115] The syntax and related semantics of residual encoding / decoding used for transform skip modes in the current VVC draft specification are illustrated in Tables 5 and 2, respectively. How to read Table 5 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0116] Table 5. Syntax for Residual Encoding and Decoding for Transform Skip Mode

[0117]

[0118]

[0119]

[0120]

[0121] Quantification

[0122] In the current VVC, the maximum quantization parameter (QP) value has been expanded from 51 to 63, and the signal representation of the initial QP has also changed accordingly. When encoding non-zero values ​​of slice_qp_delta, the initial value of SliceQpY can be modified at the slice layer. For transform skip blocks, the minimum allowed quantization parameter (QP) is defined as 4, because when QP equals 4, the quantization step size becomes 1.

[0123] Furthermore, the same HEVC scalar quantization uses a new concept called dependent scalar quantization. Dependent scalar quantization is a method where a set of allowable reconstructed values ​​of the transform coefficients depends on the values ​​of the transform coefficient levels preceding the current transform coefficient level in the reconstruction order. The main effect of this method is that, compared to the traditional independent scalar quantization used in HEVC, the allowable reconstructed vectors are denser in the N-dimensional vector space (where N represents the number of transform coefficients in the transform block). This means that for a given average number of allowable reconstructed vectors per N-dimensional unit volume, the average distortion between the input vector and the nearest reconstructed vector is reduced. The dependent scalar quantization method is implemented by: (a) defining two scalar quantizers with different reconstruction levels, and (b) defining a procedure for switching between these two scalar quantizers.

[0124] The two scalar quantizers used are as follows: Figure 9 The diagram uses Q0 and Q1 (described below). The location of the available reconstruction level is uniquely specified by the quantization step size Δ. The scalar quantizer used (Q0 or Q1) is not explicitly represented as a signal in the bitstream. Instead, the quantizer used for the current transform coefficient is determined by the parity of the transform coefficient levels preceding the current transform coefficient in the encoding / decoding / reconstruction order.

[0125] Figure 9 A diagram of two scalar quantizers used in the proposed dependency quantization method is shown.

[0126] like Figure 10A and Figure 10BAs shown (described below), the switching between these two scalar quantizers (Q0 and Q1) is implemented via a state machine with four quantizer states (QState). QState can take four different values: 0, 1, 2, and 3. It is uniquely determined by the parity of the transform coefficient levels preceding the current transform coefficient in the encoding / decoding / reconstruction order. At the start of dequantization of the transform block, the state is set to 0. The transform coefficients are reconstructed in scan order (i.e., the same order in which they were entropy decoded). After the current transform coefficient is reconstructed, the state is updated, as shown in Figure 10, where k represents the value of the transform coefficient level.

[0127] Figure 10A The diagram illustrates the state transitions for the dependency quantification proposed in the illustration.

[0128] Figure 10B A table showing the quantizer selection for the dependency quantization proposed in the illustration is presented.

[0129] It also supports using signals to represent default and user-defined scaling matrices. In DEFAULT mode, the scaling matrix remains unchanged; all elements of the TB size are equal to 16. Intra Block Copy (IBC) and intra codec modes currently share the same scaling matrix. Therefore, for the USER_DEFINED matrix, the number of MatrixType and MatrixType_DC are updated as follows:

[0130] MatrixType: 30 = 2 (2 for intra-frame and IBC / inter-frame) × 3 (Y / Cb / Cr components) × 5 (square TB size: luminance from 4×4 to 64×64, chrominance from 2×2 to 32×32).

[0131] MatrixType_DC: 14 = 2 (2 for intra-frame and IBC / inter-frame, Y component is 1) × 3 (TB size: 16×16, 32×32, 64×64) + 4 (2 for intra-frame and IBC / inter-frame, Cb / Cr component is 2) × 2 (Tb size: 16×16, 32×32).

[0132] For the following scaling matrices: 16×16, 32×32, and 64×64, the DC values ​​are encoded separately. For TBs smaller than 8×8, all elements in a scaling matrix are represented by signals. If the size of the TB is greater than or equal to 8×8, only 64 elements in an 8×8 scaling matrix are represented by signals as the basic scaling matrix. To obtain square matrices larger than 8×8, the 8×8 basic scaling matrix is ​​upsampled (by element-wise copying) to the corresponding square size (i.e., 16×16, 32×32, 64×64). When zeroing is applied to the high-frequency coefficients of the 64-point transform, the corresponding high frequencies of the scaling matrix are also zeroed. That is, if the width or height of the TB is greater than or equal to 32, only the coefficients in the left or upper half are retained, and the remaining coefficients are assigned zero. Furthermore, the number of elements represented by signals for a 64×64 scaling matrix is ​​reduced from 8×8 to three 4×4 submatrices because the bottom right 4×4 elements are never used.

[0133] Context modeling for transform coefficient encoding and decoding

[0134] The choice of the probabilistic model for the grammatical elements related to the absolute value of the transform coefficient level depends on the value of the absolute level in the local neighborhood or the absolute level of the partially reconstructed model. The template used... Figure 11 The diagram is shown below.

[0135] Figure 11 This diagram illustrates the template used to select the probabilistic model. The black squares specify the current scan position, and the squares marked with "x" represent the local neighborhood used.

[0136] The chosen probabilistic model depends on the sum of the absolute levels (or partially reconstructed absolute levels) in the local neighborhood and the number of absolute levels greater than 0 in the local neighborhood (given by the number of sig_coeff_flags equal to 1). Context modeling and binarization depend on the following metrics of the local neighborhood:

[0137] numSig: The number of non-zero levels in the local neighborhood;

[0138] sumAbs1: The sum of absolute levels of the partial reconstruction after the first scan in the local neighborhood (absLevel1);

[0139] sumAbs: The sum of the absolute levels of reconstructions in the local neighborhood; and

[0140] Diagonal position (d): The sum of the horizontal and vertical coordinates of the current scan position within the transform block.

[0141] Based on the values ​​of numSig, sumAbs1, and d, a probabilistic model is selected for encoding sig_coeff_flag, abs_level_gt1_flag, and par_level_flag. The Rice parameter is selected for binarizing abs_remainder and dec_abs_level based on the values ​​of sumAbs and numSig.

[0142] In the current VVC, the simplified 32-point MTS (also known as RMTS32) is based on skipping high-frequency coefficients and is used to reduce the computational complexity of 32-point DST-7 / DCT-8. Furthermore, it is accompanied by changes in coefficient encoding and decoding, including all types of zeroing (i.e., the existing zeroing of high-frequency components in RMTS32 and DCT2). Specifically, the binarization of the last non-zero coefficient position encoding and decoding is based on a reduced TU size, and the choice of the context model for the last non-zero coefficient position encoding and decoding is determined by the original TU size. Additionally, 60 context models are used to encode the sig_coeff_flag of the transform coefficients. The context model index is selected based on the sum of the absolute levels of the five largest previous partial reconstructions (called locSumAbsPass1) and the dependency quantization state QState, as follows:

[0143] If cIdx equals 0, then ctxInc is obtained as follows:

[0144] ctxInc=12*Max(0,QState–1)+

[0145] Min((locSumAbsPass1+1)>>1,3)+(d<2?8:(d<5?4:0))

[0146] Otherwise (cIdx is greater than 0), ctxInc is obtained as follows:

[0147] ctxInc=36+8*Max(0,QState-1)+

[0148] Min((locSumAbsPass1+1)>>1,3)+(d<2?4:0)

[0149] Palette mode

[0150] The basic idea behind the palette mode is that samples in the CU are represented by a small set of representative color values. This set is called the palette. Color values ​​excluded from the palette can also be indicated by signaling escape colors; the values ​​of the three color components of the escape color are directly represented as signals in the bitstream. Figure 12 The diagram illustrates this point.

[0151] Figure 12 An example of a block that has been encoded in palette mode is shown. Figure 12 Includes coded block 1210 and palette 1220 in palette mode.

[0152] exist Figure 12 In this context, the palette size is 4. The first three samples are reconstructed using palette entries 2, 0, and 3, respectively. Black samples represent escape symbols. The CU level flag `palette_escape_val_present_flag` indicates whether any escape symbols exist in the CU. If escape symbols exist, the palette size is increased by one, and the last index is used to indicate the escape symbol. Therefore, in Figure 12 In the middle, index 4 is assigned to the escape symbol.

[0153] In order to decode the encoded blocks of the palette, the decoder needs the following information: the palette table; and the palette index.

[0154] If the palette index corresponds to an escape symbol, then additional overhead is indicated by a signal to indicate the corresponding color value of the sample.

[0155] Additionally, on the encoder side, it is necessary to obtain a suitable color palette for use with the CU.

[0156] For the derivation of the lossy encoding / decoding palette, an improved k-means clustering algorithm was used. The first sample of a block is added to the palette. Then, for each subsequent sample in the block, the sum of absolute differences (SAD) between the sample and each current palette color is calculated. If the distortion of each component is less than a threshold corresponding to the minimum SAD of the palette entry, the sample is added to the cluster belonging to that palette entry. Otherwise, the sample is added as a new palette entry. When the number of samples mapped to a cluster exceeds the threshold, the centroid of that cluster is updated and becomes the palette entry for that cluster.

[0157] In the next step, the clusters are sorted in descending order of usage. Then, the palette entry corresponding to each entry is updated. Typically, the centroid of the cluster is used as the palette entry. However, when considering the encoding / decoding cost of palette entries, rate-distortion analysis is performed to analyze whether any entry from the palette predictions is more suitable to replace the centroid as the updated palette entry. This process continues until all clusters have been processed or the maximum palette size is reached. Finally, if a cluster has only one sample and the corresponding palette entry is not in the palette predictions, that sample is converted to an escape symbol. Additionally, duplicate palette entries are removed and their clusters are merged.

[0158] After palette derivation, each sample in the block is assigned the index of the nearest (SAD) palette entry. The sample is then assigned to either the 'INDEX' or 'COPY_ABOVE' mode. For each sample, either 'INDEX' or 'COPY_ABOVE' mode can be used. The encoding / decoding cost of the mode is then calculated. The mode with the lower cost is selected.

[0159] For encoding and decoding of palette entries, palette prediction values ​​are maintained. The maximum size of the palette and the palette prediction values ​​are represented by signals in SPS. Palette prediction values ​​are initialized at the beginning of each CTU line, each strip, and each tile.

[0160] For each entry in the palette prediction, a reuse flag is signaled to indicate whether it is part of the current palette. Figure 13 The diagram illustrates this point.

[0161] Figure 13 This demonstrates the use of palette prediction values ​​to represent palette entries with signals. Figure 13 This includes the previous color palette 1310 and the current color palette 1320.

[0162] The reuse flag is transmitted using zero-run-length encoding / decoding. Following this, zero-order exponential Golomb code is used to signal the number of new palette entries. Finally, the component values ​​of the new palette entries are signaled.

[0163] like Figure 14A and Figure 14B As shown, the palette index is encoded using both horizontal and vertical traversal scans. The scan order is explicitly indicated by a signal using `palette_transpose_flag` in the bitstream.

[0164] Figure 14A The horizontal traversal scan is shown. Figure 14B The vertical traversal scan is shown.

[0165] For the codec palette index, a palette mode based on coefficient groups (CGs) is used, which divides the CU into multiple segments with 16 samples each based on a traversal scan mode, such as... Figure 15A and Figure 15B As shown, for each CG, the index run, palette index value, and quantized color for escape mode are encoded / parsed sequentially.

[0166] Figure 15A This illustrates a sub-block-based index mapping scan for a color palette. Figure 15B This illustrates a sub-block-based index mapping scan for a color palette.

[0167] Palette indices are encoded using two main palette sample modes: 'INDEX' and 'COPY_ABOVE'. As previously explained, escape symbols are assigned an index equal to the maximum palette size. In 'COPY_ABOVE' mode, the palette index of the sample from the previous row is copied. In 'INDEX' mode, the palette index is explicitly represented by a signal. The encoding and decoding sequence of the palette within each segment is as follows:

[0168] For each pixel, a context-coded binary bit `run_copy_flag = 0` is used to indicate whether the pixel has the same pattern as the previous pixel; that is, whether both the previous scanned pixel and the current pixel have the run type `COPY_ABOVE` or whether both the previous scanned pixel and the current pixel have the run type `INDEX` and the same index value. Otherwise, `run_copy_flag = 1` is used.

[0169] If a pixel has a different mode than the previous pixel, a context-encoded binary bit, copy_above_palette_indices_flag, indicating the pixel's run type (i.e., INDEX or COPY_ABOVE), is used. If the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not need to resolve the run type because the INDEX mode is used by default. Similarly, if the previously resolved run type is COPY_ABOVE, the decoder also does not need to resolve the run type.

[0170] After the pixels in a segment are encoded and decoded using a palette, the index value of the INDEX mode (palette_idx_idc) and the quantized escape color (palette_escape_val) are bypassed and encoded.

[0171] Inefficient video decoding

[0172] In VVC, when encoding and decoding transform coefficients, a unified (identical) rice parameter (RicePara) derivation is used to represent the syntax of abs_remainder and dec_abs_level using signals. The only difference is that the base level is set to 4 and 0 for encoding and decoding abs_remainder and dec_abs_level, respectively. The rice parameter is determined not only based on the sum of the absolute levels of five adjacent transform coefficients in the local template, but also based on the corresponding base level, as follows:

[0173] RicePara=RiceParTable[max(min(31,sumAbs–5*baseLevel),0)]

[0174] In other words, the binary codewords of the syntax elements `abs_remainder` and `dec_abs_level` are adaptively determined based on the level information of adjacent coefficients. Since this codeword determination is performed on a per-sample basis, additional logic is required to handle the codeword's adaptation to coefficient encoding and decoding.

[0175] Similarly, when encoding and decoding residual blocks in transform skip mode, the binary codeword of the syntax element abs_remainder is adaptively determined based on the level information of adjacent residual samples.

[0176] Furthermore, when encoding or decoding syntax elements related to residual encoding / decoding or transform coefficient encoding / decoding, the choice of probabilistic model depends on the level information of adjacent levels, which requires additional logic and additional context models.

[0177] In the current design, the binarization of escape samples is achieved by calling the third-order exponential Golomb binarization process. There is room for further performance improvement.

[0178] In the current VVC, there are two different level mapping schemes, applied to regular transforms and transform skipping, respectively. Each level mapping scheme is associated with different conditions, mapping functions, and mapping positions. For blocks applying regular transforms, one level mapping scheme is used after the number of context-coded bits (CCB) exceeds the limit. The mapping position, represented by ZeroPos[n], and the mapping result, represented by AbsLevel[xC][yC], are obtained as specified in Table 2. For blocks applying transform skipping, another level mapping scheme is used before the number of context-coded bits (CCB) exceeds the limit. The mapping position, represented by predCoeff, and the mapping result, represented by AbsLevel[xC][yC], are obtained as specified in Table 5. From a standardization perspective, this non-uniform design may not be optimal.

[0179] For HEVC profiles exceeding 10 bits, an extended_precision_processing_flag of 1 indicates the use of an extended dynamic range for coefficient resolution and inverse transform processing. In current VVC, residual encoding / decoding or transform skip encoding / decoding of transform coefficients exceeding 10 bits is reportedly a significant performance degradation issue. There is still room for performance improvement.

[0180] The proposed method

[0181] This disclosure presents several methods for addressing the low video decoding efficiency issue mentioned in this section. Note that the following methods can be used individually or in combination.

[0182] According to a first aspect of this disclosure, it is proposed to use a fixed set of binary codewords for encoding and decoding certain syntax elements (e.g., abs_remainder) in residual encoding and decoding. Different methods can be used to form the binary codewords. Some exemplary methods are listed below.

[0183] First, use the same procedure used in the current VVC to determine the codeword for abs_remainder, but always with a selected fixed rice parameter (e.g., 1, 2, or 3).

[0184] Second, fixed-length binarization.

[0185] Third, truncate Rice binarization.

[0186] Fourth, truncate the binary (TB) binarization process.

[0187] Fifth, the k-order exponent Golomb binarization process (EGk).

[0188] Sixth, finite k-order exponent Golomb binarization.

[0189] According to a second aspect of this disclosure, it is proposed to use a fixed codeword set for encoding and decoding certain syntax elements (e.g., abs_remainder and dec_abs_level) in transform coefficient encoding and decoding. Different methods can be used to form the binary codewords. Some exemplary methods are listed below.

[0190] First, use the same procedure as the current VVC for determining the codewords for abs_remainder and dec_abs_level, but with a fixed rice parameter, such as 1, 2, or 3. The value of baseLevel can still differ for abs_remainder and dec_abs_level used in the current VVC. (For example, baseLevel could be set to 4 and 0 respectively for encoding and decoding abs_remainder and dec_abs_level).

[0191] Second, use the same procedure as currently used in VVC to determine the codewords for abs_remainder and dec_abs_level, but with fixed rice parameters, such as 1, 2, or 3. The baseLevels values ​​for abs_remainder and dec_abs_level are chosen to be the same, for example, 0 for both or 4 for both.

[0192] Third, fixed-length binarization.

[0193] Fourth, truncate Rice binarization.

[0194] Fifth, truncate the binary (TB) binarization process.

[0195] Sixth, the k-th order exponent Golomb binarization process (EGk).

[0196] Seventh, finite k-order exponent Golomb binarization.

[0197] According to a third aspect of this disclosure, it is proposed to use a single context to encode and decode syntax elements (e.g., abs_level_gtx_flag) related to residual encoding / decoding or coefficient encoding / decoding, and context selection based on adjacent decoding level information can be removed.

[0198] According to a fourth aspect of this disclosure, it is proposed to use a variable set of binary codewords to encode and decode certain syntax elements (e.g., abs_remainder) in residual encoding and decoding, and to determine the selection of the binary codeword set based on certain encoded information of the current block, such as quantization parameters (QP) associated with TB / CB and / or slices, the prediction mode of CU (e.g., IBC mode or intra-frame or inter-frame mode), and / or slice type (e.g., I-slice, P-slice, or B-slice). Different methods can be used to obtain the variable set of binary codewords; some exemplary methods are listed below.

[0199] First, use the same procedure as the current VVC for determining the codewords for abs_remainder, but with different rice parameters.

[0200] Second, the k-order exponent Golomb binarization process (EGk).

[0201] Third, finite k-order exponent Golomb binarization.

[0202] Table 6. Rice parameter determination based on QP value

[0203]

[0204]

[0205] The same method explained in the fourth aspect is also applicable to transform efficient coding and decoding. According to the fifth aspect of the present disclosure, it is proposed to use a variable set of binary codewords to encode and decode certain syntax elements (such as abs_remainder and dec_abs_level) in transform coefficient coding and decoding, and to determine the selection of the binary codeword set according to certain encoded information of the current block, such as the quantization parameter (QP) associated with the TB / CB and / or the slice, the prediction mode of the CU (such as the IBC mode or intra or inter-frame), and / or the slice type (such as I-slice, P-slice or B-slice). Again, different methods can be used to obtain the variable set of binary codewords, and some exemplary methods are listed below.

[0206] First, use the same process for determining the codeword of abs_remainder as that used in the current VVC, but adopt different rice parameters.

[0207] Second, the k-th order exponential Golomb binarization process (EGk).

[0208] Third, the finite k-th order exponential Golomb binarization.

[0209] In these above methods, different rice parameters can be used to obtain different variable sets of binary codewords. For a given residual sample block, the rice parameter used is determined according to the CU QP represented as QP CU rather than the adjacent level information. A specific example is shown in Table 6, where TH1 to TH4 are predefined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predefined rice parameters. It should be noted that the same logic can have different implementation manners in practice. For example, the same rice parameter can also be obtained from the QP value of the current CU using certain equations or look-up tables, as shown in Table 6.

[0210] According to the fifth aspect of the present disclosure, a parameter set and / or a threshold associated with the codeword determination of the syntax elements in transform coefficient coding and decoding and / or transform skip residual coding are signaled into the bitstream. When encoding and decoding the syntax elements through an entropy codec (e.g., arithmetic coding), the determined codeword is used as the binarized codeword.

[0211] Note that the parameter set and / or threshold can be a complete set or subset of all parameters and thresholds associated with the codeword determination of the syntax elements. The parameter set and / or threshold can be represented by signals at different levels in the video bitstream. For example, these parameter sets and / or thresholds can be represented by signals at the sequence level (e.g., sequence parameter set), picture level (e.g., picture parameter set and / or picture header), stripe level (e.g., stripe header), coding tree unit (CTU) level, or coding unit (CU) level.

[0212] In one example, the rice parameters, represented by signals, are used in the strip header, picture header, picture parameter set (PPS), and / or sequence parameter set (SPS) to determine the codewords used for encoding and decoding the syntax `abs_remainder` in transform-skip residual encoding and decoding. When the CU is encoded in transform-skip mode and the CU is associated with the aforementioned strip header, picture header, PPS, and / or SPS, the rice parameters, represented by signals, are used to determine the codewords used for encoding and decoding the syntax `abs_remainder`.

[0213] According to a sixth aspect of this disclosure, a set of parameters and / or thresholds associated with the codeword determination shown in the first and second aspects are used for transforming the syntax elements of coefficient encoding / decoding and / or transforming the skip residual encoding / decoding. Different sets may be used depending on whether the current block contains luminance residuals / coefficients or chrominance residuals / coefficients. The determined codewords are used as binarized codewords when the syntax elements are encoded / decoded by an entropy codec (e.g., arithmetic encoding / decoding).

[0214] In one example, the codeword of abs_remainder associated with the transform residual codec used in the current VVC is used for both the luma and chroma blocks, but the luma and chroma blocks use different fixed rice parameters. (For example, K1 represents the luma block and K2 represents the chroma block, where K1 and K2 are integers.)

[0215] According to a seventh aspect of this disclosure, the set of parameters and / or thresholds associated with the determination of codewords for syntax elements in transform coefficient encoding / decoding and / or transform skip residual encoding / decoding are represented as signals in a bitstream. Different sets can be represented as signals for luma blocks and chroma blocks. When syntax elements are encoded / decoded using an entropy codec (e.g., arithmetic encoding / decoding), the determined codewords are used as binarized codewords.

[0216] The same approach explained above also applies to escape value encoding and decoding in palette mode, for example, palette_escape_val.

[0217] According to the eighth aspect of this disclosure, different k-order exponent Golomb binarizations can be used to obtain different sets of binary codewords for encoding and decoding escape values ​​in palette mode. In one example, for a given escape sample block, the exponent Golomb parameter used (i.e., the value of k) is based on the block's QP value (denoted as QP). CU The value of parameter k can be determined using the same example shown in Table 6 when obtaining the value of parameter k based on a given QP value for a block. Although four different thresholds (from TH1 to TH4) are listed in this example, and the value can be determined based on these thresholds and QP... CU Five different k values ​​(from K0 to K4) are obtained, but it's worth noting that the number of thresholds is for illustrative purposes only. In practice, the entire QP value range can be partitioned into different numbers of QP value segments using different numbers of thresholds, and for each QP value segment, a different k value can be used to obtain the corresponding binary codeword for encoding and decoding the escape values ​​of the encoded blocks in palette mode. It is also worth noting that the same logic can be implemented differently in practice. For example, the same rice parameter can be obtained using certain equations or lookup tables.

[0218] According to a ninth aspect of this disclosure, the parameter set and / or threshold associated with the codeword determination of the syntax elements of the escape sample are represented as signals in a bitstream. When the syntax elements of the escape sample are encoded or decoded (e.g., arithmetic encoding or decoding) by an entropy codec, the determined codeword is used as a binarized codeword.

[0219] Note that the parameter set and / or threshold can be a complete set or subset of all parameters and thresholds associated with the codeword determination of the syntax elements. The parameter set and / or threshold can be represented by signals at different levels in the video bitstream. For example, these parameter sets and / or thresholds can be represented by signals at the sequence level (e.g., sequence parameter set), picture level (e.g., picture parameter set and / or picture header), stripe level (e.g., stripe header), coding tree unit (CTU) level, or coding unit (CU) level.

[0220] In one example according to this aspect, k-order exponent Golomb binarization is used to determine the codewords for encoding and decoding the palette_escape_val syntax in palette mode, and the value of k is signaled to the decoder in the bitstream. The value of k can be signaled at different levels, for example, it can be signaled in strip headers, picture headers, PPS and / or SPS, etc. When the CU is encoded in palette mode and the CU is associated with the aforementioned strip headers, picture headers, PPS and / or SPS, the signaled exponent Golomb parameter is used to determine the codewords used for encoding and decoding the palette_escape_val syntax.

[0221] Coordination of level mapping for transition skip mode and regular transition mode

[0222] According to the tenth aspect of this disclosure, the same conditions for applying level mapping are used for both the transform skip mode and the regular transform mode. In one example, it is proposed to apply level mapping after the number of context-coded bits (CCB) exceeds the limits of both the transform skip mode and the regular transform mode. In another example, it is proposed to apply level mapping before the number of context-coded bits (CCB) exceeds the limits of both the transform skip mode and the regular transform mode.

[0223] According to the eleventh aspect of this disclosure, the same method for deriving the mapping positions in the level mapping is used for both the transform skip mode and the regular transform mode. In one example, it is proposed that the method for deriving the mapping positions in the level mapping used in the transform skip mode is also applied to the regular transform mode. In another example, it is proposed that the method for deriving the mapping positions in the level mapping used in the regular transform mode is also applied to the transform skip mode.

[0224] According to the twelfth aspect of this disclosure, the same level mapping method is applied to both the transform skip mode and the regular transform mode. In one example, it is proposed that the level mapping function used in the transform skip mode is also applied to the regular transform mode. In another example, it is proposed that the level mapping function used in the regular transform mode is also applied to the transform skip mode.

[0225] Simplified derivation of Rice parameters in residual encoding and decoding

[0226] According to the thirteenth aspect of this disclosure, it is proposed that when encoding and decoding the syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes, simple logic (such as shift or division operations) should be used instead of a lookup table to obtain the rice parameter. According to this disclosure, the lookup table specified in Table 4 can be removed. In one example, the rice parameter cRiceParam is obtained as follows: cRiceParam = (locSumAbs >> n), where n is a positive number, such as 3. It is worth noting that in practice, the same result can be achieved using other different logic, for example, by performing division with a value equal to the power of 2. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0227] Table 7. Decoding Process

[0228]

[0229]

[0230] According to the fourteenth aspect of this disclosure, it is proposed to use fewer adjacent positions to obtain the rice parameter when encoding and decoding the syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes. In one example, it is proposed to use only two adjacent positions to obtain the rice parameter when encoding and decoding the syntax elements of abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0231] Table 8. Decoding Process

[0232]

[0233]

[0234] In another example, it is proposed that when encoding and decoding the syntax elements of abs_remainder / dec_abs_level, only one adjacent position is used to obtain the rice parameter. The corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements that are deleted or removed from the decoding process.

[0235] Table 9. Decoding Process

[0236]

[0237]

[0238] According to the fifteenth aspect of this disclosure, when encoding and decoding the syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes, it is proposed to use different logic to adjust the value of locSumAbs based on the value of baseLevel to derive the rice parameter. In one example, additional scaling and offset operations are applied in the form of "(locSumAbs-baseLevel*5)*alpha+beta". When alpha is 1.5 and beta is 1, the corresponding decoding process based on the VVC draft is as follows.

[0239] Table 10. Decoding Process

[0240]

[0241]

[0242] According to the sixteenth aspect of this disclosure, it is proposed to use Golomb-Rice codes to remove the clipping operation used to obtain the rice parameter in the syntax elements of abs_remainder / dec_abs_level. An example of the decoding process in the VVC draft according to this disclosure is shown below, where strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0243] Table 11. Decoding Process

[0244]

[0245]

[0246] According to this disclosure, an example of the decoding process in the VVC draft is shown below, where strikethrough text indicates steps or elements that are deleted or removed from the decoding process.

[0247] Table 12. Decoding Process

[0248]

[0249]

[0250] According to the seventeenth aspect of this disclosure, when encoding and decoding the syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes, the initial value of locSumAbs is changed from 0 to a non-zero integer to obtain the rice parameter. In one example, an initial value of 1 is assigned to locSumAbs, and the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0251] Table 13. Decoding Process

[0252]

[0253] According to the eighteenth aspect of this disclosure, when encoding and decoding syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes, the rice parameter is obtained by using the maximum value among adjacent positional level values ​​instead of their sum. An example of the corresponding decoding process based on the VVC draft is illustrated below.

[0254] Table 14. Decoding Process

[0255]

[0256] According to the nineteenth aspect of this disclosure, when encoding and decoding syntax elements of abs_remainder / dec_abs_level using Golomb-Rice codes, a rice parameter is proposed based on the relative amplitude of each AbsLevel value at adjacent positions and the base level value. In one example, the rice parameter is obtained based on the number of AbsLevel values ​​at adjacent positions that are greater than the base level. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0257] Table 15. Decoding Process

[0258]

[0259] In another example, the rice parameter is derived from the sum of (AbsLevel - baseLevel) values ​​of adjacent positions where the AbsLevel value is greater than the base level. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0260] Table 16. Decoding Process

[0261]

[0262] According to this disclosure, an example of the decoding process in the VVC draft is shown below, where strikethrough text indicates steps or elements that are deleted or removed from the decoding process.

[0263] Table 17. Decoding Process

[0264]

[0265] Simplified derivation of level mapping positions in residual encoding and decoding

[0266] According to the twentieth aspect of this disclosure, it is proposed to remove QState from the derivation of ZeroPos[n], so that ZeroPos[n] is obtained only from cRiceParam. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements that are deleted or removed from the decoding process.

[0267] Table 18. Decoding Process

[0268]

[0269] According to the twenty-first aspect of this disclosure, ZeroPos[n] is obtained based on the value of locSumAbs. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0270] Table 19. Decoding Process

[0271]

[0272]

[0273] According to the twenty-second aspect of this disclosure, ZeroPos[n] is proposed to be obtained based on the values ​​of AbsLevel at adjacent positions. In one example, ZeroPos[n] is obtained based on the maximum value of AbsLevel[xC+1][yC] and AbsLevel[xC][yC+1]. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0274] Table 20. Decoding Process

[0275]

[0276]

[0277] According to aspect twenty-three of this disclosure, cRiceParam and ZeroPos[n] are derived by maximizing the value of all AbsLevel values ​​at adjacent positions. An example of the corresponding decoding process based on the VVC draft is shown below, where strikethrough text indicates steps or elements deleted or removed from the decoding process.

[0278] Table 21. Decoding Process

[0279]

[0280] The methods explained above also apply to obtaining predCoeff in residual encoding and decoding of transform skip modes. In one example, the variable predCoeff is obtained as follows:

[0281] predCoeff=Max(absLeftCoeff,absAboveCoeff)+1

[0282] Residual encoding and decoding of transform coefficients

[0283] In the present disclosure, to solve the problems pointed out in the "low efficiency of video decoding" section, methods for simplifying and / or further improving the existing designs of residual encoding and decoding are provided. Generally, the main features of the technologies proposed in the present disclosure are summarized as follows.

[0284] First, adjust the rice parameter derivation used under conventional residual encoding and decoding based on the current design.

[0285] Second, change the binary method used under conventional residual encoding and decoding.

[0286] Third, change the rice parameter derivation used under conventional residual encoding and decoding.

[0287] Rice parameter derivation in residual encoding and decoding based on the current design

[0288] According to the twenty-fourth aspect of the present disclosure, it is proposed that in residual encoding and decoding, various rice parameter derivation methods are used to encode and decode certain syntax elements (e.g., abs_remainder / dec_abs_level), and the selection is determined according to certain encoded information of the current block (e.g., quantization parameter or encoding and decoding bit depth associated with TB / CB and / or stripe / profile) and / or according to a new flag associated with the TB / CB / stripe / picture / sequence level (e.g., extended_precision_processing_flag). Different methods can be used to obtain the rice parameter, and some exemplary methods are listed below.

[0289] First, cRiceParam = (cRiceParam << a) + (cRiceParam >> b) + c, where a, b, and c are positive numbers, e.g., {a, b, c} = {1, 1, 0}. It should be noted that in practice, the same result can be achieved with other different logics, e.g., performing multiplication operations with values equal to the nth power of 2.

[0290] Second, cRiceParam = (cRiceParam << a) + b, where a and b are positive numbers, e.g., {a, b} = {1, 1}. It should be noted that in practice, the same result can be achieved with other different logics, e.g., performing multiplication operations with values equal to the nth power of 2.

[0291] Third, cRiceParam = (cRiceParam * a) + b, where a and b are positive numbers, e.g., {a, b} = {1.5, 0}. It should be noted that in practice, the same result can be achieved with other different logics, e.g., performing multiplication operations with values equal to the nth power of 2.

[0292] The following diagram illustrates an example of the corresponding decoding process based on the VVC draft. Changes to the VVC draft are shown in bold and italics in Table 22, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that the same logic can be implemented differently in practice. For example, the same rice parameter can also be obtained from the bit depth value of the current CU / sequence using certain equations or lookup tables.

[0293] Table 22. Decoding Process

[0294]

[0295]

[0296]

[0297] In another example, when the bit depth is greater than or equal to a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is obtained as follows: cRiceParam = (cRiceParam <<a)+(cRiceParam> >b)+c, where a, b, and c are positive numbers, such as 1. The corresponding decoding process based on the VVC draft is shown below. Changes to the VVC draft are shown in bold and italics in Table 23, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that the same logic can be implemented differently in practice. For example, the same rice parameter can also be obtained from the bit depth value of the current CU / sequence using certain equations or lookup tables.

[0298] Table 23. Decoding Process

[0299]

[0300]

[0301] Binary methods in residual encoding and decoding of profiles exceeding 10 bits

[0302] According to the twenty-fifth aspect of the present disclosure, it is proposed to encode and decode certain syntax elements (e.g., abs_remainder / dec_abs_level) using a variable set of binary codewords in residual encoding and decoding, and to determine the selection according to certain encoded information of the current block (e.g., quantization parameter or encoding / decoding bit depth associated with TB / CB and / or stripe / profile) and / or according to a new flag associated with the TB / CB / stripe / picture / sequence level (e.g., extended_precision_processing_flag). Different methods can be used to obtain the variable set of binary codewords, and some exemplary methods are listed below.

[0303] First, use the same procedure for determining the codeword for abs_remainder as that used in the current VVC, but always adopt a selected fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8). According to certain encoded information of the current block (e.g., quantization parameter or encoding / decoding bit depth associated with TB / CB and / or stripe / profile), and / or according to a syntax element associated with the TB / CB / stripe / picture / sequence level (e.g., rice_parameter_value), the fixed value can be different under different conditions. A specific example is shown in Table 24, where TH1 to TH4 are predefined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predefined Rice parameters. It should be noted that the same logic can have different implementation manners in practice. For example, certain equations or look-up tables can also be used to obtain the same Rice parameter from the bit depth values of the current CU / sequence, as shown in Table 24.

[0304] Second, fixed-length binarization.

[0305] Third, truncated Rice binarization.

[0306] Fourth, truncated binary (TB) binarization process.

[0307] Fifth, k-th order exponential Golomb binarization process (EGk).

[0308] Sixth, finite k-th order exponential Golomb binarization.

[0309] Table 24. Determination of Rice parameter based on bit depth

[0310]

[0311]

[0312] In one example, when the new flag (e.g., extended_precision_processing_flag) equals 1, the Rice parameter cRiceParam is fixed to n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value can vary under different conditions. An example of the corresponding decoding process based on the VVC draft is illustrated below. Changes to the VVC draft are shown in bold and italics in Table 25, and strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0313] Table 25. Decoding Process

[0314]

[0315] In another example, it is proposed that when a new flag (e.g., extended_precision_processing_flag) equals 1, only a fixed value is used for the rice parameter when encoding and decoding syntax elements of abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below. Changes to the VVC draft are shown in bold and italics in Table 26, and strikethrough text indicates steps or elements that have been removed or eliminated from the decoding process.

[0316] Table 26. Decoding Process

[0317]

[0318]

[0319] In yet another example, when the bit depth is greater than or equal to a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed at n, where n is a positive number, such as 4, 5, 6, 7, or 8. The fixed value can vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). Changes to the VVC draft are shown in bold and italics in Table 27, and strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0320] Table 27. Decoding Process

[0321]

[0322] In yet another example, it is proposed that when the bit depth is greater than a predetermined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), only a fixed value for the rice parameter is used when encoding and decoding syntax elements of abs_remainder / dec_abs_level. The corresponding decoding process based on the VVC draft is shown below, where TH is a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). Changes to the VVC draft are shown in bold and italics in Table 28, and strikethrough text indicates steps or elements that have been removed or eliminated from the decoding process.

[0323] Table 28. Decoding Process

[0324]

[0325] Derivation of rice parameters in residual encoding and decoding

[0326] According to the twenty-sixth aspect of this disclosure, various rice parameter derivation methods are proposed for encoding and decoding certain syntax elements (e.g., abs_remainder / dec_abs_level) in residual encoding and decoding, and the selection is determined based on certain encoded information of the current block (e.g., quantization parameters or encoding / decoding bit depth associated with TB / CB and / or stripe / profile) and / or based on new flags associated with TB / CB / strip / picture / sequence level (e.g., extended_precision_processing_flag). Different methods can be used to obtain the rice parameters; some exemplary methods are listed below.

[0327] First, a counter is proposed to obtain the rice parameter. The counter is determined based on the values ​​of the encoded coefficients and some encoded information of the current block (e.g., part ID). In a specific example, the rice parameter riceParameter = counter / a, where a is a positive number, such as 4, and it maintains 2 counters (divided by luminance / chrominance). These counters are reset to 0 at the beginning of each stripe. Once encoded, if this is the first coefficient encoded in a sub-TU, the counters are updated as follows:

[0328] if(coeffValue>=(3< <rice))counter++

[0329] if (((coeffValue<<1)<(1<<riceParameter))&&(counter> 0))counter--;

[0330] Second, a shift operation is proposed to be added to the rice parameter derivation in VVC. The shift is determined based on the value of the encoded coefficients. An example of the corresponding decoding process based on the VVC draft is shown below, and the shift is determined according to the counter of Method 1. The changes to the VVC draft are shown in bold and italic in Table 29, and the steps or elements removed or deleted from the decoding process are shown in strikethrough font.

[0331] Table 29. Decoding Process

[0332]

[0333] Third, a shift operation is proposed to be added to the rice parameter derivation in VVC. The shift is determined based on certain encoded information of the current block, such as the codec bit depth associated with the TB / CB and / or stripe profile (e.g., a 14-bit or 16-bit profile). An example of the corresponding decoding process based on the VVC draft is shown below, and the shift is determined according to the counter of Method 1. Changes to the VVC draft are shown in bold and italic font in Table 30, and strikethrough font indicates steps or elements that have been removed or eliminated from the decoding process.

[0334] Table 30. Decoding Process

[0335]

[0336] Transform skipped residual encoding and decoding

[0337] According to the twenty-seventh aspect of this disclosure, it is proposed to use a set of variables of binary codewords to encode and decode certain syntax elements (e.g., abs_remainder) in transform skip residual encoding and decoding, and to determine the selection based on certain encoded information of the current block (e.g., quantization parameters or encoding / decoding bit depth associated with TB / CB and / or stripe / profile) and / or based on new flags associated with TB / CB / strip / picture / sequence level (e.g., extended_precision_processing_flag). Different methods can be used to obtain the set of variables of binary codewords; some exemplary methods are listed below.

[0338] First, use the same procedure for determining the codeword of abs_remainder as that used in the current VVC, but always adopt a selected fixed Rice parameter (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value can be different under different conditions according to certain encoded information of the current block (e.g., quantization parameter or coding / decoding bit depth associated with TB / CB and / or stripe / profile) and / or according to syntax elements associated with the TB / CB / stripe / picture / sequence level (e.g., rice_parameter_value). A specific example is shown in Table 7, where TH1 to TH4 are predefined thresholds satisfying (TH1 < TH2 < TH3 < TH4), and K0 to K4 are predefined Rice parameters. It should be noted that the same logic can have different implementations in practice. For example, certain equations or look-up tables can also be used to obtain the same Rice parameter from the bit depth value of the current CU / sequence, as shown in Table 7.

[0339] Second, fixed-length binarization.

[0340] Third, truncated Rice binarization.

[0341] Fourth, truncated binary (TB) binarization process.

[0342] Fifth, k-th order exponential Golomb binarization process (EGk).

[0343] Sixth, finite k-th order exponential Golomb binarization.

[0344] An example of the corresponding decoding process based on the VVC draft is shown below. Changes to the VVC draft are shown in bold and italic fonts in Table 31, and strikethrough fonts show steps or elements deleted or removed from the decoding process. It should be noted that the same logic can have different implementations in practice. For example, certain equations or look-up tables can also be used to obtain the same Rice parameter.

[0345] Table 31. Decoding process

[0346]

[0347]

[0348] In another example, it is proposed that when a new flag (e.g., extended_precision_processing_flag) equals 1, only a fixed value is used for the rice parameter when encoding and decoding syntax elements of abs_remainder. The corresponding decoding process based on the VVC draft is shown below. Changes to the VVC draft are shown in bold and italics in Table 32, and strikethrough text indicates steps or elements that have been removed or eliminated from the decoding process.

[0349] Table 32. Decoding Process

[0350]

[0351]

[0352] In yet another example, when the new flag (e.g., extended_precision_processing_flag) equals 1, the Rice parameter cRiceParam is fixed to n, where n is a positive number (e.g., 2, 3, 4, 5, 6, 7, or 8). The fixed value can vary under different conditions. The following illustrations show examples of the corresponding decoding process based on the VVC draft. Changes to the VVC draft are shown in bold and italics in Table 33, and strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0353] Table 33. Decoding Process

[0354]

[0355]

[0356] In yet another example, when the bit depth is greater than or equal to a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16), the Rice parameter cRiceParam is fixed at n, where n is a positive number, such as 4, 5, 6, 7, or 8. The fixed value can vary under different conditions. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predefined threshold (e.g., 10, 11, 12, 13, 14, 15, or 16). Changes to the VVC draft are shown in bold and italics in Table 34, and strikethrough text indicates steps or elements removed or deleted from the decoding process.

[0357] Table 34. Decoding Process

[0358]

[0359] In yet another example, a control flag is signaled in the stripe header to indicate whether signaling of the Rice parameter for the transform skip block is enabled or disabled. When the control flag is signaled as enabled, a further signaled syntax element is used for each transform skip strip to indicate the Rice parameter for that strip. When the control flag is signaled as disabled (e.g., set to equal to "0"), no further signaled syntax element is used at a lower level to indicate the Rice parameter for the transform skip strip; instead, a default Rice parameter (e.g., 1) is used for all transform skip strips. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predefined value (e.g., 0, 1, 2). Changes to the VVC draft are shown in bold and italic in Table 35, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that sh_ts_residual_coding_rice_index can be encoded and decoded in different ways and / or may have a maximum value. For example, u(n) using an n-bit unsigned integer or f(n) using an n-bit (from left to right) fixed-pattern bit string written with the left bit first can also be used to encode / decode the same syntax elements.

[0360] Table 35. Strip Header Syntax

[0361]

[0362] A value of 1 for `sh_ts_residual_coding_rice_flag` indicates that `sh_ts_residual_coding_rice_index` may exist in the current slice. A value of 0 for `sh_ts_residual_coding_rice_flag` indicates that `sh_ts_residual_coding_rice_flag` does not exist in the current slice. When `sh_ts_residual_coding_rice_flag` does not exist, it is inferred that the value of `sh_ts_residual_coding_rice_flag` is equal to 0.

[0363] sh_ts_residual_coding_rice_index specifies the rice parameter used in the residual_ts_coding() syntax structure.

[0364] Table 36. Decoding Process

[0365]

[0366]

[0367] Figure 16A method for video decoding is illustrated. For example, this method can be applied to a decoder.

[0368] In step 1610, the decoder may receive a video bitstream. The video bitstream may include encoded video information and information for decoding the encoded video information.

[0369] In step 1612, the decoder may receive control flags at the stripe header level. These control flags can be signals indicating whether the rice parameter is enabled. For example, the control flags can be used to decode encoded video information, and the rice parameter can be used to decode the syntax of abs_remainder and dec_abs_level.

[0370] In step 1614, the decoder may receive at least one syntax element at the stripe header level. This at least one syntax element may be signaled for a transform skip strip and indicate the rice parameter. For example, a signaled syntax element may be used for each transform skip strip to indicate the rice parameter for that strip. In another example, when no signaled syntax element is used at a lower level to indicate the rice parameter for a transform skip strip, a default rice parameter is used for all transform skip strips.

[0371] In step 1616, the decoder may perform entropy decoding on the video bitstream based on control flags and at least one syntax element. For example, the decoder may use control flags and at least one syntax element to obtain quantized coefficient levels and prediction-related information to decode the encoded video information.

[0372] The entropy encoding of the quantization index of the transform / transform skip block can be called transform / transform skip coefficient encoding / decoding.

[0373] In one or more embodiments, the encoder may determine that the residual codec disable flag is equal to 0. The encoder may also signal a residual codec rice flag. The at least one syntax element may include the residual codec rice flag. The encoder may also determine that the residual codec rice flag is equal to 1. The encoder may further signal a residual codec rice index flag. The at least one syntax element includes the residual codec rice index flag.

[0374] In one or more embodiments, the decoder may determine that the residual codec disable flag is equal to 0. The decoder may receive a residual codec rice flag. The at least one syntax element may include the residual codec rice flag. The decoder may also determine that the residual codec rice flag is equal to 1. The decoder may further receive a residual codec rice index flag. The at least one syntax element includes the residual codec rice index flag.

[0375] In yet another example, a control flag is signaled in the sequence parameter set (or in the sequence parameter set range extension syntax) to indicate whether signal transmission of the Rice parameter for the transform skip block is enabled or disabled. When the control flag is signaled as enabled, a syntax element is further signaled for each transform skip strip to indicate the Rice parameter for that strip. When the control flag is signaled as disabled (e.g., set to equal to "0"), no further syntax element is signaled at a lower level to indicate the Rice parameter for the transform skip strip; instead, a default Rice parameter (e.g., 1) is used for all transform skip strips. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predefined value (e.g., 0, 1, 2). Changes to the VVC draft are shown in bold and italic in Table 37, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that sh_ts_residual_coding_rice_idx can be encoded and decoded in different ways and / or may have a maximum value. For example, u(n) using an n-bit unsigned integer or f(n) using an n-bit (from left to right) fixed-pattern bit string written with the left bit first can also be used to encode / decode the same syntax elements.

[0376] Table 37. Sequence Parameter Set (RBSP) Syntax

[0377]

[0378] `sps_ts_residual_coding_rice_present_in_sh_flag` equal to 1 indicates that `sh_ts_residual_coding_rice_idx` may exist in the SH syntax structure referencing SPS. `sps_ts_residual_coding_rice_present_in_sh_flag` equal to 0 indicates that `sh_ts_residual_coding_rice_idx` does not exist in the SH syntax structure referencing SPS. When `sps_ts_residual_coding_rice_present_in_sh_flag` does not exist, it is inferred that the value of `sps_ts_residual_coding_rice_present_in_sh_flag` is equal to 0.

[0379] Table 38. Strip Header Syntax

[0380]

[0381] sh_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax structure.

[0382] Table 39. Decoding Process

[0383]

[0384] In yet another example, for each transform skipped strip, a signal is used to represent a syntax element to indicate the Rice parameter of that strip. An example of the corresponding decoding process based on the VVC draft is illustrated below. Changes to the VVC draft are shown in bold and italics in Table 40, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that sh_ts_residual_coding_rice_idx can be encoded / decoded in different ways and / or may have a maximum value. For example, u(n) using an n-bit unsigned integer or f(n) using an n-bit (from left to right) fixed-pattern bit string written left-bit-first can also be used to encode / decode the same syntax element.

[0385] Table 40. Strip Header Syntax

[0386]

[0387] `sh_ts_residual_coding_rice_idx` specifies the `rice` parameter used in the `residual_ts_coding()` syntax structure. If `sh_ts_residual_coding_rice_idx` does not exist, it is assumed that its value is 0.

[0388] Table 41. Decoding Process

[0389]

[0390] Figure 17 A method for video decoding is illustrated. For example, this method can be applied to a decoder.

[0391] In step 1710, in response to determining that the residual codec disable flag is equal to 0, the decoder may receive the residual codec rice flag. The at least one syntax element includes the residual codec rice flag.

[0392] In step 1712, in response to determining that the residual codec rice flag is equal to 1, the decoder may receive the residual codec rice index flag. The at least one syntax element includes the residual codec rice index flag.

[0393] In yet another example, a control flag is signaled in the picture parameter set range extension syntax to indicate whether signal transmission of the Rice parameter for the transform skip block is enabled or disabled. When the control flag is signaled as enabled, a syntax element is further signaled to indicate the Rice parameter for that picture. When the control flag is signaled as disabled (e.g., set to equal to "0"), no further syntax element is signaled at a lower level to indicate the Rice parameter for the transform skip strip; instead, a default Rice parameter (e.g., 1) is used for all transform skip strips. An example of the corresponding decoding process based on the VVC draft is shown below, where TH is a predefined value (e.g., 0, 1, 2). Changes to the VVC draft are shown in bold and italic in Table 42, and strikethrough text indicates steps or elements removed or deleted from the decoding process. It is worth noting that pps_ts_residual_coding_rice_idx can be encoded and decoded in different ways and / or may have a maximum value. For example, u(n) using an n-bit unsigned integer or f(n) using an n-bit (from left to right) fixed-pattern bit string written with the left bit first can also be used to encode / decode the same syntax elements.

[0394] Table 42. Syntax for Expanding Image Parameter Set Range

[0395]

[0396] A value of 1 for `pps_ts_residual_coding_rice_flag` indicates that `pps_ts_residual_coding_rice_index` may exist in the current image. A value of 0 for `pps_ts_residual_coding_rice_flag` indicates that `pps_ts_residual_coding_rice_idx` does not exist in the current image. When `pps_ts_residual_coding_rice_flag` does not exist, it is assumed that the value of `pps_ts_residual_coding_rice_flag` is equal to 0.

[0397] pps_ts_residual_coding_rice_idx specifies the rice parameter used in the residual_ts_coding() syntax structure.

[0398] Table 43. Decoding Process

[0399]

[0400]

[0401] The methods described above can be implemented using an apparatus comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus can be used in conjunction with other hardware or software components for performing the methods described above. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using the one or more circuits.

[0402] Considering the specification and practice of the disclosure herein, other examples of this disclosure will be apparent to those skilled in the art. This application is intended to cover any changes, uses, or adaptations of the disclosure made in accordance with its general principles, including deviations from this disclosure within the scope of known or customary practice in the art. The specification and examples are intended to be illustrative only.

[0403] It should be understood that this disclosure is not limited to the exact examples shown in the description and figures above, and various modifications and changes can be made without departing from its scope.

[0404] Figure 18 A computing environment 1810 coupled to a user interface 1860 is shown. The computing environment 1810 may be part of a data processing server. The computing environment 1810 includes a processor 1820, memory 1840, and I / O interfaces 1850.

[0405] Processor 1820 typically controls the overall operation of computing environment 1810, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1820 may include one or more processors to execute instructions to perform all or some of the steps described above. Furthermore, processor 1820 may include one or more modules that facilitate interaction between processor 1820 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, GPU, etc.

[0406] Memory 1840 is configured to store various types of data to support the operation of computing environment 1810. Memory 1840 may include predefined software 1842. Examples of such data include instructions for any application or method operating on computing environment 1810, video datasets, image data, etc. Memory 1840 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0407] I / O interface 1850 provides an interface between processor 1820 and peripheral interface modules (such as keyboards, click wheels, buttons, etc.). Buttons may include, but are not limited to, home buttons, start scan buttons, and stop scan buttons. I / O interface 1850 can be coupled to encoders and decoders.

[0408] In some embodiments, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs, such as those included in memory 1840, executable by processor 1820 in computing environment 1810, for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0409] The non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the video encoding / decoding method described above.

[0410] In some embodiments, the computing environment 1810 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0411] The description of this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.

[0412] Examples have been chosen and described to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure and to best utilize the basic principles and various embodiments with modifications suitable for the intended particular use. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of this disclosure.

Claims

1. A method for video decoding, comprising: The decoder receives control flags at the strip header level and control flags at the sequence picture set SPS level, wherein the control flags at the strip header level and the control flags at the SPS level are used to determine whether at least one syntax element related to the Rice parameter is signaled for a transform skip strip. In response to determining that the at least one syntax element is represented by a signal, the decoder receives the at least one syntax element at the stripe header level, wherein the at least one syntax element is used to determine the Rice parameter; and The decoder performs entropy decoding on the video bitstream based on the strip header-level control flags, the SPS-level control flags, and at least one syntax element.

2. The method as described in claim 1, in, The control flag at the SPS level is the residual encoding / decoding Rice flag; Among them, the control flag at the strip header level is the residual encoding / decoding disable flag, and The at least one syntax element at the stripe header level received by the decoder includes: In response to determining that the residual codec disable flag is equal to 0 and the residual codec Rice flag is equal to 1, a residual codec Rice index flag is received, wherein the at least one syntax element includes the residual codec Rice index flag.

3. The method as described in claim 2, wherein, When the residual codec Rice flag is equal to 1, the residual codec Rice flag indicates that the residual codec Rice index flag exists in the current stripe.

4. The method of claim 2, wherein, When the residual codec Rice flag is equal to 0, the residual codec Rice flag indicates that the residual codec Rice index flag does not exist in the current stripe.

5. The method of claim 4, further comprising: In response to determining that the residual codec Rice flag does not exist, it is inferred that the value of the residual codec Rice flag is equal to 0.

6. The method of claim 2, further comprising: In response to determining that the residual codec Rice flag is equal to 1, the transform skip flag is equal to 1, and the residual codec disable flag is equal to 0, the Rice parameter is set to the value of the residual codec Rice index flag plus a predefined value.

7. The method of claim 2, further comprising: In response to determining that the transform skip flag is equal to 1 and the residual codec disable flag is equal to 0, the Rice parameter is set to the value of the residual codec Rice index flag plus 1.

8. A computing device, comprising: One or more processors; as well as A memory configured to store instructions executable by the one or more processors and a video bitstream to be decoded, wherein the one or more processors, when executing the instructions, are configured to perform the method as described in any one of claims 1 to 7 using the video bitstream.

9. A computer-readable storage medium storing a plurality of programs for execution by a computing device having one or more processors and a video bitstream to be decoded, wherein, When executed by the one or more processors, the plurality of programs cause the computing device to perform the method as described in any one of claims 1 to 7 using the video bitstream.

10. A computer program product comprising a plurality of programs, which, when executed by one or more processors, cause the one or more processors to perform the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Coefficient level coding in video coding

    CN107925764A

  • Variable Length Coding and Decoding Methods and Devices for Grouped Pixels

    US20170318314A1